﻿id	summary	reporter	owner	description	type	status	priority	milestone	component	version	severity	resolution	keywords	cc	focuses
66207	Shortcodes: Update in harmony with HTML API	dmsnell		"WordPress’ Shortcode system has been gradually de-ephasized, but remains in use by many plugins, themes, and Core itself. There are many ways that it is ripe for updating, particularly in harmony with the HTML API.

 - It needs a performance overhaul to replace memory-heavy and inefficient processor.
 - It needs to harmonize with HTML parsing now that WordPress has the ability to understand HTML content.
 - Ambiguities in the grammar needs better explicit documentation. These ambiguities are results of the PCRE patterns which //define// how Shortcodes are parsed.
 - The multiple ways of finding and processing Shortcodes need to be combined into a convenient and single semantic API.

This is a tracking issue to organize the long-term project.

== Notes

=== Parsing is extremely dynamic.

The parsing syntax for Shortcodes is itself dependent on the runtime system when the request is run and when functions like `do_shortcodes()` is called. There may be historic conflation between “only process registered shortcodes” and “only recognize syntax based on registered shortcodes.” This is because //only registered Shortcodes are recognized as Shortcode syntax//.

{{{#!php
<?php

add_shortcode( 's0', fn () => '' );
echo do_shortcodes( 'a[s0]b[s1]c' );
// 'ab[s1]c'
}}}

This dynamism can create surprising behavior, for example, because escaped Shortcodes are not recognized for non-registered Shortcodes. In the following snippet, one //might// expect that `[s0]` should not be rendered because it appears to be part of the content of the non-registered `[s1]` Shortcode, which itself is escaped, meaning that the entire input string would come out unaltered by `do_shortcodes()`.

{{{#!php
<?php

add_shortcode( 's0', fn () => '' );
echo do_shortcodes( '[[s1]a[s0]b[/s1]]' );
// '[[s1]ab[/s1]]'
}}}

This dynamism also implies that one cannot “find” the Shortcodes in an HTML document in isolation. In fact, a given document may have a different set of Shortcodes in it based on who is asking, when, and proverbially what direction the wind is blowing.

=== Parsing only matches positive assertive patterns.

Due to the way that Shortcode handling is based on PCRE patterns and functions, there is no inherent error-handling for malformed Shortcode syntax. Added to this, the matcher is written as a single PCRE pattern with several noteworthy characteristics;

 - Escaping brackets are captured on either side, regardless if they are balanced. Shortcodes which were meant to be escaped can be swallowed because of a missing escaping bracket at close.
 - Whether a Shortcode expects a closing tag is based on whether the closing tag appears in the same document. Non-registered Shortcodes can interact unexpectedly with closing tags, creating them when they were not intended to exist.
 - Rules for parsing switch between affirmative lists (e.g. “a tag name may contain these characters”) and negative lists (e.g. “to parse a tag name, match until these characters appear”) and this leaves room for different code to have different understandings of where Shortcodes exist.
 - Shortcodes without closing tags, or Shortcodes whose closing tags appear far later in a document, cause an undue performance burden during parsing.

{{{#!php
<?php
$input = 'a[s0]b[s1]c[/s0]d[/s1]e';
add_shortcode( 's1', fn ( $attrs, $content ) => ""[s1:{$content}]"" );
do_shortcodes( $input );
add_shortcode( 's0', fn ( $attrs, $content ) => ""[s0:{$content}]"" );
do_shortcodes( $input );

//   input: a[s0]b[s1]c[/s0]d[/s1]e
//      s1: a[s0]b[s1:c[/s0]d]e
// s0 + s1: a[s0:b[s1]c]d[/s1]e
}}}


The documentation thankfully notes some of the wishy-washy nature of Shortcode handling.

 - [https://codex.wordpress.org/Shortcode_API Shortcode API]
 - [https://codex.wordpress.org/Shortcode Shortcode]

This helpfully acknowledges that defining behavior is difficult given that most of the existing behaviors are emergency due to the way the PCRE pattern happened to have been written, but it misses out on being helpful in clarifying expectations for how given Shortcode syntax will be handled. It fails to build an intuition on Shortcode handling.

=== Behavior has changed over time.

There have been changes in how Shortcodes are processed and the documentation makes clear that not all existing behaviors are guaranteed. Almost all variation occurs in abnormal cases, where markup is malformed or where there is valid-looking Shortcode syntax for non-registered Shortcodes.

Notably, an attempt was made to only parse Shortcodes within HTML syntax tokens. Unfortunately, at the time this appeared, WordPress had no reliable mechanism for parsing HTML and so a number of extant behaviors remain as emergent from those decisions.

Enforcing a new intentional set of behaviors can clarify and simplify how Shortcodes work in WordPress.

=== Finding potential Shortcode closers first be much more efficient.

Shortcode closing tags are only matched with very specific syntax. Yet, when reaching an opening tag, it’s required to scan, potentially, to the end of the document to find out //if// a closing tag exists. This occurs for every opening Shortcode tag.

However, a quick pass through the document can find a set of candidate Shortcode closing tags, and as opening tags are found, a rapid search through this candidate list can return matching closers. Advancing past a candidate closing Shortcode purges the set of candidates, maintaining search efficiency.

== Tasks

 - [ ] Replace call to `wp_kses_hair_parse()` with HTML API parsing.
 - [ ] Rebuild parser with efficient streaming parser.
 - [ ] Define changes in behavior to normalize processing.

== Related work

 - #50683 Rebuild the Shortcode parser."	enhancement	new	normal	Future Release	Shortcodes	trunk	normal		has-patch has-unit-tests		
