Make WordPress Core

Opened 6 hours ago

Last modified 6 hours ago

#66207 new enhancement

Shortcodes: Update in harmony with HTML API

Reported by: dmsnell Owned by:
Priority: normal Milestone: Future Release
Component: Shortcodes Version: trunk
Severity: normal Keywords: has-patch has-unit-tests
Cc: Focuses:

Description

WordPress’ Shortcode system has been gradually de-ephasized, but remains in use by many plugins, themes, and Core itself. There are many ways that it is ripe for updating, particularly in harmony with the HTML API.

  • It needs a performance overhaul to replace memory-heavy and inefficient processor.
  • It needs to harmonize with HTML parsing now that WordPress has the ability to understand HTML content.
  • Ambiguities in the grammar needs better explicit documentation. These ambiguities are results of the PCRE patterns which define how Shortcodes are parsed.
  • The multiple ways of finding and processing Shortcodes need to be combined into a convenient and single semantic API.

This is a tracking issue to organize the long-term project.

Notes

Parsing is extremely dynamic.

The parsing syntax for Shortcodes is itself dependent on the runtime system when the request is run and when functions like do_shortcodes() is called. There may be historic conflation between “only process registered shortcodes” and “only recognize syntax based on registered shortcodes.” This is because only registered Shortcodes are recognized as Shortcode syntax.

<?php

add_shortcode( 's0', fn () => '' );
echo do_shortcodes( 'a[s0]b[s1]c' );
// 'ab[s1]c'

This dynamism can create surprising behavior, for example, because escaped Shortcodes are not recognized for non-registered Shortcodes. In the following snippet, one might expect that [s0] should not be rendered because it appears to be part of the content of the non-registered [s1] Shortcode, which itself is escaped, meaning that the entire input string would come out unaltered by do_shortcodes().

<?php

add_shortcode( 's0', fn () => '' );
echo do_shortcodes( '[[s1]a[s0]b[/s1]]' );
// '[[s1]ab[/s1]]'

This dynamism also implies that one cannot “find” the Shortcodes in an HTML document in isolation. In fact, a given document may have a different set of Shortcodes in it based on who is asking, when, and proverbially what direction the wind is blowing.

Parsing only matches positive assertive patterns.

Due to the way that Shortcode handling is based on PCRE patterns and functions, there is no inherent error-handling for malformed Shortcode syntax. Added to this, the matcher is written as a single PCRE pattern with several noteworthy characteristics;

  • Escaping brackets are captured on either side, regardless if they are balanced. Shortcodes which were meant to be escaped can be swallowed because of a missing escaping bracket at close.
  • Whether a Shortcode expects a closing tag is based on whether the closing tag appears in the same document. Non-registered Shortcodes can interact unexpectedly with closing tags, creating them when they were not intended to exist.
  • Rules for parsing switch between affirmative lists (e.g. “a tag name may contain these characters”) and negative lists (e.g. “to parse a tag name, match until these characters appear”) and this leaves room for different code to have different understandings of where Shortcodes exist.
  • Shortcodes without closing tags, or Shortcodes whose closing tags appear far later in a document, cause an undue performance burden during parsing.
<?php
$input = 'a[s0]b[s1]c[/s0]d[/s1]e';
add_shortcode( 's1', fn ( $attrs, $content ) => "[s1:{$content}]" );
do_shortcodes( $input );
add_shortcode( 's0', fn ( $attrs, $content ) => "[s0:{$content}]" );
do_shortcodes( $input );

//   input: a[s0]b[s1]c[/s0]d[/s1]e
//      s1: a[s0]b[s1:c[/s0]d]e
// s0 + s1: a[s0:b[s1]c]d[/s1]e

The documentation thankfully notes some of the wishy-washy nature of Shortcode handling.

This helpfully acknowledges that defining behavior is difficult given that most of the existing behaviors are emergency due to the way the PCRE pattern happened to have been written, but it misses out on being helpful in clarifying expectations for how given Shortcode syntax will be handled. It fails to build an intuition on Shortcode handling.

Behavior has changed over time.

There have been changes in how Shortcodes are processed and the documentation makes clear that not all existing behaviors are guaranteed. Almost all variation occurs in abnormal cases, where markup is malformed or where there is valid-looking Shortcode syntax for non-registered Shortcodes.

Notably, an attempt was made to only parse Shortcodes within HTML syntax tokens. Unfortunately, at the time this appeared, WordPress had no reliable mechanism for parsing HTML and so a number of extant behaviors remain as emergent from those decisions.

Enforcing a new intentional set of behaviors can clarify and simplify how Shortcodes work in WordPress.

Finding potential Shortcode closers first be much more efficient.

Shortcode closing tags are only matched with very specific syntax. Yet, when reaching an opening tag, it’s required to scan, potentially, to the end of the document to find out if a closing tag exists. This occurs for every opening Shortcode tag.

However, a quick pass through the document can find a set of candidate Shortcode closing tags, and as opening tags are found, a rapid search through this candidate list can return matching closers. Advancing past a candidate closing Shortcode purges the set of candidates, maintaining search efficiency.

Tasks

  • [ ] Replace call to wp_kses_hair_parse() with HTML API parsing.
  • [ ] Rebuild parser with efficient streaming parser.
  • [ ] Define changes in behavior to normalize processing.

Related work

  • #50683 Rebuild the Shortcode parser.

Change History (1)

This ticket was mentioned in ​PR #13822 on ​WordPress/wordpress-develop by ​@dmsnell.


6 hours ago
#1

  • Keywords has-patch has-unit-tests added

Trac ticket: Core-66207.

This patch explores an efficient streaming interface for Shortcode processing. It’s currently _only_ an exploration and not intended as a real proposal at this stage.

Synopsis

  • Runs a forward pass to identify potential closing tags and eliminate algorithmic waste.
  • Relies on WP_Token_Map when determining if a potential Shortcode tag ought to be recognized (based on which Shortcodes are registered).
  • May stop processing at any point (after the candidate closing tags have been found).

Current limitations

  • Searches without regard for HTML structure.
  • Does not parse Shortcode attributes.
  • Does not expose information about the matched Shortcodes.
Note: See TracTickets for help on using tickets.