AEM ENGINEERING NOTES / 003

Why Is This AEM Query Traversing the Repository?

Oak & Querying · Diagnostic guide

Learn why AEM Oak queries fall back to repository traversal, how to inspect query plans, understand index selection, and design queries and indexes for performance.

Mohammed Boudoun · Published

In this engineering note

01 — Introduction

Repository traversal is not just a slow query problem. It usually signals a mismatch between the query and the available Oak indexes. The query is correct and returns the right nodes — but it gets there by walking content rather than consulting an index, and that difference is the whole performance story.

On a development instance with a few thousand nodes, a traversing query looks fine. On a production author or publish instance holding millions of nodes, the same query reads a large fraction of the repository on every execution. The result is identical; the cost is not. Traversal is the runtime telling you it could not find an index that covers the query you asked.

This article is a diagnostic procedure for AEM developers, backend engineers and Oak practitioners: how to recognise traversal, read the execution plan, understand how Oak selects an index, and then fix the real problem by refining either the query, the index, or the search architecture behind it.

02 — What is query traversal?

A JCR query in AEM is parsed by Oak's query engine, which evaluates the available indexes, estimates the cost of each, and picks the cheapest plan. If a suitable index exists, Oak reads matching entries directly from it. If no index covers the query, Oak falls back to traversal: it walks nodes under the query's path scope and evaluates the restrictions in memory.

Traversal is a correctness guarantee, not a failure mode — Oak will always return the right result, even with no usable index. That is precisely why it is dangerous: nothing breaks, nothing errors, and the query keeps working as content grows. The cost scales with the number of nodes visited, so a query that traverses a subtree grows linearly with that subtree while an indexed query stays roughly flat.

The practical rule: traversal on a small, tightly path-restricted subtree can be acceptable; traversal across a large or unbounded scope is a latent production incident. The goal of diagnosis is to know which one you have.

  • Indexed query: Oak reads matching entries from an index — cost tied to result size
  • Traversal: Oak walks nodes under the path scope — cost tied to repository size
  • Both return correct results; only the cost model differs
  • Traversal cost grows with content, so the problem appears over time, not at launch

03 — How Oak chooses an index

Oak does not guess. Each index implementation inspects the query and returns an estimated cost; the query engine selects the lowest-cost plan that can answer it. Understanding what feeds that estimate is what lets you design queries an index can actually serve.

The query shape drives everything. Property restrictions (equality on an indexed property), path restrictions (a specific subtree rather than the whole repository), node type constraints, and fulltext predicates each map to different index capabilities. A property index answers equality and range on a declared property; a Lucene index answers fulltext and richer property combinations; neither helps if the query restricts on a property nothing indexes.

When no index can satisfy the restrictions, every index returns an effectively infinite cost and the traversal plan wins by default. Reading traversal as 'Oak chose the only option left' is more accurate than 'Oak chose badly'.

  • Query shape: which restrictions are present and how selective they are
  • Property restrictions: equality and range on declared properties
  • Path restrictions: scope the query to a subtree instead of the whole tree
  • Node types: nt:/sling: type constraints an index can key on
  • Fulltext constraints: contains() predicates served by a Lucene index
  • Index cost: the estimate each index returns; the engine picks the cheapest

04 — Common causes of traversal

Most traversal reduces to a small set of recurring causes. Work through them against the specific query that is slow, not queries in general.

  • No property index exists for the property being restricted on
  • A Lucene index definition exists but does not declare the queried property
  • The query restricts on an unindexed or dynamically-named property
  • Leading-wildcard or open-ended LIKE patterns that no index can seek on
  • Broad path scope — querying from / or a very large subtree
  • OR-heavy queries where each branch needs a different index
  • Functions or expressions on a property that prevent a direct index lookup
  • Node type assumptions that do not match the indexed type
  • A stale, disabled, or misconfigured index that is not actually serving the query
  • Missing path restriction, so even a selective property scan widens into a walk

05 — Query examples

The clearest way to see the effect is to compare a traversing query with an indexed alternative that returns the same nodes. Two patterns dominate: adding a path restriction, and restricting on a property that an index actually covers.

In the QueryBuilder form below, the first query scans everything of a type with no path scope; the second narrows the scope and restricts on an indexed property, which lets Oak seek instead of walk. The SQL2 pair shows the same principle — a leading wildcard that cannot use an index versus an equality predicate that can.

// QueryBuilder — traversing: no path scope, unindexed property
type=cq:Page
property=jcr:content/customStatus
property.value=approved

// QueryBuilder — improved: path scope + indexed property
path=/content/site/en
type=cq:Page
property=jcr:content/customStatus
property.value=approved
// (customStatus declared in a property or Lucene index)

-- SQL2 — traversing: leading wildcard cannot seek an index
SELECT * FROM [cq:Page] AS p
WHERE ISDESCENDANTNODE(p, '/content/site')
  AND p.[jcr:content/title] LIKE '%report%'

-- SQL2 — improved: equality on an indexed property
SELECT * FROM [cq:Page] AS p
WHERE ISDESCENDANTNODE(p, '/content/site/en')
  AND p.[jcr:content/category] = 'report'

06 — Explain the query plan

Do not guess whether a query traverses — ask Oak. The query engine can explain its plan, exposing which index it selected, the estimated cost, and whether it fell back to traversal. In AEM this is available through the query tooling in the operations console and via the EXPLAIN prefix in SQL2.

Prefix a SQL2 statement with EXPLAIN to return the chosen plan instead of results. The plan names the index that will serve the query; a plan that reads [nt:base] as a traversal, or warns about traversal, is the signal to stop and look at index coverage. A healthy plan names a specific property or Lucene index and a bounded path.

Read the plan for three things: the selected index, the estimated cost or node count, and the path restriction. If the cost estimate is large or the selected 'index' is really a traversal over a base type, the query is not covered — regardless of how fast it happens to run on your current content volume.

EXPLAIN SELECT * FROM [cq:Page] AS p
WHERE ISDESCENDANTNODE(p, '/content/site/en')
  AND p.[jcr:content/category] = 'report'

Look for, in the returned plan:
  - selected index        e.g. lucene:cqPageLucene(/oak:index/...)
  - estimated cost        lower is better; very large => likely traversal
  - traversal warning     'no proper index' / traversal over [nt:base]
  - path restriction      a bounded subtree, not /

07 — Oak indexes in AEM

AEM ships on Oak with two index families you work with most: property indexes and Lucene indexes. Property indexes are synchronous and best for exact, selective restrictions on a single declared property — fast to maintain, narrow in scope. Lucene indexes are asynchronous and cover fulltext plus richer combinations of properties, at the cost of a short indexing lag and heavier definitions.

Index definitions live under /oak:index as content, which is what makes them deployable and version-controllable alongside the application. A custom index definition declares the properties it indexes, the node types it applies to, and whether it is synchronous or async. Async indexing means a newly written node becomes queryable once the background indexer processes it — usually seconds, but not instantaneous.

Be deliberate about out-of-the-box indexes. Modifying a shipped index definition can affect queries you did not write and complicate upgrades, since Adobe may change those definitions between versions. Prefer a well-scoped custom index for application-specific queries over editing an OOTB one, and treat every index change as something to validate rather than assume.

  • Property index: synchronous, ideal for selective equality/range on one property
  • Lucene index: asynchronous, covers fulltext and multi-property queries
  • Index definitions are content under /oak:index — deployable and reviewable
  • Async indexing introduces a short lag before new content is queryable
  • Avoid editing OOTB index definitions; prefer scoped custom indexes

08 — AEM upgrade / migration angle

Traversal problems often appear without the query changing at all. A query that was indexed and fast can start traversing after an upgrade, a migration, or simply enough content growth — and because nothing errors, the regression is easy to miss until latency climbs. This is why query and index behaviour belong in the validation plan for any AEM 6.5 LTS modernization or platform migration.

Upgrades change Oak versions and sometimes OOTB index definitions. An index your query relied on may be redefined, renamed, or have different property coverage; a custom index carried over from an older line may no longer match how the new Oak plans the query. Custom code changes that alter a query's shape — a new restriction, a reordered predicate, a changed node type — can likewise move it off its index.

The underappreciated driver is content volume. A query that traverses a small subtree is invisible in testing and painful in production, so the failure often surfaces months after a release as the repository fills. Validate queries against representative content volumes, not the handful of nodes on a fresh environment.

  • Oak version changes can alter index selection and cost estimation
  • Upgrades may redefine or rename OOTB index definitions
  • Custom index definitions may no longer match a newer Oak's planning
  • Code changes that reshape a query can move it off its index
  • Content growth turns a tolerable subtree walk into a production problem

09 — Search vs. repository query

Not every search belongs in Oak. JCR queries are the right tool for structured, selective retrieval — find the pages of this type under this path with this property value. They are the wrong tool for open-ended functional search: relevance ranking, faceting, typo tolerance, synonyms, aggregations and free-text exploration across large corpora.

When requirements drift toward that functional-search shape, forcing them through Oak leads to ever-broader Lucene indexes and queries that are hard to keep off traversal. A dedicated search platform such as Elasticsearch is built for exactly those access patterns, which is the architecture behind an Elasticsearch-backed AEM search: AEM remains the governed content source, content is indexed into the search system, and the application queries that system for relevance-driven results.

This is an architecture decision, not a workaround. Keep selective, structured lookups in Oak where they are cheap and transactional; move relevance-ranked, faceted, high-cardinality search to a system designed for it. Trying to serve the second category from raw JCR queries is a common root cause of unavoidable traversal.

10 — Performance troubleshooting flow

A slow query has a short, repeatable investigation. The point of the flow below is to replace guessing — adding indexes speculatively, or broadening a Lucene definition until the symptom disappears — with reading the plan and acting on what it names.

Start from the query and the plan, not from the index. Only once you know which index was selected (or that none was) do you decide whether to refine the query, extend an index, or move the use case to a search platform. Then validate the change at a representative content volume, because that is where traversal actually hurts.

Slow query
   |
Inspect the query (shape, restrictions, path scope)
   |
Check the execution plan (EXPLAIN / query tool)
   |
Identify the selected index
   |
Traversal?  --- no --> cost acceptable at scale? --> done
   | yes
Review index coverage (property / Lucene definition)
   |
Refine the query OR the index (or move to a search platform)
   |
Validate on representative content volume

11 — Anti-patterns

Most traversal remediation goes wrong in predictable ways. Each of these makes a symptom disappear locally while degrading the platform or hiding the real cause.

  • Creating a new index for every query instead of reshaping the query to fit existing coverage
  • Broad, catch-all custom Lucene indexes that index far more than any query needs
  • Querying the full repository from / instead of restricting to a known subtree
  • Ignoring path restrictions because 'it is fast enough' on current content
  • Solving functional, relevance-driven search with raw JCR queries when a search platform fits better
  • Validating only on a small environment, so traversal never shows up before production

12 — Production considerations

Index and query changes are deployments with real operational cost, and they behave differently in production than on a laptop. Treat them with the same rigour as any platform change, and tie their validation to production observability so a regression surfaces as a metric rather than a support ticket.

Content scale and query frequency together decide impact: a moderately expensive query executed on every page render is a bigger problem than an expensive query run occasionally. Async Lucene indexing means a newly deployed index is not instantly populated — a reindex can be expensive on a large repository and must be planned, not triggered casually during peak load.

Validate changes safely: confirm the new plan with EXPLAIN, reindex in a controlled window, and verify against representative content volume before trusting results from a sparse environment. The goal is to know the query is indexed and bounded before it meets production traffic, not to discover it afterward.

  • Weigh query cost against execution frequency, not cost alone
  • Async indexing and reindex cost scale with repository size — plan the window
  • Confirm the post-change plan with EXPLAIN before relying on it
  • Validate on representative content volume, not a sparse instance

13 — My perspective

Treat the query and the index as one design problem. A query is only as good as the index behind it, and an index only earns its cost if real queries use it. Writing them separately — a query here, an index added reactively later — is how repositories accumulate both traversal and unused indexes at the same time.

In practice I read the plan first, every time. The plan converts 'this feels slow' into 'this query selected no index and walks /content' — a specific, fixable statement. From there the decision is honest: tighten the path and restriction, extend a scoped index, or accept that the use case is functional search and belongs in a system built for it.

The habit that pays off is validating at scale. Traversal is invisible on a fresh instance and unmistakable on a full one, so the engineering value is in testing against representative content before the regression reaches users — and in designing the query and index together so it never traverses in the first place.

14 — Conclusion

Traversal is a signal, not a verdict. It tells you the query and the available indexes have diverged — so the fix is rarely 'add an index and move on' and usually 'revisit the query shape, the index design, the content model, or the search architecture' until the two fit again.

Read the plan, name the selected index, and decide deliberately between refining the query, extending an index, and moving functional search to a dedicated platform. Done that way, a query that quietly walked millions of nodes becomes a bounded, indexed lookup — and stays one as the repository grows.

References

Related engineering work