CASE 21 / POC / AEM CONTENT / RAG ORCHESTRATION

AEM + RAG Services on Kubernetes.

PoC design / architecture exploration — not a production deployment

Explore a proof-of-concept architecture that connects governed AEM Content Fragments to retrieval and an enterprise assistant. Kubernetes supplies workload boundaries and rollout controls; the PoC must still establish permission safety, content freshness and answer quality.

Representative engineering design, not a verified client delivery record. Implementation steps and validation checks describe the proposed approach; no measured results are claimed.

CONTEXTRepresentative engineering design
TECHNICAL FOCUSAEM RAG Architecture with Kubernetes Services
STATUSPoC / architecture exploration
CONTENT TO ASSISTANT / POC
  1. Content Fragments
  2. GraphQL / API
  3. Ingestion
  4. Retrieval index
  5. RAG service
  6. LLM endpoint
  7. Assistant

Conceptual dependency flow. AEM and the LLM endpoint remain external; Kubernetes orchestrates ingestion, retrieval and RAG service workloads.

01 / CONTEXT

Use governed content as a source, not a guarantee

Structured AEM content offers fields and provenance that can help build retrievable passages. This proposed PoC studies whether that structure supports useful, source-backed answers; it does not assume an LLM will preserve the source's meaning or access restrictions.

02 / PROBLEM

Orchestration does not establish trustworthy retrieval

A technically healthy RAG service can retrieve stale content, omit the relevant passage or disclose information the caller cannot access. Those are application and governance failures that Pod health checks cannot detect.

03 / ARCHITECTURE

Separate ingestion from interactive answering

An ingestion workload fetches approved content through GraphQL or APIs, transforms it into versioned passages and updates a retrieval index. Interactive services authenticate callers, retrieve permitted evidence, build a bounded prompt and call an external LLM endpoint before returning an answer with source references.

04 / ENGINEERING APPROACH

Carry identity and provenance through the pipeline

Attach source URL, content identifier, revision and access metadata to passages. Enforce authorization before content reaches the model, propagate deletions and permission changes, and treat retrieved text as untrusted data rather than executable instructions.

  • Use stable identifiers and resumable ingestion so retries do not duplicate passages.
  • Separate service identities and network access for ingestion, retrieval and model calls.
  • Version chunking, embedding inputs, prompt templates and model configuration for reproducible evaluation.
05 / DEPLOYMENT FLOW

Evaluate behavior before promoting a service revision

Run a fixed question set against the candidate retrieval and prompt configuration, then inspect citations, abstentions and permission failures. Deploy compatible service and index revisions together, with a rollback plan that accounts for their shared data contract.

POC VALIDATION LOOP
  1. Approved content
  2. Ingest
  3. Index
  4. Evaluate
  5. Deploy candidate
  6. Review answers

Evaluation is proposed work. No answer-quality benchmark or production result is asserted.

06 / OPERATIONAL CONSIDERATIONS

Scale ingestion and answering independently

Batch ingestion and interactive requests have different latency and concurrency needs. Bound model-call budgets and retries, record retrieval timing and source revisions, and avoid storing sensitive prompts or answers in general-purpose logs.

07 / TRADEOFFS

Freshness, isolation and platform overhead

A service boundary permits independent scaling and releases but introduces more failure points and operational cost. Before choosing Kubernetes, compare a simpler managed execution model; a small PoC may not justify a cluster operating model.

Continue with the complementary design in Headless AEM: GraphQL and Sling Model Exporter Delivery Boundaries.

08 / WHAT I WOULD VALIDATE IN PRODUCTION

Define evidence before considering production use

The PoC's output should be an evaluation record and a decision about suitability, not a claim that deployment establishes trustworthy AI. Production adoption would require explicit security, content and operational acceptance.

  • Ask the same question as differently authorized users and inspect retrieved passages as well as final answers.
  • Delete or restrict an AEM fragment and verify it no longer reaches prompts after the defined propagation window.
  • Test unsupported questions, conflicting sources and injected instructions inside retrieved content.
  • Simulate LLM timeouts, ingestion interruption and incompatible index revisions; verify bounded failure and recovery.
09 / LESSONS

Treat answer quality as a release concern

Orchestration supports repeatability, but the decisive artifact is the evaluation evidence. A credible AI architecture makes permission handling, provenance and failure behavior as reviewable as its deployment manifests.

REFERENCE MATERIAL

Technical references

These sources document product behavior. The design and validation approach above are engineering proposals, not claims made by the vendors.

NEXT STEPS

Building or modernizing an AEM platform?

This representative case study explores technical trade-offs and architectural decisions for a specific engineering scenario. If you are planning a similar migration, modernization, or integration, let's discuss the engineering approach.

Start a conversation
TECHNOLOGY STACK
KubernetesAEM Content FragmentsGraphQLDockerRetrievalLLM APIsObservability