There are several strong resources, but they use overlapping terminology differently. The most useful approach is to take the concepts, not adopt any one taxonomy wholesale.
| Resource | Key idea | Why it is useful for this mental model |
|---|---|---|
| C4 Model — Simon Brown C4 Model | Explicit levels of abstraction: System → Container → Component → Code, plus dynamic and deployment diagrams. | Probably the best day-to-day approach for software teams. It directly attacks mixed abstraction levels and ambiguous boxes/arrows. (C4 model) |
| Software Systems Architecture — Rozanski & Woods Viewpoints and Perspectives | Separates viewpoints such as Functional, Information, Concurrency, Development, Deployment and Operational from cross-cutting perspectives such as security, performance and resilience. | This is probably the closest formalisation to the model we’ve developed. (Viewpoints and Perspectives) |
| arc42 arc42 overview | Practical architecture documentation organised around context, building blocks, runtime, deployment, decisions, quality requirements and risks. | Excellent bridge from architecture theory to something a team can actually maintain. (arc42) |
| SEI — Views and Beyond Views and Beyond collection | Architecture is a set of relevant views of system structures, plus information that applies across views. | Strongest rigorous treatment of why one architecture diagram cannot describe a system adequately. (Software Engineering Institute) |
| Kruchten — 4+1 View Model Original 4+1 paper | Logical, Process, Development and Physical views, validated by scenarios. | Foundational explanation of why static logical structure, runtime behaviour and physical deployment are different things. (Concordia ENCS Users) |
| ISO/IEC/IEEE 42010:2022 ISO 42010 | Formal vocabulary around architecture descriptions, stakeholders, concerns, viewpoints and model kinds. | Useful when you want precise terminology rather than an opinionated diagramming method. (ISO) |
| ArchiMate 101 — The Open Group community ArchiMate 101 | Explicit Business, Application and Technology layers plus viewpoints spanning them. | Useful for enterprise-scale systems where software needs to be related to business capabilities and infrastructure. (GitLab) |
| Michael Nygard — Architecture Decision Records Documenting Architecture Decisions | Record context → decision → consequences, not just the resulting architecture. | Adds the missing “why did we build it this way?” dimension. (Cognitect.com) |
For what you’re trying to achieve, I’d particularly recommend:
1. Rozanski & Woods. Their distinction between viewpoints and perspectives improves the model we’ve been discussing.
They essentially say:
Viewpoint
What structural aspect are we examining?
Functional
Information
Concurrency
Development
Deployment
Operational
Perspective
What system quality are we checking across those views?
Performance
Security
Availability
Resilience
Evolution
...
A security concern isn’t really another architecture level. It cuts across your application, data, runtime and deployment views. The same applies to latency, scalability and resilience. (Viewpoints and Perspectives)
That gives us a useful extension from two dimensions to three:
1. ABSTRACTION
How concrete are we?
Contextual → Conceptual → Logical → Physical
2. VIEWPOINT
What aspect are we describing?
Functional / Application
Data / Information
Runtime
Integration
Deployment
Development
...
3. QUALITY PERSPECTIVE
What properties must hold across those views?
Latency
Scalability
Reliability
Security
Observability
Evolvability
Cost
...
For an ML/search system, I think this is substantially better than trying to invent an “observability architecture”, “performance architecture”, “security architecture”, etc. as peer architecture levels.
2. C4. C4 is particularly strong on the communication problem. Its documentation explicitly calls out common failures such as mixed abstraction levels, unexplained notation, ambiguous elements and unlabelled relationships. (C4 model)
Its review checklist is surprisingly valuable. Among other things, it asks whether:
For engineering teams, that is arguably more important than picking the theoretically “correct” architecture taxonomy.
Kruchten’s famous 4+1 model uses:
Logical
Process
Development
Physical
where Physical essentially means deployment topology. (Concordia ENCS Users)
But the conceptual/logical/physical hierarchy we’ve been discussing uses:
Conceptual
Logical
Physical
where Physical means concrete implementation more generally.
Those aren’t contradictory. They are using the word physical along different dimensions.
This is exactly why I would avoid saying:
“This is the physical architecture.”
Instead say:
“This is the physical deployment view.”
or:
“This is the logical runtime view.”
or:
“This is the conceptual data view.”
That little two-word convention removes a surprising amount of ambiguity.
Traditional software architecture literature does not adequately emphasise data, models, training systems, evaluation and feedback loops. A few ML references fill that gap.
Still one of the most important ML-system architecture papers. It discusses system-level problems such as dependency entanglement, feedback loops, undeclared consumers, data dependencies and boundary erosion. (Google Research)
Its big architectural lesson is:
ML system
data ──────────────┐
features ──────────┤
configuration ─────┤
models ────────────┤
serving ───────────┤
monitoring ────────┤
evaluation ────────┤
dependencies ──────┘
model code
↑
only one part
So an ML architecture that merely shows:
API → Model → Response
is usually hiding most of the architecturally important system.
MLOps: Continuous delivery and automation pipelines
Particularly useful because it explicitly distinguishes:
It also stresses that production ML systems are much larger than the actual ML model. (Google Cloud Documentation)
This suggests that an ML system often deserves different runtime views:
Online search runtime
Query → Retrieval → Ranking → Results
Offline indexing runtime
Listings → Processing → Embeddings → Search indexes
Training runtime
Training data → Train → Evaluate → Model registry
Evaluation runtime
GT dataset → Candidate system → Metrics → Quality gate
Trying to combine those four onto one “architecture diagram” is usually a mistake.
Less about diagramming, but highly relevant architecturally. It repeatedly emphasises getting the end-to-end pipeline and infrastructure correct, keeping the initial ML simple, instrumenting metrics early, testing infrastructure separately from ML, and assigning ownership to data/features. (Google for Developers)
That reinforces an important point:
The architecture should expose the boundaries that teams need to reason about and test, rather than merely showing where the model sits.
After comparing these approaches, I’d refine our earlier model to this:
ARCHITECTURE DESCRIPTION
│
┌────────────────┼────────────────┐
│ │ │
ABSTRACTION VIEWPOINT PERSPECTIVE
│ │ │
how concrete? what aspect? what quality?
│ │ │
Contextual Context Performance
Conceptual Functional Reliability
Logical Data Security
Physical Runtime Scalability
Integration Observability
Development Evolvability
Deployment Cost
And one more dimension that diagrams alone don’t capture:
DECISIONS
Why was it designed this way?
What alternatives were considered?
What trade-offs were accepted?
That’s where ADRs fit.
For example:
Query Understanding
↓
Constraint Handling
↓
Retrieval
↙ ↓ ↘
lexical text image
↘ ↓ ↙
Fusion
↓
Ranking
It says what responsibilities exist.
QU API
↓
OpenSearch + Qdrant
↓
RRF implementation
↓
GPU cross-encoder
It says what implements them.
query
↓
understand
↓
parallel retrieval
↓
fusion
↓
rerank
↓
results
It says what happens during one request.
GKE
├─ query pods
├─ retrieval orchestrator
└─ ranking pods
OpenSearch cluster
Qdrant cluster
GPU endpoint
It says where concrete instances execute.
Listing
├─ metadata
├─ lexical representation
├─ text embeddings
└─ image embeddings
It explains a completely different concern without contaminating the runtime diagram.
Now apply this across those views:
Query understanding < 20 ms
Retrieval < 50 ms
Fusion < 10 ms
Reranking < 70 ms
──────────────────────────────────
End-to-end p95 < 200 ms
That might cause changes to:
Performance therefore isn’t merely another box or diagram. It is a constraint that cuts through multiple views.
The literature above points towards a fairly small set of conventions that could prevent most architecture-document ambiguity:
Name every diagram <abstraction> <viewpoint> view. For example, Logical Search Runtime View or Physical Search Deployment View.
State the question the diagram answers. Example: “How are candidates generated and ranked for an online query?”
Don’t casually mix abstraction levels.
Vector Retrieval and Qdrant 1.15 on GKE should normally not appear as peer boxes.
Separate static structure from runtime behaviour. A component dependency graph is not a request sequence.
Map logical → physical explicitly.
Logical responsibility Physical implementation
----------------------------------------------------
Lexical retrieval → OpenSearch
Vector retrieval → Qdrant
Reranking → Qwen reranker on GPU
This makes technology substitution much easier to reason about.
Treat data/model artefacts as first-class architectural elements in ML systems. Models, embeddings, GT datasets, indexes and features shouldn’t disappear behind generic “ML” boxes.
Use perspectives for cross-cutting qualities. Performance, reliability, observability, privacy, cost and security should be tested against relevant views.
Capture important decisions separately in ADRs. The diagram shows what is. An ADR explains why it became that way. (Cognitect.com)
Use a glossary. arc42 explicitly includes one for architecture-specific terminology; this becomes particularly important for words such as component, service, pipeline, model, index, retriever and ranker. (arc42)
Create only views that answer a real question. Both C4 and the older 4+1 work explicitly allow omitting views that add no value. (C4 model)
The overall goal isn’t to produce a comprehensive architecture taxonomy. It’s to make it difficult for two engineers to look at the same diagram and form different mental models of what the boxes and arrows mean. The combination I’d favour for a modern ML/search team is essentially C4’s disciplined abstraction + Rozanski/Woods’ viewpoints and perspectives + arc42’s lightweight documentation structure + ADRs + ML-specific data/training/evaluation/runtime views.