EVALUATING AI | Exercising Reason, Rigor, and Restraint
“AI” is now the most-discussed word in architecture, AND the least precise.
In a single meeting it might mean a foundation model, a rendering plugin, a subscription service, or a category of generative-design tools that predate the current AI boom by a decade. That imprecision is why so many firms struggle to evaluate AI-enabled technology with real discipline.
Before a firm can judge whether a tool is worth adopting, it has to know what kind of tool it’s looking at. This Brief lays out a framework for that approach, using generative planning, AI visualization, and AI-enabled production and office workflows as working case studies—not to crown winners, but to show how to think about them.
The argument: classify a technology correctly before evaluating it, and adopt
AI is Not a Product Category

Sorting a tool into its correct layer first separates an apples-to-engines comparison from a real one.
Capability is Not the Same as Application
Even within the application layer, AI describes capabilities that have almost nothing in common. It’s worth separating four:

It’s also worth resisting the reflex to treat ‘generative’ as synonymous with ‘AI.’ Parametric and rule-based systems—Grasshopper definitions, Dynamo scripts, CityEngine’s CGA rules—can generate sophisticated design alternatives without any machine learning at all, producing output algorithmically from rules and inputs.
That’s a meaningfully different technology from a model trained on millions of images to predict plausible pixels, with different failure modes and different dependence on training data. Knowing which one a firm is actually adopting changes what it should reasonably expect from it.
The Right Tool for the Right Problem for the Right Phase
ArcGIS CityEngine, Autodesk Forma, and TestFit are routinely grouped together as ‘generative planning tools’—all three produce building massing from rules and site data. But grouping them obscures more than it reveals. Each solves a different problem, at a different phase, at a different scale.

AI Beyond the Desktop: Workflow and Project Productivity
Most of the AI conversation in A/E is focused on design and visualization — but much of the realistic near-term value sits somewhere less glamorous: specification writing, construction-document review, RFI drafting, meeting notes, and general office production.
These are language and document tasks, not geometry tasks, and call for the same discipline — against a different set of capabilities and a very different stakes profile.

General-purpose office assistants—Copilot and Gemini in Microsoft 365 and Google Workspace, transcription and scheduling tools—sit at the low-stakes, high-volume end of this spectrum, worth adopting broadly precisely because a wrong meeting summary is cheap to catch and fix. That asymmetry is itself a diligence signal: it should set how much human-in-the-loop review a use case actually warrants, independent of how impressive the underlying model is.
Evaluate the Ecosystem, Not Just the Tool
Once a tool is correctly classified and matched to a real problem and phase, the harder work starts: evaluating it as a fit for the firm, not just as a product on its own terms. Six factors belong in that evaluation.

Platform durability deserves scrutiny, since “AI-enabled” spans everything from Alphabet-backed products, with a fairly low discontinuation risk, to single-founder startups launched within the past year. Neither extreme is disqualifying—but risk profile should be a named factor, not an afterthought discovered when a favorite trial goes unmaintained—see Google Delve; RIP Sidewalk Labs.
Two friction points recur often enough to flag on their own: 1) Credit- or token-based pricing—common across AI-native tools—trades the predictability of a per-seat license for a metering system where the real cost of a render or a generation isn’t obvious until the invoice arrives; it has to be budgeted from actual usage, not the price page. 2) Trial periods and learning resources are just as often mismatched to how firms evaluate software: MV+A’s 2024 review found thirty-day trials too short to reach a real verdict, and found comparable gaps in vendor tutorials across every application reviewed. Neither issue disqualifies a tool on its own, but both belong on the checklist—not discovered mid-pilot.
Adopt [AI] Workflows
The goal should never be to find a place to ‘add AI.’ It is to find where an existing practice is genuinely limited, then ask whether an emerging technology closes that gap better than the workaround in place—a question only testing on a real project can answer.
Running that sequence honestly requires holding three kinds of maturity apart, since a tool can score high on one and low on another:

The Take-Aways
None of this argues for caution over adoption—it argues for sequencing, and it applies as much to a spec-review agent as to a rendering plugin. Firms that get real value from AI won’t be the ones that adopted first or evaluated the most tools; they’ll be the ones that classified correctly, matched capability to problem and phase, and tested before committing.
-
AI is a stack, not a single product—know which layer you’re evaluating.
-
Generative doesn’t mean AI, and AI visualization doesn’t mean AI design.
-
The right question is never “which AI tool is best”—it’s “what problem, which phase, and which tool?”
-
Oversight should scale with the stakes of the output, not the sophistication of the tool.
-
Adopt workflows that happen to use AI—not AI in search of a workflow.

