Limitations of generative AI in design work

A

Anthony Massobrio

CFD Expert & AI for CAE Contributor

·

July 29, 2026

·

When engineering teams evaluate AI tools, they typically walk into a conversation shaped by general-purpose chatbots. Meanwhile, vendor pitches for physics-aware deep learning circulate in a separate register entirely. The two may be treated as the same technology, but they are not, and that confusion is the root of the most common misconceptions about the limitations of generative AI in design work.

In short, the limitations of generative AI in design work stem from one fact: general-purpose models optimize for plausibility rather than physical correctness, so they hallucinate, lose information in long documents, expose sensitive data, and face regulatory bans on high-risk use.

For the engineers and designers who evaluate these new technologies, that distinction matters because a tool that looks convincing is not the same as one that is verifiably correct, and in the near future, the cost of that gap grows as more of the design process is handed to AI.

This article covers general-purpose generative AI (GenAI), engineering AI, and their respective limitations:

  • The first half of the article answers two questions: what general-purpose AI systems can do as of 2026 and where they fail. Sources include published benchmarks and the EU regulatory framework.
  • The second half explores what engineering teams need from AI and what Engineering Intelligence delivers that an LLM cannot. Engineering Intelligence is a class of deep learning models trained on CAD and CAE data and constrained by physical laws.

The FAQ at the end covers recurring questions about generative AI in design.

"Classical" input-to-output inference pipeline for general-purpose and engineering generative AI. They share a node structure but diverge at every node | Author

Table of contents

  1. What are the top Artificial Intelligence capabilities in 2026?
  2. What are the risks and limitations of general-purpose AI in the near future?
  3. What are engineering’s needs?
  4. What is Engineering Intelligence?
  5. How to manage uncertainty?

Note:

1. What are the top Artificial Intelligence capabilities in 2026?

The top AI capabilities in 2026 are agentic systems that plan, reason, and act autonomously across complex workflows. This emerging technology is also reshaping design process expectations through rapid prototyping and experimentation, helping teams create and refine concepts faster than with traditional methods. With the generative AI market projected to reach USD 136.7 billion by 2030, organizations are paying close attention to how quickly these workflows are shifting (source: MarketsandMarkets). The innovation is frontier Large Language Models extended with multimodality, browser and workspace automation, multi-agent coordination, and advanced reasoning:

  • Native Multimodality: Frontier models are now trained as natively multimodal, sharing text, audio, images, and video within a single context window.
  • Browser and Workspace Automation: Agents now control web browsers and desktops end to end, completing multi-step tasks such as booking a flight by opening a browser, reading a calendar, navigating the interface, and paying through secure frameworks⁽²⁾.
  • Multi-Agent Ecosystems: Some systems coordinate multiple LLMs for tasks, like Anthropic’s research system, which uses a lead orchestrator and sub-agents to explore queries, with a citation agent verifying sources⁽³⁾.
  • Advanced Reasoning and Logic: Models trained with reinforcement learning generate intermediate steps before final answers⁽⁴⁾.

2. What are the risks and limitations of general-purpose AI in the near future?

The main risks and limitations of general-purpose AI are confident but false outputs (hallucinations), accuracy that decays inside long documents (context rot), leakage of sensitive corporate data, and regulatory bans on high-risk uses.

The trust and accuracy barrier

  • The “Confident Mistake” (Hallucinations): An LLM predicts the next tokens based on probabilities, producing fluent but potentially false prose. It can generate fabrications with the same fluency as facts, making falsehoods hard to spot without checking.
  • Context Rot or Lost in the Middle: Despite large context windows, model accuracy drops for info in the middle. Recall is reliable at the start and end but falls sharply in between.

The everyday use of generative AI can reduce human judgment and push teams to create less, acting more like prompt engineers or curators.

LLM hallucination: objective, storage, decoding & calibration jointly produce fluent output that is statistically plausible but unanchored to truth | Author

Data security and compliance risk

To make effective general-purpose AI, users must provide specific real-world data. However, about 8.5% of employee prompts to public AI accidentally leak sensitive corporate data, raising privacy issues, compliance exposure, and conflicts with data protection laws like GDPR. Using generative AI in design and digital products introduces risks of bias, security, and deployment when prompts, assets, or user data are sent to public systems.⁽⁹⁾

This illustration of artificial intelligence has in fact been generated by AI! | [www.europarl.europa.eu](https://www.europarl.europa.eu)

3. What are engineering’s needs?

Engineering needs three things that general-purpose AI does not provide, because it operates in a physical reality dictated by a triad of

(1) quantitative objectives / KPIs,

(2) physical constraints, and

(3) immediate, quantifiable verification.

Why is Engineering Intelligence the solution for the modern engineer rather than general-purpose artificial intelligence? We must look at the foundational architecture of Mainstream generative AI’s interactions with reality.

Mainstream generative AI operates within an Internet of Words. When a user prompts a Large Language Model (LLM) to write an essay, design a marketing campaign, or generate an image, it can produce new content or original content, but that does not mean it has the engineering knowledge needed for real-world design work. The primary metric of success is plausibility, not absolute mathematical truth⁽¹¹⁾.

Engineering does not live in the Internet of Words

Engineering lives in a physical reality dictated by a strict triad:

  1. quantitative objectives,
  2. physical / engineering constraints
  3. quantifiable verification.

Engineering projects require structured knowledge, constraints, and validation criteria before AI can contribute useful solutions. The primary metric of success is plausibility only after teams have AI that can expand feasible design options within those constraints, not merely produce plausible output.

There are six categories of engineering constraints: physical, technical, economic, time, regulatory, and environmental. Each shape determines which solutions are even admissible before optimization begins.

The specifics of an engineering prompt

An enterprise engineer’s prompt is a multi-dimensional boundary condition: target objectives, manufacturing limits, and performance thresholds expressed as explicit numerical parameters, at the opposite end of the spectrum from everyday AI use, where free-form natural language is the entire input.

DimensionEveryday AI TechnologyEngineering Intelligence
Objective ScopeBroad, creative, subjectiveNiche, deterministic, specific
Error ToleranceHighZero
Boundary ConditionsFlexibleHard-constrained

The hallucinated design

Engineering AI must be physically accurate, unlike consumer AI, which tolerates hallucinations. Systems must do more than generate designs that merely look convincing: they have to verify that the result is physically valid. Neural Concept embeds physics-aware deep learning in CAD for rapid verification, flagging uncertainty, and invoking CAE only when necessary.

Engineering Intelligence uses physics-aware models to evaluate geometry changes instantly, preventing infeasible forms. Trained on real data and physics-constrained, it acts as a real-time sanity check, flagging violations early and rejecting issues.

4. What is Engineering Intelligence?

Engineering Intelligence is a specialized AI field that incorporates physics-awareness, spatial reasoning, and 3D geometry into product development. It serves as a layer on top of existing engineering software, such as CAD and CAE, to revolutionize product design and validation.

 Engineering Intelligence as a layer on top of existing engineering software

What are the key components of Engineering Intelligence?

  • Physics-Awareness: Unlike LLMs, Engineering Intelligence grasps physics laws. Deep Learning pattern recognition predicts engineering outcomes starting from CAD geometry description. This takes place in milliseconds, rather than hours, as with traditional simulations.
  • Geometric Intelligence (3D Deep Learning) gives AI spatial reasoning: the capacity to interpret complex 3D CAD geometries directly, recognize patterns, and propose modifications that are manufacturable without leaving the CAD environment.
  • Human-AI Collaboration (Copilots) describes workflows in which AI Design Copilots act as virtual assistants, enabling engineers to explore candidate solutions and navigate trade-offs among performance objectives (weight vs. durability) in real time. AI can also help create more personalized and contextualized experiences by analyzing user data and preferences, but it still requires human judgment. They still depend on domain knowledge and human judgment: without expert knowledge encoded into the process, they do not unlock the full potential of engineering workflows.

Practical applications in design

Several production deployments span automotive, aerospace, and Formula 1, for instance:

  • Automotive: Optimizing external aerodynamics and crash safety (used by Subaru and GM).
  • Aerospace: Predicting pressure fields on airplane bodies in 30 ms (used by Airbus, see figure).
  • Formula 1: Making real-time design adjustments where milliseconds of performance determine race outcomes (used by Visa Cash App Racing Bulls).
AI reaching very high level of accuracy: comparison of CFD and AI prediction

What are the key metrics for AI prediction?

Performance metrics such as , MAE, and (R)MSE are used to determine whether a predictive model is production-ready.

  • represents the percentage of the physical variance in the simulation data that the AI model can explain. Technical case studies (such as those with Airbus or LS Electric) emphasize achieving high R² values.
  • Mean Absolute Error (MAE). While R² gives a relative score, MAE provides the error in the actual physical units.
  • (Root) Mean Squared Error (R)MSE. We use MSE and RMSE when the cost of a “large mistake” is high, for example, in structural integrity, where a single underprediction could lead to a part failure.
  • MSE and RMSE are based on the L2 norm, which penalizes larger errors more heavily. A worked sensitivity analysis on satellite thermal design, with 27 geometric parameters and explicit L1/L2 reporting, is set out in a white paper on Smarter Satellite Design for Thermal Constraints.
relative error on predictions of satellite thermal fields. Training phase mean 0.64% (blue), testing phase mean 1.17% (orange).

These metrics translate to public, verifiable results.

  • On MIT’s DrivAerNet++ drag-coefficient task, Neural Concept’s Geometric Regressor leads the previously published best (TripNet) on all three measures defined above: R², MAE, and MSE. The same model also leads the public leaderboard on surface pressure, wall shear stress, and volumetric velocity; see the full benchmark report.
Comparison of the ground truth pressure field (left), the AI model prediction (middle), and the corresponding error for a representative test sample (right).

5. How to manage uncertainty?

Reliable Engineering Intelligence manages uncertainty through four complementary methods: iterative-convergence estimates, Bayesian optimization with confidence intervals, geometry- and physics-aware constraints, and a human-in-the-loop workflow.

1. The “Iterative Convergence” Approach

A unique method developed involves convergence rates as a proxy for uncertainty: “iterative neural networks“ refine a prediction over internal steps at a lower computational cost than traditional methods.

2. Bayesian Optimization & Confidence Intervals

Predictive models can be embedded in a Bayesian Optimization framework that provides a confidence interval. The user can see the range of potential outcomes and decide whether to rerun CAE to obtain more data. See the paper on the importance of uncertainty quantification.

3. Geometric & Physics Awareness

Traditional Artificial Intelligence fails when it makes “impossible” predictions (the "Black Box" problem). Engineering AI manages this uncertainty by embedding physical laws into the loss function. The physics-aware loss keeps predictions within physical possibility⁽¹²⁾. See also best practices for improved simulation.

4. Human-Artificial Collaborative Workflow

The Design Copilot based on physics-aware AI empowers engineers with quick exploration. Traditional CAE for final validation balances speed with physics-based results.

FAQ

What are the limitations of generative AI interfaces in design?

Generative AI can struggle to produce highly complex or high-resolution designs due to computational intensity that may exceed current technology. The model may not fully understand human aesthetics and cultural significance, which are crucial in the design creation process, indicating a need for more advanced algorithms and training data.

What is a hallucination in LLMs?

Generative AI can produce inaccurate or false information, a phenomenon known as hallucination, which raises concerns about their reliability in providing truthful answers. The phenomenon known as hallucination presents inaccuracies in a confident tone, making it difficult to discern truth from fiction.

What are the main ethical challenges posed by generative AI models?

The lack of accountability in generative AI outputs poses ethical challenges, as users must verify the accuracy of AI-generated content, potentially leading to misinformation and undermining trust.

For instance, the algorithms used in generative AI often reproduce biases present in their training datasets, leading to outputs that may reinforce harmful stereotypes and misrepresent minority voices.

Generative AI models lack source traceability, raising concerns about academic integrity and the ability to reference and credit the original authors of the content used in training.

Generative AI tools can inadvertently reinforce harmful stereotypes and biases present in their training, perpetuating discrimination in their outputs.

What are the potential copyright issues of GenAI?

There is no clear legal framework for the ownership of GenAI work, leading to intellectual property ambiguity.

Generative AI raises significant ethical concerns regarding copyright and intellectual property, as it could relies on information sourced from the web without explicit consent from the original creators, which can lead to the exploitation of their work.

Generative AI interfaces frequently lack contextual understanding, which can lead to confusing or unrelated answers, as they struggle to maintain the broader context of conversations.

What is the environmental impact of generative AI?

The environmental impact of generative AI is a growing concern, with estimates suggesting that AI technologies could consume as much electricity as entire countries do, raising questions about sustainability amid global energy needs.

Sources

(1) https://www.digitalengineering247.com/article/neural-concept-debuts-physics-aware-ai-design-copilot

(2) https://www.mastercard.com/global/en/news-and-trends/stories/2025/agentic-commerce-framework.html

(3) https://www.anthropic.com/engineering/built-multi-agent-research-system

(4) https://arxiv.org/abs/2501.12948

(5) https://sqmagazine.co.uk/llm-hallucination-statistics/

(6) https://huggingface.co/spaces/vectara/leaderboard

(7) https://arxiv.org/abs/2401.01301

(8) https://arxiv.org/abs/2307.03172

(9) https://www.csoonline.com/article/3819170/nearly-10-of-employee-gen-ai-prompts-include-sensitive-data.html

(10) https://www.europarl.europa.eu/topics/en/article/20230601STO93804/eu-ai-act-first-regulation-on-artificial-intelligence

(11) https://arxiv.org/abs/2307.02511

(12) https://link.springer.com/article/10.1007/s10462-025-11322-7

A

Anthony Massobrio

CFD Expert & AI for CAE Contributor

Anthony has been a CFD expert since 1990, working initially as a senior researcher, then moved to Engineering, acting also as technical director in a challenging Automotive Tier 1 supplier environment. Since 2001, Anthony has worked in Software & Engineering Consultancy as a Sales Engineer and manager. In 2020, Anthony fell in love with AI and has worked since then in the field of “AI for CAE” at Neural Concept and as an independent contributor.

Connect on LinkedIn