top of page

AI in Defence and Security: Demystifying AI

Dr Stephen Anning
Oct 20, 2025
7 min read

Artificial Intelligence (AI) holds extraordinary potential for enhancing operational capabilities across Defence and Security. Yet, for many, AI remains entangled in hype, misunderstood definitions, and questions over what it can realistically achieve. In response, this explainer aims to clarify three key paradigms of AI – Strong AI, Narrow AI, and Frontier AI – and reflect on the governance needed to ensure their responsible, effective use in Defence and Security contexts. The narrative here is designed to equip personnel to navigate these paradigms, challenge AI hype, and appreciate how thoughtful governance can enable success.

The Market Context: The Gartner Hype Cycle

Figure 1: The Gartner Hype Cycle

Mainstream discussions of AI follow a familiar arc known as the “Gartner Hype Cycle,” illustrating how an initial innovation trigger yields exaggerated promises (the peak of inflated expectations) before colliding with reality in the trough of disillusionment. From there, a more grounded slope of enlightenment emerges, culminating in the plateau of productivity. The latest version of the hype cycle places AI in the “trough of disillusionment”.

This analysis aims to provide a narrative for taking Defence from the “trough of disillusionment” to the “slope of enlightenment,” equipping practitioners and decision-makers with realistic expectations. Once illusions are shed, the real benefits of AI can be harnessed responsibly.

A strong feature of this hype is the conflation of “AI” with machine learning or the reference to “simulating human intelligence.” In reality, “AI” is an umbrella term that means different things in different contexts. AI can indeed be advanced mathematics and coding, but it can also involve textual reasoning, experiential factors, or pattern recognition. The danger is that if we view AI purely as a set of formulae or lines of code, we miss subtler human aspects – especially for tasks requiring judgement or moral nuance, such as interpreting intelligence reports about people and populations. Another hazard arises from the mismatch between purely numeric data approaches and the complexity of real operations involving people, culture, and unpredictable human behaviour.

This analysis explores three paradigms of AI: Strong AI, Narrow AI, and Frontier AI. Confusion arises when we use these paradigms interchangeably, so clarifying them is essential for operational success.

Strong AI (Artificial General Intelligence)

Also known as Artificial General Intelligence, “Strong AI” imagines a machine that exhibits human-equivalent intelligence, capable of reasoning, learning, and understanding across any domain. This vision often resides in philosophy and science fiction because it asks: can machines truly think or become conscious? One iconic basis for Strong AI is Alan Turing’s “Turing Test,” a thought experiment he introduced in 1950 to determine if a machine could convincingly mimic human conversation to the extent that a tester could not distinguish it from a human interlocutor.

Philosophical thought experiments like “Mary’s Room” probe the notion of human intelligence. In this thought experiment, Mary is a scientist who knows every theoretical fact about colour but has never experienced it outside her black-and-white room. When she finally sees colour, does she learn anything new?Mary’s Room thought experiment highlights a tension between purely formal or data-driven knowledge and the lived, subjective experience humans hold. Advocates of Strong AI suppose that one day, advanced computation will capture these experiential and emotional facets of intelligence. Conversely, sceptics question whether the more intangible aspects of human consciousness could ever be captured in the inner workings of a machine.

For now, Strong AI remains speculative. Philosophical questions of machine consciousness seldom directly shape near-term procurement or operational systems. Nonetheless, they remind us that human emotion, subjective experience, and moral reasoning often underpin operational decisions about people, aspects not easily digitised. This caution is relevant if we over-delegate critical missions to machines. Even sophisticated systems risk excluding the intangible dimension of soldierly intuition, empathic leadership, or evolving morale on the ground.

Narrow AI (Machine Learning)

In contrast to Strong AI’s lofty ambitions, Narrow AI (sometimes called “Weak AI”) is the workhorse of most AI systems in operational use. These systems excel at specific tasks, from image classification to speech recognition, but they lack broader reasoning or general awareness. The mathematics behind narrow AI often involves machine learning algorithms trained on large data sets. For instance, an image-recognition model can learn to identify vehicles or terrain features relevant to an intelligence briefing but cannot autonomously direct strategic objectives.

The key to Narrow AI is data. Models learn patterns from historical examples through supervised, unsupervised, semi-supervised or reinforcement learning.

  • Supervised learning trains a model on labelled data, learning to map inputs to known outputs.

  • Unsupervised learning finds hidden patterns or structures in unlabelled data without predefined outcomes.

  • Semi-supervised learning combines a small amount of labelled data with a large amount of unlabelled data to improve learning accuracy.

  • Reinforcement learning teaches an agent to make decisions by rewarding desirable actions and penalising undesirable ones through trial and error.

Learning from data using machine learning algorithms reveals limitations: an algorithm inherits the flaws of biased or incomplete training data. When Narrow AI addresses human subjects, such as classifying the underlying beliefs and motivations of people, ethical pitfalls arise quickly. A mathematically robust method might produce harmful results in real-world contexts if intangible factors such as morale or culture are misquantified.

For Defence and Security, Narrow AI can be transformative. It underpins many detection and alert systems, from spotting anomalies in satellite imagery to triaging vast intelligence queues. The challenge is that these systems, unless carefully governed, can produce “weapons of math destruction” – a term coined by the data scientist Cathy O’Neil to describe AI that inadvertently causes injustice or intensifies inequality. Procurement officers, decision-makers and analysts must ensure that each new AI deployment includes robust training data, performance monitoring, and bias testing.

Frontier AI (Generative or Foundational Models)

The third paradigm, Frontier AI, captures the recent surge in “generative AI,” including large language models (LLMs) such as OpenAI’s ChatGPT, Google’s Gemini, or Microsoft’s Copilot. Generative AI harnesses massive computational power and scaled training data to create foundational or general-purpose models that tackle tasks from summarising text to generating images. Many believe these highly capable models blur the lines between Narrow AI and genuine human-like intelligence, although others emphasise that a “prediction engine” for words or images does not equate to consciousness or authenticity.

Frontier AI’s significance stems from its potential to handle an array of tasks swiftly, albeit with occasional “hallucinations”. Hallucinations in frontier AI systems are outputs that are syntactically correct but factually untrue. Models rely on next-token prediction, selecting the likeliest subsequent words based on their training data. If that data is biased, out-of-date, or incomplete, the model’s output can be misleading or harmful. Despite such pitfalls, Frontier AI can be harnessed in Defence and Security for fast analysis of open-source intelligence or for real-time speech translation, provided it is embedded in an appropriate governance structure.

Where authoritative information is crucial, “retrieval-augmented generation” pairs a large language model with factual databases, thus grounding responses in verifiable data. Even so, hazards remain if we uncritically trust the machine’s output. The scale alone demands robust oversight. LLMs can ingest entire corpora of text from the public domain, inadvertently codifying both the valuable knowledge and the misinformation present in that corpus.

Ethical & Practical Governance

The operational examples above underline the critical need for oversight. People in the field may be swayed by a software system that appears to “understand” them. As illustrated by Joseph Weizenbaum’s Eliza therapy bot in the 1960s, we often over-ascribe intelligence to code that parrots patterns. Known as the “Eliza Effect,” this vulnerability amplifies the risk that soldiers, analysts, or commanders might anthropomorphise AI and cede too much authority to an algorithm.

Just as firearms training ensures safe usage, AI literacy ensures that operators appreciate how AI learns, how it might err, and what checks are needed. Fundamental guidelines include:

  • Thoroughly vet training data. If a model is used for human assessment, be certain its historical data does not systematically exclude or misrepresent minority groups.

  • Continuously update the data and model. Shifts in context – e.g. “terrorist” to “peace negotiator” – can flip labels if the world changes but the system’s knowledge remains static.

  • Provide feedback loops. If no mechanism exists for capturing mistakes or questionable outputs, then mislabelling and biases persist. This is especially dangerous in generative AI, which can “hallucinate” plausible but false statements.

  • Pair the model with deterministic or rules-based checks for high-stakes decisions. If you can define a rule around dietary restrictions or legally binding constraints, consider coding it explicitly.

  • Clarify lines of authority. In high-risk tasks, no soldier should say “I did it because the machine said so.” Accountability remains with the chain of command.

Conclusion

Defence and Security communities stand on the threshold of widespread AI adoption across these paradigms: Strong AI, though largely philosophical, pushes us to consider the intangible or emotional dimensions of humanity in modern operations in Defence and Security. Narrow AI, built upon machine learning, brings immediate operational value to narrowly defined tasks. If integrated responsibly, Frontier AI’s generative or foundational models can multiply productivity and insight. In the following years, Frontier AI is expected to continue evolving rapidly, intensifying debates about bridging the gap between advanced pattern prediction and actual general intelligence.

Yet no matter how advanced AI becomes, best practices involve a synergy of human oversight, robust datasets, iterative model refinement, and unwavering attention to bias or hallucination. This synergy forms the basis of “responsible AI”: the freedom to exploit cutting-edge capabilities tempered by a framework of accountability that fosters trust. Like firearms, AI demands both skill and constraint. That dual imperative underpins its potential to transform how we plan, fight, and ultimately protect our way of life.

References

  • Gartner (2025) Gartner Hype Cycle Identifies Top AI Innovations in 2025

  • O’Neil, C. (2017). Weapons of math destruction: How big data increases inequality and threatens democracy

  • Jackson, F. (1986). What Mary didn't know. The journal of philosophy, 83(5), 291-295.

  • Weizenbaum, J. (1966). ELIZA—a computer program for the study of natural language communication between man and machine. Communications of the ACM, 9(1), 36-45.

 
 
 

Recent Posts

See All
How to Write an Annotation Schema

Writing an annotation schema for qualitative research involves translating your research aims and theoretical concepts into a clear, structured system for labelling and interpreting data. The process

 
 
 
What is an Annotation Schema?

An annotation schema for qualitative research is a structured framework that sets out how researchers label, categorise, and interpret data such as text, audio, video, or images. Its purpose is to en

 
 
 

Comments


bottom of page