Blog

AI AUDEX: Advancing Autonomous Audio Processing for the DoD

blogpic-080326

By Timothy & Jeremy, Research Scientists
Four-minute read

Across the Department of Defense, artificial intelligence continues to reshape how organizations process information, accelerate workflows and improve mission effectiveness.

Large Language Models (LLMs) and multi‑agent AI systems are now appearing in everyday operational environments—from AFRL’s NIPRGPT pilot to the Army’s deployment of Ask Sage and CamoGPT. These technologies are helping analysts automate repetitive tasks, streamline logistics and make sense of increasingly complex data streams.

At AIS, our Independent Research & Development (IR&D) effort, AI Agents for Audio Exploitation (AI AUDEX), explored how multi‑agent AI teams could autonomously enhance speech audio with minimal human intervention. This work builds directly on our SIGINT experience and complements upcoming capabilities for NASIC’s Haystack modernization, where audio‑aware reasoning and summarization tools will play a major role.

This project offered a unique opportunity: combine advanced agent orchestration with practical mission challenges in audio processing, and determine how far autonomy can be pushed while preserving the quality and intelligibility of speech.

The Challenge: Automating a Tedious Workflow

AIS’s Audio Group currently uses the Rapid Audio Batch Tool (RABT), a distributed processing capability that allows analysts to perform a wide variety of audio enhancements. RABT is powerful—but its workflow construction is entirely manual. Analysts must drag and connect processing components one by one, often repeating the same patterns across large batches of files.

AI AUDEX asked a simple question with large implications:
Could autonomous AI Agents learn how to build and execute audio enhancement workflows on their own—without requiring analysts to design them manually?

To answer it, our team developed and tested two novel approaches to multi‑agent speech enhancement.

Two Approaches to Multi‑Agent Reasoning

AI AUDEX evaluated two distinct models for autonomous speech enhancement:

1.

Multi‑Iteration LangGraph Agents
These agents reasoned visually over spectrograms and iteratively selected enhancement tools. Their design emphasized multi‑step reasoning and cross‑agent collaboration.

2.

Single‑Iteration Deep Agents
This approach removed spectrograms entirely. Agents made decisions using only audio statistics, operating in a single-pass selection cycle to improve speed.

Both techniques were tested against 100 noisy speech files at five different noise levels.

How We Measured Performance

Speech enhancement is complex, and no single metric can fully capture how humans perceive improvements. To evaluate the agent teams, AIS used three industry‑standard measures:

STOI and PESQ are complementary. STOI focuses on clarity of words; PESQ focuses on perceived audio quality. Modern audio research always uses both because enhancements often create trade-offs. Improving intelligibility can reduce naturalness, and removing noise can introduce distortions.

Key Findings

AI AUDEX produced several important insights that will shape future R&D and customer-facing solutions.

1.

Both agent approaches improved SNR significantly.
Noise reduction was strong across all test files and noise levels.

2.

STOI and PESQ degraded under most conditions.
This reflects a common challenge in speech processing: aggressive cleanup often harms intelligibility or introduces artifacts. Both agent teams tended to “over-process” because they were biased toward selecting the most tools available to them.

3.

Restricting agents improved results.
When guardrails limited how many enhancement steps agents could apply, performance improved across all three metrics.

4.

Deep Agents outperformed LangGraph Agents.
They achieved:

  • Faster processing times (about 190 seconds per file)
  • Slightly better STOI and PESQ
  • Comparable or better SNR
  • And they accomplished this without the need for spectrograms.

Why This Work Matters for AIS and DoD Customers

AI AUDEX expands AIS’s capabilities in multi‑agent autonomy, audio analysis, and multi‑modal reasoning. The tooling, insights, and performance benchmarks developed through this IR&D now serve as:

  • Past performance for autonomous audio processing
  • Reusable patterns for multi‑agent orchestration
  • Building blocks for future work on Haystack and other SIGINT modernization efforts
  • Evidence that AIS can safely automate complex workflows previously requiring significant analyst time

The effort also highlighted important opportunities for future research, including integrating agents with state-of-the-art multimodal audio models capable of directly consuming waveform input. These emerging models may help resolve the STOI and PESQ trade-offs seen in this study, enabling agents to make more human‑like decisions about audio quality and intelligibility.

Looking Ahead

As the DoD continues to adopt AI tools to streamline operations and accelerate mission decision-making, autonomous multi‑agent systems will play a larger role in data-heavy environments like SIGINT. The AI AUDEX project positions AIS at the forefront of that evolution.

Our next step is applying these findings to real mission workflows — reducing analyst burden while improving speed, consistency and autonomy in complex audio processing tasks.

Privacy Settings
We use cookies to enhance your experience while using our website. If you are using our Services via a browser you can restrict, block or remove cookies through your web browser settings. We also use content and scripts from third parties that may use tracking technologies. You can selectively provide your consent below to allow such third party embeds. For complete information about the cookies we use, data we collect and how we process them, please check our Privacy Policy
Youtube
Consent to display content from - Youtube
Vimeo
Consent to display content from - Vimeo
Google Maps
Consent to display content from - Google
Spotify
Consent to display content from - Spotify
Sound Cloud
Consent to display content from - Sound
Assured Information Security
Privacy Overview

This website uses cookies so that we can provide you with the best user experience possible. Cookie information is stored in your browser and performs functions such as recognising you when you return to our website and helping our team to understand which sections of the website you find most interesting and useful.