Speaker Identification via Audio-Text Segmentation in Tactical Comms

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Tactical communications systems face challenges in speaker identification due to low-rate voice encoders that obscure voice quality and lack sender validation in digital-speech based communications, and SMS-based communications that do not provide audio queues for speaker verification.

Innovation Solution

The system stores reference audio files with speaker information on communication devices, captures and compares audio messages to determine matching files, generates secure text messages with appended speaker information, and optionally compresses and transmits them, ensuring secure and verified identity verification and message reformulation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If low rate voice encoders are used to reduce bits required to represent voice, then data transmission efficiency is improved, but speaker identification capability deteriorates due to obscured voice quality

Engineering Contradiction:
Improvebits required to represent voiceVSAvoidspeaker identification capability
Core Design Contradiction:
Quantity of substanceVSMeasurement precision

Solution Approach 1:

The system segments the voice communication into two independent components: (1) the audio signal itself for conveying message content, and (2) a separate text-based identifier that explicitly states the speaker's identity. This segmentation allows the audio to be compressed for efficiency while the speaker identification is maintained through the separate text component, resolving the contradiction between bit reduction and speaker recognition.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system introduces a text-based identifier as an intermediary element between the audio signal and the receiver. This intermediary carries the speaker identification information independently of the audio quality, allowing the audio to be heavily compressed or even replaced while maintaining reliable speaker verification through the text intermediary.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Device complexity

If SMS based communication is used to transmit voice as text, then data transmission is simplified, but sender validation capability deteriorates due to lack of audio queues for speaker verification

Engineering Contradiction:
Improvecommunication system complexityVSAvoidsender validation capability
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The system merges the advantages of text-based communication (simplicity, reliability) with audio-based communication (speaker verification) by combining both modalities in a single message structure. The text portion provides simple transmission and speaker identification, while the audio portion provides voice quality verification, together resolving the contradiction between simplicity and reliability.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system performs preliminary action by pre-associating speaker identifiers with audio queues before communication occurs. This allows the receiver to have pre-established reference audio samples for comparison, enabling reliable speaker verification without requiring complex real-time analysis, thus maintaining simplicity while improving reliability.

Inventive Principle:
Principle #10Preliminary action

3Reliability

If audio messages are transmitted to ensure speaker identification, then speaker recognition reliability is improved, but data transmission size increases

Engineering Contradiction:
Improvespeaker recognition reliabilityVSAvoiddata transmission size
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The system applies partial action by transmitting only the necessary audio information for speaker verification rather than the complete audio message. The audio is processed to extract key speaker identification features, and only these condensed features are transmitted, achieving reliable speaker recognition with minimal data overhead rather than transmitting full audio quality.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentEP3910508B1System and methods for speaker identification, message compression and/or message replay in a communications environment
Publication Date: 2024.01.31 L3HARRIS GLOBAL COMMUNICATIONS INC
  • EP3910508B1 patent drawingFigure 1
  • EP3910508B1 patent drawingFigure 2~4
  • EP3910508B1 patent drawingFigure 5~7

AI summary

Systems (100) and methods (800) for communicating information. The methods comprise: storing message sets in Communication Devices ("CDs") so as to be respectively associated with speaker information; performing operations, by a first CD, to capture an audio message spoken by an individual and to convert the audio message into a message audio file; comparing the message audio file to each reference audio file in the message sets to determine whether one of the reference audio files matches the message audio file by a certain amount; converting the audio message into a text message when a determination is made that a reference audio file does match the message audio file by a certain amount; generating a secure text message by appending the speaker information that is associated with the matching reference audio file to the text message, or by appending other information to the text message; transmitting the secure text message.