Speaker Identification via Audio-Text Segmentation in Tactical Comms
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Tactical communications systems face challenges in speaker identification due to low-rate voice encoders that obscure voice quality and lack sender validation in digital-speech based communications, and SMS-based communications that do not provide audio queues for speaker verification.
Innovation Solution
The system stores reference audio files with speaker information on communication devices, captures and compares audio messages to determine matching files, generates secure text messages with appended speaker information, and optionally compresses and transmits them, ensuring secure and verified identity verification and message reformulation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If low rate voice encoders are used to reduce bits required to represent voice, then data transmission efficiency is improved, but speaker identification capability deteriorates due to obscured voice quality
Solution Approach 1:
The system segments the voice communication into two independent components: (1) the audio signal itself for conveying message content, and (2) a separate text-based identifier that explicitly states the speaker's identity. This segmentation allows the audio to be compressed for efficiency while the speaker identification is maintained through the separate text component, resolving the contradiction between bit reduction and speaker recognition.
Solution Approach 2:
The system introduces a text-based identifier as an intermediary element between the audio signal and the receiver. This intermediary carries the speaker identification information independently of the audio quality, allowing the audio to be heavily compressed or even replaced while maintaining reliable speaker verification through the text intermediary.
2Device complexity
If SMS based communication is used to transmit voice as text, then data transmission is simplified, but sender validation capability deteriorates due to lack of audio queues for speaker verification
Solution Approach 1:
The system merges the advantages of text-based communication (simplicity, reliability) with audio-based communication (speaker verification) by combining both modalities in a single message structure. The text portion provides simple transmission and speaker identification, while the audio portion provides voice quality verification, together resolving the contradiction between simplicity and reliability.
Solution Approach 2:
The system performs preliminary action by pre-associating speaker identifiers with audio queues before communication occurs. This allows the receiver to have pre-established reference audio samples for comparison, enabling reliable speaker verification without requiring complex real-time analysis, thus maintaining simplicity while improving reliability.
3Reliability
If audio messages are transmitted to ensure speaker identification, then speaker recognition reliability is improved, but data transmission size increases
Solution Approach 1:
The system applies partial action by transmitting only the necessary audio information for speaker verification rather than the complete audio message. The audio is processed to extract key speaker identification features, and only these condensed features are transmitted, achieving reliable speaker recognition with minimal data overhead rather than transmitting full audio quality.
Data Source
Figure 1
Figure 2~4
Figure 5~7
AI summary
Systems (100) and methods (800) for communicating information. The methods comprise: storing message sets in Communication Devices ("CDs") so as to be respectively associated with speaker information; performing operations, by a first CD, to capture an audio message spoken by an individual and to convert the audio message into a message audio file; comparing the message audio file to each reference audio file in the message sets to determine whether one of the reference audio files matches the message audio file by a certain amount; converting the audio message into a text message when a determination is made that a reference audio file does match the message audio file by a certain amount; generating a secure text message by appending the speaker information that is associated with the matching reference audio file to the text message, or by appending other information to the text message; transmitting the secure text message.