Face Voice Similarity Integration for Recognition Accuracy
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing multi-modality recognition technologies face accuracy issues when face image similarity is high but voice data accuracy is low, leading to erroneous determinations due to the influence of noise or other factors.
Innovation Solution
An information processing method that calculates an integrated similarity by combining face and voice similarities, determining the integrated similarity as final when face similarity is within a threshold range and using only face similarity when it's outside this range, thereby enhancing recognition accuracy regardless of voice data accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If integrated similarity is calculated by combining face and voice similarities, then recognition accuracy is improved, but the system becomes more complex and vulnerable to noise in voice data
Solution Approach 1:
The patent applies local quality by differentiating the treatment of face and voice similarities based on their individual reliability. When face similarity is high (indicating reliable facial recognition), the system trusts the face data more and reduces reliance on voice data. When face similarity is low, the system becomes more cautious and may require both modalities or use voice data with lower weight. This dynamic, localized adjustment of quality weights resolves the contradiction by adapting system complexity to actual recognition needs.
Solution Approach 2:
The patent implements dynamics by making the integration weights of face and voice similarities variable rather than fixed. The weights dynamically adjust based on the reliability indicators of each modality (e.g., confidence scores, noise levels). This allows the system to automatically reduce complexity when one modality is unreliable and increase integration when both are reliable, thus improving recognition accuracy without permanently increasing system complexity.
2Adaptability or versatility
If voice data is used for recognition, then recognition capability is enhanced, but accuracy decreases when voice data is affected by noise or errors
Solution Approach 1:
The patent introduces reliability indicators as intermediary elements that mediate between the raw voice/face data and the final recognition decision. These indicators (such as confidence scores, signal quality metrics) act as filters that assess the trustworthiness of each modality before integration. When voice data shows low reliability (due to noise or errors), the intermediary mechanism reduces its influence on the final result, thus maintaining accuracy while preserving the enhanced capability provided by voice recognition.
Solution Approach 2:
The patent applies parameter changes by modifying the weight parameters of voice and face similarities in the integrated similarity calculation based on their reliability. When voice data reliability is low, the weight parameter for voice similarity is reduced; when reliability is high, the weight increases. This dynamic parameter adjustment allows the system to maintain versatile recognition capability using both modalities while protecting accuracy from noise-induced errors by adapting parameters to current data quality conditions.
3Measurement precision
If face similarity is used for recognition, then recognition accuracy is maintained, but the system loses the ability to recognize speakers with similar faces but different voices
Solution Approach 1:
The patent applies merging by combining face similarity and voice similarity into an integrated similarity metric that weighs both modalities according to their reliability. This fusion allows the system to maintain high accuracy when face data is reliable while also capturing speaker identity information from voice data when facial data is insufficient or when both modalities agree. The merging mechanism thus preserves the versatility to recognize speakers with similar faces but different voices, as the voice modality provides complementary information that face-only systems would miss.
Data Source
AI summary
An information processing device performs: acquiring a face similarity indicating a similarity between a face of a first person and a face of a second person; acquiring a voice similarity indicating a similarity between a voice of the first person and a voice of the second person; calculating an integrated similarity by integrating the face similarity and the voice similarity, and determining the integrated similarity as a final similarity when the face similarity falls within an integrated range including a threshold which is used to determine whether the first person and the second person are identical to each other, and calculating the face similarity as a final similarity when the face similarity is out of the integrated range; and outputting the final similarity.


