A personnel identity recognition method, device and medium based on audio fusion

CN120260577BActive Publication Date: 2026-08-21MOBILE TECH COMPANY CHINA TRAVELSKY HLDG +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510410901.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2026-08-21
Estimated Expiration
2045-04-02

AI Technical Summary

Technical Problem

[0002]民航机舱内的高动态声学场景常伴随多机组人员同步语音交互(如飞行员与管制员重叠通话、客舱乘务组协同广播),现有声纹识别技术在此类多人混叠语音场景中存在显著瓶颈:传统单一声纹匹配模型在频谱混叠区域(能量交叠频段>30%)难以有效分离说话人身份,导致目标声纹特征被非目标语音掩盖(信噪比下降至-5dB以下);基于独立语音分离的预处理方法虽能部分抑制串扰,但受限于固定声源数假设与计算时延(通常>200ms),无法实时解析动态变化的多人对话(如3人以上快速指令交替);此外,现有声纹库缺乏多人联合声学特征建模能力,当关键指令需多人协同认证时(如紧急代码双重确认),系统无法从混合语音流中同步提取并验证多重身份,造成航空安全规程执行漏洞

Benefits of technology

本发明的基于音频融合的人员身份识别方法,融合声纹特征库与动态匹配机制,有效解决多人混叠语音的身份识别难题:针对重叠语音的频谱干扰,采用多维度声纹融合与子空间分解技术,实现混叠场景下个体声纹特征的精准剥离与匹配;通过自适应声源分析模型,动态解析并发说话人数量并同步提取特征,避免传统方法对声源数量的依赖;融合声纹特征库支持多人联合生物特征建模,可从单段混合语音中完成多重身份协同验证,满足民航安全规程中的多人员操作认证需求;优化后的实时处理架构显著降低计算负载,确保高噪声环境下多人语音流的快速身份识别与响应,提升航空通信系统的可靠性与安全性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120260577B_ABST
    Figure CN120260577B_ABST
Patent Text Reader

Abstract

This invention provides a method, device, and medium for personnel identification based on audio fusion, relating to the field of personnel identification technology. The method includes: if the x-th fine-grained human voice segment HC to be identified is not determined... x For the corresponding user, the preset fused voiceprint feature list HB is obtained; the HC is obtained. x The first similarity with each fused voiceprint feature in HB is used to obtain HC. x The corresponding first similarity list HD x According to HD x Obtain the first target similarity HD x =MAX(HD) x ); where MAX() is the preset maximum value function; if HD x If '≥H', then HC is determined. x The corresponding users are several known users; where H' is a preset voiceprint feature similarity threshold; the present invention can ensure rapid identity recognition and response of multi-person voice streams in high-noise environments, and improve the reliability and security of aviation communication systems.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of personnel identification technology, and in particular to a personnel identification method, device and medium based on audio fusion. Background Technology

[0002] High-dynamic acoustic scenarios in civil aviation cabins often involve synchronous voice interactions among multiple crew members (such as overlapping conversations between pilots and air traffic controllers, and coordinated announcements by cabin crew). Existing voiceprint recognition technologies face significant bottlenecks in such multi-person aliased voice scenarios: traditional single voiceprint matching models struggle to effectively separate speaker identities in spectral aliasing regions (energy overlap bands > 30%), causing target voiceprint features to be masked by non-target speech (signal-to-noise ratio drops below -5dB); while preprocessing methods based on independent speech separation can partially suppress crosstalk, they are limited by the assumption of a fixed number of sound sources and computational latency (usually > 200ms), making it impossible to analyze dynamically changing multi-person dialogues in real time (such as rapid instruction exchanges among 3 or more people); furthermore, existing voiceprint databases lack the ability to model multi-person joint acoustic features. When critical instructions require multi-person collaborative authentication (such as double confirmation of emergency codes), the system cannot simultaneously extract and verify multiple identities from the mixed speech stream, creating loopholes in the enforcement of aviation safety procedures. Summary of the Invention

[0003] To address the aforementioned technical problems, the technical solution adopted by this invention is as follows: According to a first aspect of this application, a person identification method based on audio fusion is provided, the method comprising the following steps: R100, if the xth fine-grained human voice segment to be identified is not determined (HC) x For the corresponding user, the preset fused voiceprint feature list HB = (HB1, HB2, ..., HB...) is obtained. e ,…,HB g ), e = 1, 2, ..., g; where HB e HB represents the preset e-th fused voiceprint feature, and g represents the preset number of fused voiceprint features. e It is obtained by fusing voice segments from at least two known users; R200, obtain HC x The first similarity with each fused voiceprint feature in HB is used to obtain HC. x The corresponding first similarity list HD x =(HD) x,1 HD x,2 , ..., HD x,e ,…,HD x,g ); where HD x,e For HC x With HB e The first similarity between them; R300, according to HD x Obtain the first target similarity HD x =MAX(HD) x ); where MAX() is the preset function for finding the maximum value; R400, if HD x If '≥H', then HC is determined. x The corresponding users are several known users; where H' is a preset voiceprint feature similarity threshold.

[0004] According to another aspect of this application, a non-transitory computer-readable storage medium is also provided, wherein at least one instruction or at least one program is stored in the storage medium, and the at least one instruction or at least one program is loaded and executed by a processor to implement the above-described audio fusion-based personnel identification method.

[0005] According to another aspect of this application, an electronic device is also provided, including a processor and the aforementioned non-transitory computer-readable storage medium.

[0006] The present invention has at least the following beneficial effects: This invention presents an audio fusion-based personnel identification method that integrates a voiceprint feature database and a dynamic matching mechanism to effectively solve the problem of identification in multi-person aliased speech. To address spectral interference from overlapping speech, multi-dimensional voiceprint fusion and subspace decomposition techniques are employed to achieve accurate extraction and matching of individual voiceprint features in aliased scenarios. An adaptive sound source analysis model dynamically analyzes the number of concurrent speakers and extracts features synchronously, avoiding the dependence of traditional methods on the number of sound sources. The integrated voiceprint feature database supports multi-person joint biometric modeling, enabling multi-person collaborative verification from a single segment of mixed speech, meeting the multi-person operation authentication requirements of civil aviation safety regulations. The optimized real-time processing architecture significantly reduces computational load, ensuring rapid identification and response of multi-person speech streams in high-noise environments, and improving the reliability and security of aviation communication systems. Attached Figure Description

[0007] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0008] Figure 1 A flowchart of a personnel identification method based on audio fusion provided in an embodiment of the present invention. Detailed Implementation

[0009] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0010] It should be noted that, based on this disclosure, those skilled in the art will understand that one aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of aspects set forth herein can be used to implement the device and / or practice the method. Furthermore, this device and / or practice the method can be implemented using other structures and / or functionalities besides one or more of the aspects set forth herein.

[0011] Example 1: The following will refer to Figure 1 The flowchart shown illustrates a person identification method based on audio fusion, which introduces such a method.

[0012] This audio fusion-based personnel identification method includes the following steps: R100, if the xth personal voice segment to be identified is not determined HC x For the corresponding user, the preset fused voiceprint feature list HB = (HB1, HB2, ..., HB...) is obtained. e ,…,HB g ), e = 1, 2, ..., g; where HB e HB represents the preset e-th fused voiceprint feature, and g represents the preset number of fused voiceprint features. e It is obtained by fusing speech segments from at least two known users.

[0013] In this embodiment, the existing method of comparing individual user voiceprints can be used to compare HC. x Perform corresponding user identification; if the xth personal voice segment HC to be identified cannot be determined... x The corresponding user represents HC. x The corresponding user is not in the known user list, or HC x If there is more than one corresponding user, further confirmation is needed.

[0014] Furthermore, HB is obtained through the following steps: R110, obtain the voice segments of each known user to obtain a list of known user voice segments HE = (HE1, HE2, ..., HE...). h ,…,HE k), h=1,2,…,k; where HE h Let k be the audio segment of the h-th known user, and k be the number of known users.

[0015] In this embodiment, the method is applied in the cockpit of an aircraft. For the same aircraft, the users who frequently appear in the cockpit are known. Therefore, it is possible to obtain the voice segments of each known user and thus obtain the HE (Head-to-Head) message.

[0016] R120, fuse at least two known user speech segments in HE to obtain a fused speech segment list HF = (HE1, HE2, ..., HE...). e ,…,HE g ); where HE e The e-th fused speech segment is obtained through fusion; g=2 k -k-1.

[0017] R130 extracts fused speaker features from each fused speech segment in HF to obtain HB.

[0018] In this embodiment, for example: if there are 3 known users, there are 3 ways to combine them in pairs and 1 way to combine them in groups of three, for a total of 4 combinations; the voice segments of two users are fused, the voice segments of three users are fused, and so on, until the voice segments of k users are fused, thus obtaining g fused voice segments; the fusion method can be the superposition of voice segments to simulate the scene of multiple people speaking at the same time; the voiceprint features of the fused voice segments are extracted to obtain HB.

[0019] The aforementioned HB can be obtained in advance and stored in a preset storage location, and can be directly called when needed to improve the efficiency of voice recognition.

[0020] R200, obtain HC x The first similarity with each fused voiceprint feature in HB is used to obtain HC. x The corresponding first similarity list HD x =(HD) x,1 HD x,2 , ..., HD x,e ,…,HD x,g ); where HD x,e For HC x With HB e The first similarity between them.

[0021] In this embodiment, it should be noted that those skilled in the art can use existing vector similarity acquisition methods to obtain HC according to actual needs. x The first similarity with each fused voiceprint feature in HB is not elaborated here.

[0022] R300, according to HD x Obtain the first target similarity HD x =MAX(HD) x ); where MAX() is the preset function for finding the maximum value.

[0023] R400, if HD x If '≥H', then HC is determined. x The corresponding users are several known users; where H' is a preset voiceprint feature similarity threshold.

[0024] In this embodiment, HD x '≥H', means HD x 'The corresponding fused speech segment's sound features and HC x The sound characteristics are the same, which can confirm that HC x The corresponding users are known users; for example: if the known users are user A and user B, then HC x The corresponding users are user A and user B, i.e., HC x The audio clip generated when user A and user B speak simultaneously.

[0025] In this embodiment, before identifying a person's identity through voice segments, it is necessary to determine which segmented human voice segments are included in the voice segment to be identified in order to improve recognition efficiency. This can be achieved through the following steps: Furthermore, after step R400, the method further includes the following steps: R500, if HD x If '<H', then HC is determined. x The corresponding users include users other than the known users.

[0026] In this embodiment, if GD x <H' indicates QC x If none of the initial fused vocal segments in GB can be matched, then QC can be determined. x The corresponding user is a user other than the known users, or a QC user. x Among the multiple users, there are other users besides the known users, thus achieving the purpose of identifying other users; after identifying other users besides the known users, a prompt message can be generated to notify the management personnel that there are other personnel in the cabin.

[0027] In this embodiment, the integration of a voiceprint feature library and a dynamic matching mechanism effectively solves the problem of identity recognition in multi-person aliased speech: To address spectral interference from overlapping speech, multi-dimensional voiceprint fusion and subspace decomposition techniques are employed to achieve accurate extraction and matching of individual voiceprint features in aliased scenarios; an adaptive sound source analysis model dynamically analyzes the number of concurrent speakers and extracts features synchronously, avoiding the dependence of traditional methods on the number of sound sources; the integrated voiceprint feature library supports multi-person joint biometric modeling, enabling multi-identity collaborative verification from a single segment of mixed speech, meeting the multi-person operation authentication requirements in civil aviation safety regulations; the optimized real-time processing architecture significantly reduces computational load, ensuring rapid identity recognition and response for multi-person speech streams in high-noise environments, and improving the reliability and security of aviation communication systems.

[0028] Example 2: Before identifying the person corresponding to the voice segment to be recognized, it is necessary to identify the human voice segment of the voice segment to be recognized. The following methods are provided: R010: Obtain the speech segment to be recognized for a preset duration.

[0029] In this embodiment, the speech segment to be identified can be a speech segment collected by an audio acquisition device in the aircraft cabin; it should be noted that the speech segment includes human voice segments and non-human voice segments, and human voice recognition needs to be performed on the human voice segment to be identified.

[0030] R020 divides the speech segment to be recognized into several consecutive speech frames of the same duration to obtain a speech frame list SA = (SA1, SA2, ..., SA2). i SA n ), i=1, 2,...,n; among them, SA i Let be the i-th speech frame obtained by dividing the speech segment to be recognized, and n be the number of speech frames obtained by dividing the speech segment to be recognized.

[0031] In this embodiment, the speech segment to be recognized can be divided into speech frames according to actual needs. For example, if the duration of each speech frame is set to 10ms, a speech segment to be recognized with a duration of 60s can be divided into 6000 speech frames.

[0032] Furthermore, the duration of the speech frame is determined through the following steps: S210, Obtain the short-time energy corresponding to each sampling point of the speech segment to be recognized, so as to obtain the short-time energy list SK = (SK1, SK2, ..., SK3) corresponding to the speech segment to be recognized. dx , ..., SK dy ), dx = 1, 2, ..., dy; where SK dxLet dx be the short-time energy of the dx-th sampling point corresponding to the speech segment to be recognized, and dy be the number of sampling points corresponding to the speech segment to be recognized.

[0033] In this embodiment, the speech segment to be identified can be digitally processed, and the sampling frequency can be set to obtain the short-time energy corresponding to each sampling point.

[0034] S220, iterate through SK, if SK dx If ≥SK', then SK will be... dx Determined to have short-term audible energy; otherwise, SK dx The energy level is determined to be a short-term energy level without sound; SK' is the preset threshold for a short-term energy level with sound.

[0035] In this embodiment, a higher short-time energy corresponding to a sampling point indicates the presence of sound at that sampling point; otherwise, no sound is present. SK' is an empirical value that can be set according to actual needs or based on the short-time energy corresponding to the average ambient noise.

[0036] S230, divide several consecutive short-time energies with sound in SK into a group to obtain a list of short-time energies with sound SU = (SU1, SU2, ..., SU...). de , ...,SU df ), de = 1, 2, ..., df; where SU de Let df be the de-th short-time energy group with sound obtained from the division, and df be the number of short-time energy groups with sound obtained from the division.

[0037] The speech segment corresponding to the short-time energy group with sound indicates that the speech segment contains sound, while the speech segment corresponding to the short-time energy group without sound indicates that the speech segment does not contain sound.

[0038] S240, obtain the time interval between two adjacent short-time energy groups with sound in SU, to obtain the time interval list Δst = (Δst1, Δst2, ..., Δst dr , …, Δst df-1 ), dr=1,2,…,df-1; where Δst dr For SU dr with SU dr+1 The time interval between them.

[0039] S250, obtain the target time interval st'=MAX(Δst); where MAX() is the preset maximum value function.

[0040] S260, determine the duration SH of the voice frame according to st'; where, if st' < st1, then SH = PA, otherwise, SH = PA × (1 + (st' - st1) / st1) × PA; st1 is a preset time interval threshold, and PA is a preset minimum duration of the voice frame.

[0041] In this embodiment, through the above steps, the dynamic frame length adapts to the speech rhythm. The frame length is adjusted according to the longest pause st', increasing the frame length when speaking slowly (long pause) and decreasing the frame length when speaking quickly (short pause). Long frames: improve frequency domain resolution, suitable for analyzing low-frequency speech features; short frames: improve temporal domain resolution, suitable for capturing rapidly changing voiceless sounds or consonants.

[0042] The system initially separates spoken and silent segments using a short-time energy threshold SK'; then it filters isolated noise points using continuous grouping (S230). This reduces the probability of transient noise (such as coughing or tapping) being misidentified as speech.

[0043] The time interval Δst is used to identify speech segment boundaries; the maximum interval st' helps to distinguish between natural pauses and the end of speech.

[0044] Improve the accuracy of Voice Endpoint Detection (VAD) and avoid truncation or redundancy.

[0045] Dynamic adjustment: Short frames (high real-time performance) and long frames (low computational load) are switched on demand.

[0046] Real-time communication: Short frames are prioritized to ensure low latency; Offline analysis: Longer frames reduce computational resource consumption.

[0047] Furthermore, SK' and st1 can be updated in real time based on environmental noise. Specifically, WebRTC's noise estimation algorithm can be used to determine SK' and st1.

[0048] R030, input SA into the preset VAD model to obtain the voice confidence list corresponding to SA: SA' = (SA'1, SA'2, ..., SA') i , …, SA' n ); where SA' i for SA i The confidence level of the human voice frame.

[0049] In this embodiment, VAD (Voice Activity Detection) can detect whether each speech frame is a human voice. During the detection process, the confidence level of each speech frame in SA (Speech Activity Detection) can be obtained, thus obtaining SA'. It should be noted that those skilled in the art can use the VAD model to obtain the confidence level of human voice according to actual needs, which will not be elaborated here.

[0050] R040, perform a first sliding window operation on SA'; wherein, the first sliding window includes m voice confidence scores.

[0051] In this embodiment, since VAD performs human voice detection on each speech frame, and the duration of each speech frame is very short, a single human voice confidence score cannot directly determine whether the corresponding speech frame is a human voice frame. Therefore, a first sliding window is set to combine multiple consecutive human voice confidence scores to judge human voice frames, so as to improve the accuracy of the judgment. The duration of the first sliding window can be 25ms.

[0052] Furthermore, step R040 includes the following steps: R041, obtain the first preset value SN=1.

[0053] R042 controls the first sliding window so that the confidence score of the first human voice within the first sliding window is SA'. SN .

[0054] R043, if SN < n+1-m, then obtain SN = SN+1 and proceed to S420; otherwise, exit the current processing.

[0055] In this embodiment, the sliding window is operated through the above steps, so that the step size of the sliding window is 1. Steps S500 and S600 are executed once for each sliding window to determine whether the first voice frame in the sliding window is a human voice frame.

[0056] R050: Each time the first sliding window is slid, obtain the number QN of voice confidence scores within the first sliding window that are greater than the preset first voice confidence threshold.

[0057] In this embodiment, the number QN of voice confidence scores greater than a preset first voice confidence threshold within the first sliding window can be determined by comparing each one individually. The first voice confidence threshold is an empirical value that can be obtained through the analysis of a large amount of voice data.

[0058] R060, if QN / m>η1, then the speech frame corresponding to the first human voice confidence score in the first sliding window is determined to be a human voice frame; otherwise, it is determined to be a non-human voice frame; where η1 is the preset first weight.

[0059] In this embodiment, if QN / m > η1, it means that there are a large number of voice confidence scores greater than the first voice confidence threshold within the sliding window. Therefore, it can be determined that most of the corresponding speech frames within the sliding window are voice frames, and thus the speech frame corresponding to the first voice confidence score within the first sliding window can be directly determined to be a voice frame. By combining adjacent voice confidence scores for joint judgment, the situation where a single speech frame error leads to an overall judgment error can be avoided when judging by a single speech frame, thereby improving the accuracy of the judgment.

[0060] R070: Based on the human voice frames and non-human voice frames in SA, determine the human voice segment corresponding to the speech to be recognized.

[0061] Furthermore, step R070 includes the following steps: R071 sets the human voice frames in SA to the first preset character and the non-human voice frames to the second preset character, to obtain the character list ZA corresponding to SA = (ZA1, ZA2, ..., ZA...). j , ..., ZA n+1-m ), j=1, 2,…, n+1-m; ZA j for SA j The corresponding preset character, n+1-m is the number of times the first sliding window is slid; the first preset character and the second preset character are different.

[0062] In this embodiment, the first preset character can be 1 and the second preset character can be 0, thereby realizing the encoding of human voice frames and non-human voice frames. It should be noted that since the length of the first sliding window is m and the total number of voice frames and non-voice frames is n, the first sliding window can only slide n+1-m times, and the remaining m-1 voice frames need to be judged in conjunction with the next voice segment to be recognized.

[0063] R072, iterate through ZA, and divide several consecutive first preset characters in ZA into groups to obtain a list of first preset character groups LA = (LA1, LA2, ..., LA...). p , ..., LA q ), p=1,2,…,q; where LA p Let q be the p-th first preset character group obtained by dividing the first preset character in ZA, and let q be the number of first preset character groups obtained by dividing the first preset character in ZA.

[0064] In this embodiment, since the human voice frame has a certain continuity, the first preset character and the second preset character in ZA may be distributed in a continuous manner. Several consecutive first preset characters in ZA can be divided into a group to obtain LA.

[0065] R073, iterate through LA, if NU p If >NU', then LA will be... p The corresponding recognized speech segment was identified as a human voice segment; among which, NU p For LA p The number of the first preset characters; NU' is the preset first threshold.

[0066] In this embodiment, the user speaks for a certain duration, therefore, when NU p LA will only be applied when >NU'.p The corresponding recognized speech segment is identified as a human voice segment; thus, it can avoid the situation where the confidence judgment of individual speech frames is incorrect, which leads to the error of the overall result, and improve the accuracy of the judgment.

[0067] Furthermore, after step R073, the method further includes the following steps: R074 inputs the determined vocal segments into the preset ASR model to obtain a list of text corresponding to each vocal segment.

[0068] In this embodiment, after obtaining the human voice segments through the above steps, each human voice segment can be input into a preset ASR model to obtain a list of text corresponding to each human voice segment, i.e., text information; thereby realizing cabin human voice recognition.

[0069] Furthermore, step R070 includes the following steps: R071 sets the human voice frames in SA to the first preset character and the non-human voice frames to the second preset character, to obtain the character list ZA corresponding to SA = (ZA1, ZA2, ..., ZA...). j , ..., ZA n+1-m ), j=1, 2,…, n+1-m; ZA j for SA j The corresponding preset character, n+1-m is the number of times the first sliding window is slid; the first preset character and the second preset character are different.

[0070] In this embodiment, the first preset character can be 1 and the second preset character can be 0, thereby realizing the encoding of human voice frames and non-human voice frames. It should be noted that since the length of the first sliding window is m and the total number of voice frames and non-voice frames is n, the first sliding window can only slide n+1-m times, and the remaining m-1 voice frames need to be judged in conjunction with the next voice segment to be recognized.

[0071] R072, iterate through ZA, and divide several consecutive first preset characters in ZA into groups to obtain a list of first preset character groups LA = (LA1, LA2, ..., LA...). p , ..., LA q ), p=1,2,…,q; where LA p Let q be the p-th first preset character group obtained by dividing the first preset character in ZA, and let q be the number of first preset character groups obtained by dividing the first preset character in ZA.

[0072] In this embodiment, since the human voice frame has a certain continuity, the first preset character and the second preset character in ZA may be distributed in a continuous manner. Several consecutive first preset characters in ZA can be divided into a group to obtain LA.

[0073] R073, iterate through LA, if NU p If >NU', then LA will be... p The corresponding recognized speech segment was identified as a human voice segment; among which, NU p For LA p The number of the first preset characters; NU' is the preset first threshold.

[0074] In this embodiment, the user speaks for a certain duration, therefore, when NU p LA will only be applied when >NU'. p The corresponding recognized speech segment is identified as a human voice segment; thus, it can avoid the situation where the confidence judgment of individual speech frames is incorrect, which leads to the error of the overall result, and improve the accuracy of the judgment.

[0075] Furthermore, after step R073, the method further includes the following steps: R074 inputs the determined vocal segments into the preset ASR model to obtain a list of text corresponding to each vocal segment.

[0076] In this embodiment, after obtaining the human voice segments through the above steps, each human voice segment can be input into a preset ASR model to obtain a list of text corresponding to each human voice segment, i.e., text information; thereby realizing cabin human voice recognition.

[0077] Furthermore, after step R072 and before step R073, the method further includes the following steps: R71, iterate through LA, if NU p >MU', then LA p The second preset character is then modified to the first preset character in a first preset number of consecutive sequences; where MU' is the preset second threshold.

[0078] In this embodiment, there is a situation in the above steps where, if the time between two adjacent voice segments is long, the last few voice frames in the previous voice segment may be judged as non-voice frames because a large number of non-voice frames exist within the first sliding window. To avoid this situation, the first preset number of second preset characters can be modified to the first preset character. The first preset number is an empirical value that can be obtained through the analysis and statistics of a large amount of voice data.

[0079] R72 adds the modified first preset number of first preset characters to LA. p In the middle, the modified first preset character group list LA' = (LA'1, LA'2, ..., LA') corresponding to LA is obtained. p , ..., LA' q); where LA' p For LA p The corresponding modified first preset character group.

[0080] R73, iterate through LA, if SU p If >NU', then LA' will be... p The corresponding recognized speech segment was identified as a human voice segment; SU p For LA' p The number of the first preset characters.

[0081] In this embodiment, the initially determined voice segment is extended for a certain duration through the above steps, so that the extended voice segment can completely cover the speech generated when talking to the user, thereby improving the accuracy of voice recognition.

[0082] Furthermore, after step R070, the method further includes the following steps: R080, in response to acquiring a new speech segment to be recognized; wherein, the duration of the new speech segment to be recognized is ST'=ST-ΔST; ΔST is the duration corresponding to m-1 speech frames.

[0083] R081 adds the last m-1 speech frames in SA before the first speech frame in the speech frame list corresponding to the new speech segment to be recognized, so as to obtain the speech frame list corresponding to the new speech segment to be recognized.

[0084] R082: Determine the corresponding human voice segment based on the list of voice frames corresponding to the new voice segment to be recognized.

[0085] In this embodiment, the remaining m-1 speech frames were not judged in the above steps, while the speech segments to be recognized were all consecutive. Therefore, the remaining m-1 speech frames were added before the first speech frame in the speech frame list corresponding to the new speech segment to be recognized, and then the human voice recognition in the above steps was performed.

[0086] In this embodiment, by combining the VAD model and the sliding window statistical method, the influence of instantaneous noise or brief non-human voice interference can be effectively reduced, improving the accuracy of human voice frame judgment. Using multi-frame confidence statistics instead of single-frame threshold judgment reduces the probability of misjudgment in high-noise environments, making it particularly suitable for complex acoustic environments such as aircraft cabins. Dynamic analysis of the confidence distribution of continuous speech frames allows for more precise location of the start and end points of human voice segments, avoiding segment truncation or redundancy problems caused by fixed windows or thresholds in traditional methods. By first screening high-quality human voice frames, invalid calculations in subsequent ASR processing can be reduced, improving the overall operating efficiency of the speech recognition system. By adjusting parameters such as the sliding window size and confidence threshold, it can adapt to different cabin environments, such as passenger cabins and cockpits, for speech detection needs. The solution of this invention, while improving the accuracy and stability of human voice detection, also has high scalability and applicability.

[0087] Furthermore, although the steps of the method in this disclosure are described in a specific order in the accompanying drawings, this does not require or imply that the steps must be performed in that specific order, or that all the steps shown must be performed to achieve the desired result. Additional or alternative steps may be omitted, multiple steps may be combined into one step, and / or a step may be broken down into multiple steps.

[0088] Embodiments of the present invention also provide a non-transitory computer-readable storage medium that can be disposed in an electronic device to store at least one instruction or at least one program related to implementing a method in the method embodiments, wherein the at least one instruction or the at least one program is loaded and executed by the processor to implement the method provided in the above embodiments.

[0089] The program product may employ any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0090] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium, capable of sending, propagating, or transmitting programs for use by or in conjunction with an instruction execution system, apparatus, or device.

[0091] The program code contained on the readable medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination thereof.

[0092] Program code for performing the operations of this application can be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java and C++, and conventional procedural programming languages ​​such as C or similar languages. The program code can execute entirely on the user's computing device, partially on the user's device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0093] Embodiments of the present invention also provide an electronic device, including a processor and the aforementioned non-transitory computer-readable storage medium.

[0094] The electronic device is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments in this application.

[0095] Electronic devices are manifested in the form of general-purpose computing devices. Components of an electronic device may include, but are not limited to: at least one processor, at least one memory, and a bus connecting different system components (including memory and processor).

[0096] The memory stores program code that can be executed by the processor, causing the processor to perform the steps in the various embodiments described in this specification.

[0097] The memory may include readable media in the form of volatile memory, such as random access memory (RAM) and / or cache memory, and may further include read-only memory (ROM).

[0098] The memory may also include programs / utilities having a set (at least one) of program modules, including but not limited to: an operating system, one or more application programs, other program modules, and program data, each or some combination of these examples may include an implementation of a network environment.

[0099] A bus can represent one or more of several bus structures, including a memory bus or memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus that uses any of the various bus structures.

[0100] The electronic device can also communicate with one or more external devices (e.g., keyboards, pointing devices, Bluetooth devices, etc.), one or more devices that enable a user to interact with the electronic device, and / or any device that enables the electronic device to communicate with one or more other computing devices (e.g., routers, modems, etc.). This communication can be performed via input / output (I / O) interfaces. Furthermore, the electronic device can communicate with one or more networks (e.g., local area networks (LANs), wide area networks (WANs), and / or public networks, such as the Internet) via a network adapter. The network adapter communicates with other modules of the electronic device via a bus. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with the electronic device, including but not limited to: microcode, device drivers, redundant processors, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0101] From the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, terminal device, or network device, etc.) to execute the methods according to the embodiments of this disclosure.

[0102] Embodiments of the present invention also provide a computer program product including program code, which, when the program product is run on an electronic device, causes the electronic device to perform the steps of the methods described above in various exemplary embodiments of the present invention.

[0103] While specific embodiments of the invention have been described in detail by way of examples, those skilled in the art should understand that the examples are for illustrative purposes only and are not intended to limit the scope of the invention. Those skilled in the art should also understand that various modifications can be made to the embodiments without departing from the scope and spirit of the invention.

Claims

1. A method for personnel identification based on audio fusion, characterized in that, The method includes the following steps: R010, acquire the speech segment to be recognized with a preset duration; R020 divides the speech segment to be recognized into several consecutive speech frames of the same duration to obtain a speech frame list SA = (SA1, SA2, ..., SA2). i SA n ), i=1, 2,...,n; among them, SA i Let be the i-th speech frame obtained by dividing the speech segment to be recognized, and n be the number of speech frames obtained by dividing the speech segment to be recognized. R030, input SA into the preset VAD model to obtain the voice confidence list corresponding to SA: SA' = (SA'1, SA'2, ..., SA') i , …, SA' n ); where SA' i for SA i The confidence level of the human voice frame; R040, Perform a first sliding window operation on SA'; wherein, the first sliding window includes m voice confidence scores; R050, each time the first sliding window is slid, obtain the number QN of voice confidence scores within the first sliding window that are greater than the preset first voice confidence threshold; R060, if QN / m>η1, then the speech frame corresponding to the first human voice confidence score in the first sliding window is determined to be a human voice frame; otherwise, it is determined to be a non-human voice frame; where η1 is the preset first weight; R071 sets the human voice frames in SA to the first preset character and the non-human voice frames to the second preset character, to obtain the character list ZA corresponding to SA = (ZA1, ZA2, ..., ZA...). j , ..., ZA n+1-m ), j=1, 2,…, n+1-m; ZA j for SA j The corresponding preset character, n+1-m is the number of times the first sliding window is slid; the first preset character and the second preset character are different; R072, iterate through ZA, and divide several consecutive first preset characters in ZA into groups to obtain a list of first preset character groups LA = (LA1, LA2, ..., LA...). p , ..., LA q ), p=1,2,…,q; where LA p Let q be the p-th first preset character group obtained by dividing the first preset character in ZA, and q be the number of first preset character groups obtained by dividing the first preset character in ZA. R71, iterate through LA, if NU p >MU', then LA p The second preset character is then modified to the first preset character after a first preset number of consecutive occurrences; where MU' is the preset second threshold; NU p For LA p The number of the first preset characters; R72 adds the modified first preset number of first preset characters to LA. p In the middle, the modified first preset character group list LA' = (LA'1, LA'2, ..., LA') corresponding to LA is obtained. p , ..., LA' q ); where LA' p For LA p The corresponding modified first preset character group; R73, iterate through LA, if SU p If >NU', then LA' will be... p The corresponding recognized speech segment was identified as a human voice segment; SU p For LA' p The number of the first preset characters; NU' is the preset first threshold; R073, iterate through LA, if NU p If >NU', then LA will be... p The corresponding recognized speech segment was identified as a human voice segment; R100, if the xth personal voice segment to be identified is not determined HC x For the corresponding user, the preset fused voiceprint feature list HB = (HB1, HB2, ..., HB...) is obtained. e ,…,HB g ), e = 1, 2, ..., g; where HB e HB represents the preset e-th fused voiceprint feature, and g represents the preset number of fused voiceprint features. e It is obtained by fusing voice segments from at least two known users; R200, obtain HC x The first similarity with each fused voiceprint feature in HB is used to obtain HC. x The corresponding first similarity list HD x =(HD) x,1 HD x,2 , ..., HD x,e ,…,HD x,g ); where HD x,e For HC x With HB e The first similarity between them; R300, according to HD x Obtain the first target similarity HD x =MAX(HD) x ); where MAX() is the preset function for finding the maximum value; R400, if HD x If '≥H', then HC is determined. x The corresponding users are several known users; where H' is a preset voiceprint feature similarity threshold.

2. The personnel identification method based on audio fusion according to claim 1, characterized in that, Following step R400, the method further includes the following steps: R500, if HD x If '<H', then HC is determined. x The corresponding users include users other than the known users.

3. The personnel identification method based on audio fusion according to claim 1, characterized in that, HB is obtained through the following steps: R110, obtain the voice segments of each known user to obtain a list of known user voice segments HE = (HE1, HE2, ..., HE...). h ,…,HE k ), h=1,2,…,k; where HE h Let k be the audio segment of the h-th known user, and k be the number of known users. R120, fuse at least two known user speech segments in HE to obtain a fused speech segment list HF = (HE1, HE2, ..., HE...). e ,…,HE g ); where HE e The e-th fused speech segment is obtained through fusion; g=2 k -k-1; R130 extracts fused speaker features from each fused speech segment in HF to obtain HB.

4. The personnel identification method based on audio fusion according to claim 1, characterized in that, Following step R073, the method further includes the following steps: R074 inputs the determined vocal segments into the preset ASR model to obtain a list of text corresponding to each vocal segment.

5. The personnel identification method based on audio fusion according to claim 1, characterized in that, Step R040 includes the following steps: R041, obtain the first preset value SN=1; R042 controls the first sliding window so that the confidence score of the first human voice within the first sliding window is SA'. SN ; R043, if SN < n+1-m, then obtain SN = SN+1 and proceed to S420; otherwise, exit the current processing.

6. A non-transitory computer-readable storage medium, wherein the storage medium stores at least one instruction or at least one program segment, characterized in that, The at least one instruction or the at least one program segment is loaded and executed by the processor to implement the audio fusion-based personnel identification method as described in any one of claims 1-5.

7. An electronic device, characterized in that, Includes a processor and the non-transitory computer-readable storage medium as described in claim 6.

Citation Information

Patent Citations

  • Speech data processing method and device, computer device and storage medium

    CN108877775A

  • Automatic identity recognition method based on voiceprint information of speaker

    CN113113022A

  • Voice separation method based on multi-speaker voice detection

    CN116935884A

  • Speech recognition method and device

    CN117392984A

  • Application of VAD method based on deep learning in speech recognition system

    CN119091932A