Personnel identity recognition method and device based on audio fusion and medium
Through the personnel identity recognition method of audio fusion, multi-dimensional voiceprint fusion and subspace decomposition technology are used to solve the identity recognition problem of multi-person aliased voice in civil aviation cabins, and fast identity recognition and multi-person collaborative verification in high-noise environments are realized, which improves the security and reliability of the aviation communication system.
Patent Information
- Application Number
- CN202510410901.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-04-02
AI Technical Summary
The existing voiceprint recognition technology is difficult to separate the speaker's identity in the multi-person aliased voice scene in the civil aviation cabin, and cannot analyze dynamically changing multi-person conversations in real time, and lacks the ability to model multiple-person joint acoustic feature, resulting in loopholes in the execution of aviation safety procedures.
The personnel identity recognition method based on audio fusion is adopted, and the number of speakers is dynamically analyzed through multi-dimensional voiceprint fusion and subspace decomposition technology, and the adaptive sound source analysis model is used to model multiple people's joint biometric features to achieve accurate peeling and matching of individual voiceprint features in aliasing scenarios.
Realize fast identity recognition and response of multi-person voice streams in high-noise environments, improve the reliability and security of the aviation communication system, and meet the needs of multi-person operation authentication.
Smart Images

Figure CN120260577A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of personnel identity recognition, and in particular to a personnel identity recognition method, device and medium based on audio fusion. Background Art
[0002] The high-dynamic acoustic scene in a civil aviation cabin is often accompanied by multi-crew synchronous voice interactions (such as pilots overlapping calls with air traffic controllers, cabin crew collaborative broadcasts). Existing speaker recognition technologies have significant bottlenecks in such multi-person overlapping voice scenarios: traditional single-speaker matching models are difficult to effectively separate speaker identities in the spectral overlapping region (energy overlapping frequency band > 30%), resulting in target speaker features being masked by non-target speech (signal-to-noise ratio drops below -5dB); although the preprocessing method based on independent voice separation can partially suppress crosstalk, it is limited by the fixed sound source number assumption and calculation delay (usually > 200ms), and cannot real-time analyze dynamic multi-person conversations (such as rapid command alternation among more than 3 people); in addition, existing speaker databases lack the ability to model joint acoustic features of multiple people. When key commands require multi-person collaborative authentication (such as double confirmation of emergency codes), the system cannot synchronously extract and verify multiple identities from the mixed voice stream, resulting in loopholes in the implementation of aviation safety regulations. Summary of the Invention
[0003] For the above technical problems, the technical solution adopted by the present invention is as follows:
[0004] According to the first aspect of the present application, a personnel identity recognition method based on audio fusion is provided. The method includes the following steps:
[0005] R100, if the user corresponding to the x-th fine-grained human voice segment HC to be recognized is not determined, obtain the preset fused speaker feature list HB = (HB1, HB2,..., HB x ,..., HB e ,..., HB g ), where e = 1, 2,..., g; among them, HB e is the preset e-th fused speaker feature, and g is the number of preset fused speaker features; HB e is obtained by fusing voice segments of at least two known users;
[0006] R200, obtain the first similarity between HC x and each fused speaker feature in HB to obtain the first similarity list HD x corresponding to HC x = (HD x,1 , HD x,2 ,..., HD x,e ,..., HD x,g ); among them, HD x,e is the first similarity corresponding to HCx The first similarity with HB e ;
[0007] R300, obtain the first target similarity HD according to HD x , and obtain the first target similarity HD x ’ = MAX(HD x ); where MAX() is a preset maximum value function;
[0008] R400, if HD x ’ ≥ H’, then determine that the user corresponding to HC x is several known users corresponding to H’; where H’ is a preset voiceprint feature similarity threshold.
[0009] According to another aspect of the present application, there is also provided a non-transitory computer-readable storage medium, in which at least one instruction or at least one program is stored, and the at least one instruction or at least one program is loaded and executed by a processor to implement the above-mentioned personnel identity recognition method based on audio fusion.
[0010] According to another aspect of the present application, there is also provided an electronic device, including a processor and the above-mentioned non-transitory computer-readable storage medium.
[0011] The present invention has at least the following beneficial effects:
[0012] The personnel identity recognition method based on audio fusion of the present invention combines a voiceprint feature library and a dynamic matching mechanism, effectively solving the problem of identity recognition of multi-person overlapping voices: for the spectral interference of overlapping voices, multi-dimensional voiceprint fusion and subspace decomposition technologies are adopted to achieve accurate separation and matching of individual voiceprint features in overlapping scenarios; through an adaptive sound source analysis model, the number of concurrent speakers is dynamically analyzed and features are extracted synchronously, avoiding the dependence on the number of sound sources in traditional methods; the integrated voiceprint feature library supports multi-person joint biometric modeling, and multiple identity collaborative verification can be completed from a single segment of mixed voice, meeting the multi-person operation authentication requirements in civil aviation safety regulations; the optimized real-time processing architecture significantly reduces the computational load, ensuring fast identity recognition and response of multi-person voice streams in high-noise environments, and improving the reliability and security of the aviation communication system. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0014] Figure 1Flowchart of the method for identifying a person's identity based on audio fusion provided by an embodiment of the present invention. Detailed implementation manners
[0015] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts fall within the protection scope of the present invention.
[0016] It should be noted that based on this disclosure, those skilled in the art should understand that one aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of aspects described herein can be used to implement the device and / or practice the method. In addition, this device and / or this method can be implemented using other structures and / or functions in addition to one or more of the aspects described herein.
[0017] Embodiment 1:
[0018] Next, a method for identifying a person's identity based on audio fusion will be introduced with reference to Figure 1 the flowchart of the method for identifying a person's identity based on audio fusion shown.
[0019] The method for identifying a person's identity based on audio fusion includes the following steps:
[0020] R100, if the user corresponding to the x-th voice segment HC to be identified has not been determined x then obtain the preset fused voiceprint feature list HB = (HB1, HB2,..., HB e ,..., HB g ), e = 1, 2,..., g; where HB e is the preset e-th fused voiceprint feature, and g is the number of preset fused voiceprint features; HB e is obtained by fusing voice segments of at least two known users.
[0021] In this embodiment, the existing method of comparing single-user voiceprints can be used to identify the corresponding user of HC x . If the user corresponding to the x-th voice segment HC to be identified cannot be determined x , it means that the user corresponding to HC x is not in the known user list, or there is more than one user corresponding to HC x , and further determination is required.
[0022] Further, HB is obtained through the following steps:
[0023] R110. Obtain the voice segments of each known user to obtain a list of known user voice segments HE = (HE1, HE2, …, HE h , …, HE k ), where h = 1, 2, …, k; among them, HE h is the voice segment of the h-th known user, and k is the number of known users.
[0024] In this embodiment, the scenario to which the method is applied is the cockpit of an airplane. For the same airplane, the users who often appear in the cockpit are known. Therefore, the voice segments of each known user can be obtained, and then HE can be obtained.
[0025] R120. Fuse at least two known user voice segments in HE to obtain a list of fused voice segments HF = (HE1, HE2, …, HE e , …, HE g ); where HE e is the e-th fused voice segment obtained by fusion; g = 2 k - k - 1.
[0026] R130. Extract the fused voiceprint features of each fused voice segment in HF to obtain HB.
[0027] In this embodiment, for example: if there are 3 known users, the number of pairwise combinations is 3, and the number of three-person combinations is 1, for a total of 4 combinations; fuse the voice segments of two known users, then the voice segments of three known users, and so on until the voice segments of k known users are fused, so as to obtain g fused voice segments; the fusion method can be the superposition of voice segments to imitate the actual scenario of multiple people speaking simultaneously; extract the voiceprint features of the fused voice segments to obtain HB.
[0028] The above HB can be obtained in advance and stored in a preset storage location. When needed, it can be directly called to improve the efficiency of human voice recognition.
[0029] R200. Obtain the first similarity between HC x and each fused voiceprint feature in HB to obtain a list of first similarities HD x corresponding to HC x = (HD x,1 , HD x,2 , …, HD x,e , …, HD x,g ); where HD x,e is the first similarity between HC x and HB e .
[0030] In this embodiment, it should be noted that those skilled in the art can use the existing vector similarity acquisition method according to actual needs to obtain HC x The first similarity between each fused voiceprint feature in HC and HB will not be elaborated here.
[0031] R300, according to HD x , obtain the first target similarity HD x ’ = MAX(HD x ); where MAX() is a preset maximum value function.
[0032] R400, if HD x ’ ≥ H’, then determine that the user corresponding to HC x is one of several known users corresponding to H’; where H’ is a preset voiceprint feature similarity threshold.
[0033] In this embodiment, HD x ’ ≥ H’ means that the voice characteristics of the fused voice segment corresponding to HD x are the same as those of HC x , and it can be determined that the user corresponding to HC x is the known user corresponding to H’; for example: the known users corresponding to H’ are user A and user B, then, the user corresponding to HC x is user A and user B, that is, HC x is a voice segment generated by user A and user B speaking simultaneously.
[0034] In this embodiment, before identifying the identity of a person through a voice segment, it is necessary to determine which segments in the voice segment to be identified are human voice segments to improve the identification efficiency, which can be achieved through the following steps:
[0035] Further, after step R400, the method further includes the following steps:
[0036] R500, if HD x ’ < H’, then determine that the users corresponding to HC x include other users besides known users.
[0037] In this embodiment, if GD x < H’, it means that QC x cannot match any of the initial fused human voice segments in GB. At this time, it can be determined that the user corresponding to QC x is other users besides known users, or QC xAmong the corresponding multiple users, there are other users besides the known users, so as to achieve the purpose of identifying other users; after identifying other users besides the known users, a prompt message can be generated to prompt the management personnel that there are other people in the cabin.
[0038] In this embodiment, by integrating the voiceprint feature library and the dynamic matching mechanism, the problem of identity recognition of multi-person overlapping voices is effectively solved: for the spectral interference of overlapping voices, multi-dimensional voiceprint fusion and subspace decomposition technologies are adopted to achieve accurate extraction and matching of individual voiceprint features in overlapping scenarios; through the adaptive sound source analysis model, the number of concurrent speakers is dynamically analyzed and features are extracted synchronously, avoiding the dependence on the number of sound sources in traditional methods; the integrated voiceprint feature library supports multi-person joint biometric modeling and can complete multiple identity collaborative verification from a single segment of mixed voice, meeting the multi-person operation authentication requirements in civil aviation safety regulations; the optimized real-time processing architecture significantly reduces the computational load, ensuring fast identity recognition and response of multi-person voice streams in high-noise environments and improving the reliability and security of the aviation communication system.
[0039] Embodiment Two:
[0040] Before performing identity recognition on the person corresponding to the voice to be recognized, it is necessary to recognize the human voice segment of the voice segment to be recognized, and the following method is provided:
[0041] R010, obtain the voice segment to be recognized with a preset duration.
[0042] In this embodiment, the voice segment to be recognized can be the voice segment collected by the audio acquisition device in the aircraft cabin; it should be noted that this voice segment contains a human voice segment and a non-human voice segment, and it is necessary to perform human voice recognition on the human voice segment to be recognized.
[0043] R020, divide the voice segment to be recognized into several consecutive voice frames with the same duration to obtain a voice frame list SA = (SA1, SA2,..., SA i ,..., SA n ), i = 1, 2,..., n; where SA i is the i-th voice frame obtained by dividing the voice segment to be recognized, and n is the number of voice frames obtained by dividing the voice segment to be recognized.
[0044] In this embodiment, the voice segment to be recognized can be divided into voice frames according to actual needs. For example, if the duration of each voice frame is set to 10 ms, for a voice segment to be recognized with a duration of 60 s, 6000 voice frames can be obtained.
[0045] Further, the duration of the voice frame is determined through the following steps:
[0046] S210. Obtain the short-time energy corresponding to each sampling point of the speech segment to be recognized, so as to obtain the short-time energy list SK = (SK1, SK2, …, SK dx , …, SK dy ), where dx = 1, 2, …, dy; among them, SK dx is the short-time energy of the dx-th sampling point corresponding to the speech segment to be recognized, and dy is the number of sampling points corresponding to the speech segment to be recognized.
[0047] In this embodiment, the speech segment to be recognized can be digitally processed, and the sampling frequency is set to obtain the short-time energy corresponding to each sampling point.
[0048] S220. Traverse SK. If SK dx ≥SK’, then determine SK dx as the short-time energy of the voice; otherwise, determine SK dx as the short-time energy of no voice; SK’ is a preset short-time energy threshold for voice.
[0049] In this embodiment, a higher short-time energy corresponding to a sampling point indicates that there is voice at this sampling point; otherwise, there is no voice; SK’ is an empirical value and can be set according to actual needs or according to the short-time energy corresponding to the average ambient noise.
[0050] S230. Divide several consecutive short-time energies of voice in SK into a group to obtain the short-time energy group list SU = (SU1, SU2, …, SU de , …, SU df ), where de = 1, 2, …, df; among them, SU de is the de-th short-time energy group obtained by division, and df is the number of short-time energy groups obtained by division.
[0051] The speech segment corresponding to the short-time energy group of voice indicates that there is voice in this speech segment, and the speech segment corresponding to the short-time energy group of no voice indicates that there is no voice in this speech segment.
[0052] S240. Obtain the time interval between two adjacent short-time energy groups in SU to obtain the time interval list Δst = (Δst1, Δst2, …, Δst dr , …, Δst df-1 ), where dr = 1, 2, …, df - 1; among them, Δst dr is the time interval between SU dr and SU dr+1 .
[0053] S250, Obtain the target time interval st’ = MAX(Δst); where MAX() is a preset maximum value function.
[0054] S260, Determine the duration SH of the speech frame according to st’; where if st’ < st1, then SH = PA, otherwise, SH = PA × (1 + (st’ - st1) / st1) × PA; st1 is a preset time interval threshold, and PA is the minimum duration of a preset speech frame.
[0055] In this embodiment, through the above steps, the dynamic frame length adapts to the speech rhythm, adjusts the frame length according to the longest pause st’, increases the frame length when speaking slowly (long pause), and decreases the frame length when speaking quickly (short pause). Long frame: Improve the frequency domain resolution, suitable for analyzing low-frequency speech features; short frame: Improve the time domain resolution, suitable for capturing quickly changing voiceless sounds or consonants.
[0056] Preliminarily separate the voiced / unvoiced segments through the short-time energy threshold SK’; filter out isolated noise points by combining continuous grouping (S230). Reduce the probability of misjudging transient noises (such as coughing, knocking) as speech.
[0057] Use the time interval Δst to identify the speech segment boundaries; the maximum interval st’ helps to distinguish natural pauses from the end of speech.
[0058] Improve the accuracy of voice activity detection (VAD), and avoid truncation or redundancy.
[0059] Dynamic adjustment: Switch between short frames (high real-time performance) and long frames (low computational complexity) as needed.
[0060] Real-time communication: Give priority to short frames to ensure low latency; offline analysis: Use long frames to reduce computational resource consumption.
[0061] Furthermore, SK’ and st1 can be updated in real time based on the ambient noise. Specifically, the noise estimation algorithm of WebRTC can be used to determine SK’ and st1.
[0062] R030, Input SA into a preset VAD model to obtain the list of voice confidence levels SA’ = (SA’1, SA’2, …, SA’ i , …, SA’ n ); where SA’ i is the confidence level of the voice frame corresponding to SA i .
[0063] In this embodiment, VAD is voice activity detection, which can detect whether each speech frame is a human voice. During the detection process, the confidence of each speech frame in SA being a human voice frame can be obtained, thereby obtaining SA'. It should be noted that those skilled in the art can use the VAD model to obtain the confidence of the human voice according to actual needs, which will not be elaborated here.
[0064] R040, perform a first sliding window operation on SA'; wherein, the first sliding window includes m confidences of the human voice.
[0065] In this embodiment, since VAD detects the human voice for each speech frame and the duration of each speech frame is very short, a single confidence of the human voice cannot directly determine whether the corresponding speech frame is a human voice frame; therefore, a first sliding window is set to combine multiple consecutive confidences of the human voice for judging the human voice frame to improve the accuracy of the judgment; the duration of the first sliding window can be 25 ms.
[0066] Further, step R040 includes the following steps:
[0067] R041, obtain the first preset value SN = 1.
[0068] R042, control the first sliding window so that the first confidence of the human voice within the first sliding window is SA' SN 。
[0069] R043, if SN < n + 1 - m, then obtain SN = SN + 1 and enter S420; otherwise, jump out of the current process.
[0070] In this embodiment, through the above steps, the sliding window is operated so that the step size of the sliding window is 1, and steps S500 and S600 are executed once each time the sliding window slides to determine whether the first speech frame within the sliding window is a human voice frame.
[0071] R050, each time the first sliding window slides, obtain the number QN of confidences of the human voice within the first sliding window that are greater than the preset first confidence threshold of the human voice.
[0072] In this embodiment, the number QN of confidences of the human voice within the first sliding window that are greater than the preset first confidence threshold of the human voice can be determined by comparing one by one; the first confidence threshold of the human voice is an empirical value and can be obtained through analyzing a large amount of human voice data.
[0073] R060, if QN / m > η1, then determine that the speech frame corresponding to the first confidence of the human voice within the corresponding first sliding window is a human voice frame; otherwise, determine it as a non-human voice frame; wherein, η1 is the preset first weight.
[0074] In this embodiment, if QN / m > η1, it indicates that the number of voice confidence levels greater than the first voice confidence threshold within the sliding window is relatively large. Therefore, it can be determined that most of the voice frames corresponding to the sliding window are voice frames, and thus it can be directly determined that the voice frame corresponding to the first voice confidence level within the first sliding window is a voice frame. By combining adjacent voice confidence levels for joint judgment, it is possible to avoid the situation where an incorrect judgment of a single voice frame leads to an incorrect overall judgment when judging by a single voice frame, thereby improving the accuracy of the judgment.
[0075] R070. Determine the voice segment corresponding to the voice to be recognized based on the voice frames and non-voice frames in SA.
[0076] Further, step R070 includes the following steps:
[0077] R071. Set the voice frames in SA as the first preset character and the non-voice frames as the second preset character to obtain the character list ZA = (ZA1, ZA2,..., ZA j ,..., ZA n+1-m ), j = 1, 2,..., n + 1 - m; ZA j is the preset character corresponding to SA j , and n + 1 - m is the number of times the first sliding window slides; the first preset character and the second preset character are different.
[0078] In this embodiment, the first preset character can be 1 and the second preset character can be 0, so as to realize the encoding of voice frames and non-voice frames. It should be noted that since the length of the first sliding window is m and the total number of voice frames and non-voice frames is n, the first sliding window can only slide n + 1 - m times, and the remaining m - 1 voice frames need to be judged in combination with the next voice segment to be recognized.
[0079] R072. Traverse ZA and divide several consecutive first preset characters in ZA into a group to obtain the first preset character group list LA = (LA1, LA2,..., LA p ,..., LA q ), p = 1, 2,..., q; where LA p is the pth first preset character group obtained by dividing the first preset characters in ZA, and q is the number of first preset character groups obtained by dividing the first preset characters in ZA.
[0080] In this embodiment, since voice frames have a certain continuity, the first preset characters and the second preset characters in ZA may show a continuous distribution. Therefore, several consecutive first preset characters in ZA can be divided into a group to obtain LA.
[0081] R073. Traverse LA. If NUp > NU', then LA p The corresponding recognized voice segment is determined as a human voice segment; where NU p is LA p the number of the first preset characters in; NU' is a preset first threshold value.
[0082] In this embodiment, when the user is speaking, it will last for a certain period of time. Therefore, when NU p > NU', then LA p The corresponding recognized voice segment is determined as a human voice segment; thus, it can avoid the situation where the confidence judgment of individual voice frames is incorrect and leads to an incorrect overall result, and improve the accuracy of judgment.
[0083] Further, after step R073, the method further includes the following steps:
[0084] R074, input the determined human voice segment into a preset ASR model to obtain a text list corresponding to each human voice segment.
[0085] In this embodiment, after obtaining the human voice segment through the above steps, each human voice segment can be input into a preset ASR model to obtain a text list corresponding to each human voice segment, that is, text information; thus, the recognition of the human voice in the cabin is realized.
[0086] Further, step R070 includes the following steps:
[0087] R071, set the human voice frames in SA as the first preset characters and the non-human voice frames as the second preset characters to obtain a character list ZA = (ZA1, ZA2,..., ZA j ,..., ZA n+1-m ), j = 1, 2,..., n + 1 - m; ZA j is the preset character corresponding to SA j , n + 1 - m is the number of times the first sliding window slides; the first preset character and the second preset character are different.
[0088] In this embodiment, the first preset character can be 1 and the second preset character can be 0, so as to realize the encoding of the human voice frames and non-human voice frames; it should be noted that since the length of the first sliding window is m and the total number of voice frames and non-voice frames is n, the first sliding window can only slide n + 1 - m times, and the remaining m - 1 voice frames need to be combined with the next voice segment to be recognized for judgment.
[0089] R072, traverse ZA, and divide several consecutive first preset characters in ZA into a group to obtain a first preset character group list LA = (LA1, LA2,..., LA p ,..., LAq ), p = 1, 2, …, q; where LA p is the p-th first preset character group obtained by dividing the first preset characters in ZA, and q is the number of first preset character groups obtained by dividing the first preset characters in ZA.
[0090] In this embodiment, since the voice frames have a certain continuity, therefore, the first preset characters and the second preset characters in ZA may show a continuous distribution, and several consecutive first preset characters in ZA can be divided into a group to obtain LA.
[0091] R073, traverse LA, if NU p > NU', then determine the recognition voice segment corresponding to LA p as a voice segment; where NU p is the number of first preset characters in LA p ; NU' is a preset first threshold.
[0092] In this embodiment, when the user is speaking, it will last for a certain period of time. Therefore, when NU p > NU', the recognition voice segment corresponding to LA p will be determined as a voice segment; thus, it can avoid the situation where the confidence judgment of individual voice frames is incorrect and leads to an incorrect overall result, and improve the accuracy of the judgment.
[0093] Further, after step R073, the method further includes the following steps:
[0094] R074, input the determined voice segment into a preset ASR model to obtain a text list corresponding to each voice segment.
[0095] In this embodiment, after obtaining the voice segment through the above steps, each voice segment can be input into a preset ASR model to obtain a text list corresponding to each voice segment, that is, text information; thus, the recognition of cabin voices is realized.
[0096] Further, after step R072 and before step R073, the method further includes the following steps:
[0097] R71, traverse LA, if NU p > MU', then modify the next consecutive first preset number of second preset characters after LA p to the first preset character; where MU' is a preset second threshold.
[0098] In this embodiment, in the above steps, there will be a situation where, if the time interval between adjacent human voice segments is relatively long, it may cause the last few human voice frames in the previous human voice segment to be judged as non-human voice frames because a relatively large number of non-human voice frames exist within the first sliding window. To avoid the occurrence of the above situation, a first preset number of second preset characters can be modified to first preset characters; the first preset number is an empirical value and can be obtained through statistical analysis of a large amount of voice data.
[0099] R72, Add the modified first preset characters of the first preset number to LA p to obtain the modified first preset character group list LA’=(LA’1, LA’2, …, LA’ p , …, LA’ q ); where LA’ p is the modified first preset character group corresponding to LA p .
[0100] R73, Traverse LA. If SU p > NU’, then determine the recognized voice segment corresponding to LA’ p as a human voice segment; SU p is the number of first preset characters in LA’ p .
[0101] In this embodiment, through the above steps, the initially determined human voice segment is extended by a certain duration so that the extended human voice segment can completely cover the voice generated when the user is speaking, improving the accuracy of human voice recognition.
[0102] Further, after step R070, the method further includes the following steps:
[0103] R080, In response to obtaining a new voice segment to be recognized; where the duration ST’ of the new voice segment to be recognized = ST - ΔST; ΔST is the duration corresponding to m - 1 voice frames.
[0104] R081, Add the last m - 1 voice frames in SA to before the first voice frame in the voice frame list corresponding to the new voice segment to be recognized to obtain the voice frame list corresponding to the new voice segment to be recognized.
[0105] R082, Determine the corresponding human voice segment according to the voice frame list corresponding to the new voice segment to be recognized.
[0106] In this embodiment, in the above steps, the remaining m - 1 speech frames are not judged, and the speech segments to be recognized are all consecutive in sequence; therefore, add the remaining m - 1 speech frames before the first speech frame in the speech frame list corresponding to the new speech segment to be recognized, and then perform the human voice recognition in the above steps.
[0107] In this embodiment, by combining the VAD model and the sliding window statistical method, it is possible to effectively reduce the influence of instantaneous noise or short-term non-human voice interference and improve the accuracy of human voice frame judgment; by using multi-frame confidence statistics instead of single-frame threshold judgment, the probability of misjudgment in a high-noise environment is reduced, which is especially suitable for complex acoustic environments such as aircraft cabins; by dynamically analyzing the confidence distribution of consecutive speech frames, the start and end points of human voice segments can be more accurately located, avoiding problems such as segment truncation or redundancy caused by fixed windows or thresholds in traditional methods; by first screening high-quality human voice frames, the ineffective calculations in subsequent ASR processing can be reduced, improving the operating efficiency of the overall speech recognition system; by adjusting parameters such as the sliding window size and confidence threshold, different cabin environments can be adapted, such as the voice detection requirements of the passenger cabin and the cockpit; the solution of the present invention has high scalability and applicability on the basis of improving the accuracy and stability of human voice detection.
[0108] In addition, although the steps of the methods in the present disclosure are described in a specific order in the drawings, this does not require or imply that these steps must be performed in that specific order, or that all the steps shown must be performed to achieve the desired result. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution, etc.
[0109] An embodiment of the present invention also provides a non-transitory computer-readable storage medium, which can be set in an electronic device to store at least one instruction or at least one segment of a program related to a method for implementing a method in the method embodiment. The at least one instruction or the at least one segment of the program is loaded and executed by the processor to implement the method provided in the above embodiment.
[0110] The program product can adopt any combination of one or more readable media. The readable media can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0111] A computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, carrying readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. The readable signal medium may also be any readable medium other than a readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device.
[0112] The program code contained on the readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0113] The program code for performing the operations of the present application may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's device, executed as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., by using an Internet service provider to connect through the Internet).
[0114] Embodiments of the present invention also provide an electronic device, including a processor and the aforementioned non-transitory computer-readable storage medium.
[0115] The electronic device is merely an example and should not impose any limitation on the functions and scope of use of the embodiments of the present application.
[0116] The electronic device is presented in the form of a general-purpose computing device. The components of the electronic device may include but are not limited to: at least one of the aforementioned processors, at least one of the aforementioned memories, and a bus connecting different system components (including the memory and the processor).
[0117] Wherein, the memory stores program code, and the program code can be executed by the processor, so that the processor executes the steps in the various embodiments described in this specification.
[0118] The memory may include a readable medium in the form of volatile memory, such as random access memory (RAM) and / or cache memory, and may further include read-only memory (ROM).
[0119] The memory may also include program / utility with a set (at least one) of program modules, and such program modules include but are not limited to: an operating system, one or more application programs, other program modules, and program data, and an implementation of a network environment may be included in each or some combination of these examples.
[0120] The bus may represent one or more of several types of bus structures, including a memory bus or a memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any of the multiple bus structures.
[0121] The electronic device may also communicate with one or more external devices (such as a keyboard, a pointing device, a Bluetooth device, etc.), may also communicate with one or more devices that enable a user to interact with the electronic device, and / or may communicate with any device that enables the electronic device to communicate with one or more other computing devices (such as a router, a modem, etc.). Such communication may be carried out through an input / output (I / O) interface. Moreover, the electronic device may also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter. The network adapter communicates with other modules of the electronic device through the bus. It should be understood that although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device, including but not limited to: microcode, device drivers, redundant processors, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0122] Through the description of the above embodiments, those skilled in the art can easily understand that the exemplary embodiments described herein can be implemented by software, or can be implemented by a combination of software and necessary hardware. Therefore, the technical solution according to the embodiments of the present disclosure can be embodied in the form of a software product, and the software product can be stored in a non-volatile storage medium (which may be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which may be a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.
[0123] The embodiment of the present invention also provides a computer program product, which includes program code. When the program product runs on an electronic device, the program code is used to enable the electronic device to execute the steps in the method according to various exemplary embodiments of the present invention described above in this specification.
[0124] Although some specific embodiments of the present invention have been described in detail by way of examples, those skilled in the art should understand that the above examples are for illustrative purposes only and not for limiting the scope of the present invention. Those skilled in the art should also understand that various modifications can be made to the embodiments without departing from the scope and spirit of the present invention.
Claims
1. A method for identifying personnel based on audio fusion, characterized in that The method includes the following steps: R100, if the x-th human voice segment HC to be recognized is not determined x for the corresponding user, then obtain the preset fused voiceprint feature list HB = (HB1, HB2,..., HB e ,..., HB g ), e = 1, 2,..., g; where HB e is the preset e-th fused voiceprint feature, and g is the number of preset fused voiceprint features; HB e is obtained by fusing voice segments of at least two known users; R200, obtain HC x The first similarity with each fused voiceprint feature in HB to obtain HC x The corresponding first similarity list HD x =(HD x,1 , HD x,2 , …, HD x,e , …, HD x,g ); where HD x,e is the first similarity between HC x and HB e . R300, according to HD x , obtain the first target similarity HD x ' = MAX(HD x ); where MAX() is a preset maximum value function; R400, if HD x ’≥H’, then determine HC x The corresponding users are several known users corresponding to H'; where H' is a preset voiceprint feature similarity threshold.
2. The method for identifying a person based on audio fusion according to claim 1, wherein, After step R400, the method further includes the following steps: R500, if HD x ’ < H’, then determine HC x The corresponding users include other users besides the known users.
3. The method for identifying a person based on audio fusion according to claim 1, wherein HB is obtained through the following steps: R110, obtain the voice segments of each known user to obtain a list of known user voice segments HE = (HE1, HE2,..., HE h ,..., HE k ), h = 1, 2,..., k; where HE h is the voice segment of the h-th known user, and k is the number of known users; R120, fuse at least two known user voice segments in HE to obtain a list of fused voice segments HF = (HE1, HE2,..., HE e ,…, HE g ); where HE e is the e-th fused voice segment obtained by fusion; g = 2 k -k-1; R130, extracting the fused voiceprint features for each fused voice segment in HF to obtain HB.
4. The method for identifying a person based on audio fusion according to claim 1, wherein Before step R100, the method further includes the following steps: R010, obtaining a voice segment to be recognized with a preset duration; R020 divides the speech segment to be recognized into a number of consecutive speech frames of the same duration to obtain a speech frame list SA = (SA1, SA2, …, SA i , …, SA n ), where i = 1, 2, …, n; among them, SA i is the i-th speech frame obtained by dividing the speech segment to be recognized, and n is the number of speech frames obtained by dividing the speech segment to be recognized; R030, input SA into a preset VAD model to obtain a list of voice confidence levels SA' = (SA'1, SA'2,..., SA' i ,..., SA' n ); where SA' i is the confidence level of the SA i being the confidence level of the voice frame; R040, performing a first sliding window operation on SA'; wherein, the first sliding window includes m voice confidence levels; R050, each time the first sliding window slides, obtaining the number QN of voice confidence levels greater than a preset first voice confidence level threshold within the first sliding window; R060, if QN / m > η1, determining the speech frame corresponding to the first voice confidence level within the corresponding first sliding window as a voice frame; otherwise, determining it as a non-voice frame; wherein, η1 is a preset first weight; R070, determining the voice segment corresponding to the voice to be recognized according to the voice frames and non-voice frames in SA.
5. The method for identifying a person based on audio fusion according to claim 4, wherein Step R070 includes the following steps: R071, set the voice frames in SA to the first preset character and the non-voice frames to the second preset character to obtain the character list ZA = (ZA1, ZA2,..., ZA j ,..., ZA n+1-m ), j = 1, 2,..., n + 1 - m; ZA j is the preset character corresponding to SA j , n + 1 - m is the number of times the first sliding window slides; the first preset character and the second preset character are different; R072, traverse ZA, and divide several consecutive first preset characters in ZA into a group to obtain a list of first preset character groups LA = (LA1, LA2,..., LA p ,..., LA q ), where p = 1, 2,..., q; among them, LA p is the p-th first preset character group obtained by dividing the first preset characters in ZA, and q is the number of first preset character groups obtained by dividing the first preset characters in ZA; R073, traverse LA. If NU p > NU', then determine the corresponding recognized voice segment of LA p as a human voice segment; where NU p is the number of the first preset characters in LA p ; NU' is a preset first threshold value.
6. The method for identifying a person based on audio fusion according to claim 5, characterized in that After step R073, the method further includes the following steps: R074, inputting the determined voice segment into a preset ASR model to obtain a text list corresponding to each voice segment.
7. The method for identifying a person based on audio fusion according to claim 1, wherein Step R040 includes the following steps: R041, obtaining a first preset value SN = 1; R042, control the first sliding window such that the first voice confidence within the first sliding window is SA' SN ; R043, if SN < n + 1 - m, obtaining SN = SN + 1 and entering S420; otherwise, jumping out of the current process.
8. A non-transitory computer-readable storage medium storing at least one instruction or at least one program segment, characterized in that, The at least one instruction or the at least one program is loaded and executed by a processor to implement the audio fusion-based personnel identity recognition method according to any one of claims 1-7.
9. An electronic device, characterized in that, It includes a processor and the non-transitory computer-readable storage medium according to claim 8.
Citation Information
Patent Citations
Speech data processing method and device, computer device and storage medium
CN108877775A
Speech recognition method
CN109273000A
Voice endpoint detection method and device, electronic equipment and readable storage medium
CN112992191A
Automatic identity recognition method based on voiceprint information of speaker
CN113113022A
Chinese long speech recognition method and device, equipment and storage medium
CN115019780A