Multi-speaker identification method and device, equipment and storage medium
Through the multi-channel microphone array and sound source localization algorithm, the sound source information is accurately obtained, and the speech segment boundaries are optimized by combining the gating and stable window re-detection mechanism, which solves the recognition accuracy and stability problems of multi-speaker recognition in complex environments and realizes efficient speaker switching detection.
Patent Information
- Application Number
- CN202510874842.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-09-16
AI Technical Summary
Multi-speaker recognition technology lacks recognition accuracy and stability in complex acoustic environments, especially in scenarios with reverberation, noise, overlapping multiple sound sources, and dynamic azimuth changes, where the recognition accuracy drops significantly.
A multi-channel microphone array and a preset sound source localization algorithm are used to determine the spatial state sequence of the sound source information. The gating mechanism of the sound source azimuth and activity is combined to segment the speech segments. The boundaries are optimized through a stable window re-detection mechanism. The confidence weight is used to determine the matching similarity of the voiceprint feature vector to terminate the current speech segment and start the new speaker recognition.
The recall rate and accuracy of speaker switching detection are improved, missed detection and false detection are reduced, and high recognition robustness is maintained in complex scenarios.
Smart Images

Figure CN120656451A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of speech processing technology, and in particular to a multi-speaker recognition method, apparatus, device and storage medium. Background Art
[0002] With the widespread adoption of intelligent voice systems in various scenarios, such as in-vehicle, conference rooms, and homes, multi-speaker recognition technology has become an indispensable and critical component of voice interaction systems. However, in practical deployments, systems often face a series of challenges in complex acoustic environments, significantly impacting recognition accuracy and stability. First, reverberation and noise contamination distort the acoustic characteristics of speech signals, severely interfering with the voiceprint model's ability to extract individual speaker characteristics, thereby reducing overall recognition performance. Second, during dynamic interactions, speakers are often in motion, resulting in frequent changes in the sound source's orientation. This significantly interferes with sound source tracking and voiceprint recognition models based on spatial consistency. Furthermore, simultaneous speech from multiple speakers is common, leading to spectral overlap in speech signals. Traditional recognition models based on a single acoustic modality struggle to effectively distinguish between speakers, resulting in increased confusion. Furthermore, real-time voice interaction places stringent demands on the system's end-to-end processing latency, making it difficult for traditional methods to achieve both accurate and real-time performance. Recognition accuracy can still significantly decline in real-world scenarios characterized by high reverberation, overlapping multiple sound sources, and dynamic positional shifts.
[0003] As can be seen from the above, how to improve the robustness and accuracy of multi-speaker recognition in complex environments is an urgent problem that needs to be solved. Summary of the Invention
[0004] In view of this, the present invention aims to provide a multi-speaker recognition method, apparatus, device, and storage medium that can improve the robustness and accuracy of multi-speaker recognition in complex environments. The specific solution is as follows:
[0005] In a first aspect, the present application provides a multi-speaker recognition method, comprising:
[0006] Determine the spatial state sequence corresponding to the current sound source information based on a multi-channel microphone array and a preset sound source localization algorithm; the current sound source information includes the sound source azimuth, sound source number, and sound source activity;
[0007] Segmenting the current sound source into speech segments based on the sound source azimuth and the sound source activity in the spatial state sequence and using a preset gating mechanism, determining each initial speech segment boundary based on the segmentation result, and optimizing each initial speech segment boundary using a preset stable window re-detection mechanism to obtain optimized speech segment boundaries;
[0008] Determining a stability index of each speech segment corresponding to each optimized speech segment boundary using the sound source azimuth and the sound source activity, and determining a confidence weight corresponding to a time frame within each window based on the stability index and using a sliding window technique;
[0009] Voiceprint feature vectors are extracted from the optimized speech paragraphs corresponding to the boundaries of each optimized speech paragraph, and the matching similarity between each voiceprint feature vector is determined using the confidence weight. If the matching similarity meets the preset switching condition, the recognition operation of the current speech paragraph corresponding to the current speaker is terminated, and the recognition operation of the new speech paragraph corresponding to the new speaker is started to obtain a multi-speaker recognition result.
[0010] Optionally, the determining of a spatial state sequence corresponding to the current sound source information based on a multi-channel microphone array and a preset sound source localization algorithm includes:
[0011] The multi-channel audio is collected based on a multi-channel microphone array, and the collected audio signal is subjected to denoising and echo cancellation processing to obtain a processed audio signal;
[0012] Based on the processed audio signal and using a preset sound source localization algorithm, the sound source azimuth, sound source number and sound source activity are determined, and the sound source azimuth, the sound source number and the sound source activity are used to determine the spatial state sequence corresponding to the current sound source information.
[0013] Optionally, determining a sound source azimuth, a sound source number, and a sound source activity based on the processed audio signal and using a preset sound source localization algorithm, and determining a spatial state sequence corresponding to current sound source information using the sound source azimuth, the sound source number, and the sound source activity, includes:
[0014] Analyzing the time difference between the sound source arriving at different microphones based on the processed audio signal and using a preset sound source localization algorithm, and determining the azimuth of the sound source using the time difference;
[0015] Analyzing the energy intensity of the processed audio signal based on the processed audio signal and using a preset sound source localization algorithm to obtain sound source activity;
[0016] Allocating sound source numbers using the sound source azimuth and the sound source activity to obtain numbers for each sound source;
[0017] A spatial state sequence is constructed based on the sound source azimuth, the sound source activity and the sound source number in each time frame and in chronological order.
[0018] Optionally, segmenting the current sound source into speech segments based on the sound source azimuth and the sound source activity in the spatial state sequence and using a preset gating mechanism to determine each initial speech segment boundary based on the segmentation result includes:
[0019] If, in a first preset number of time frames preceding the current sound source, the sound source activity in the spatial state sequence is greater than a preset activity threshold, and the variance corresponding to the sound source azimuth is not greater than a starting gating threshold, then it is indicated that the current sound source corresponding to the first preset number of time frames is continuously emitting sound;
[0020] In a second preset number of time frames subsequent to the current sound source, if the sound source activity in the spatial state sequence is less than a preset silence threshold, and the variance corresponding to the sound source azimuth is greater than a termination gating threshold, then it is indicated that the current sound source corresponding to the second preset number of time frames is silent, and speech segments are segmented for the current sound source based on the second preset number of time frames to determine boundaries of each initial speech segment;
[0021] The starting gating threshold is smaller than the ending gating threshold.
[0022] Optionally, the optimizing each of the initial speech paragraph boundaries by using a preset stable window re-detection mechanism to obtain an optimized speech paragraph boundary includes:
[0023] Determining whether the sound source activity within a preset time period corresponding to each of the speech segment boundaries meets a preset sound source recovery condition;
[0024] If so, determining whether the sound source number in the preset time period has changed;
[0025] If the sound source number in the preset time period does not change, the mis-segmented segment is determined based on the preset time period, and the initial speech paragraph boundary is optimized using the mis-segmented segment to obtain an optimized speech paragraph boundary.
[0026] Optionally, the determining of a stability index of each speech segment corresponding to each optimized speech segment boundary by using the sound source azimuth and the sound source activity, and determining a confidence weight corresponding to a time frame within each window by using a sliding window technique based on the stability index, includes:
[0027] Determining the stability index of each speech segment corresponding to each optimized speech segment boundary by using the degree of change of the sound source azimuth and the sound source activity;
[0028] Based on the sound source azimuth, a sliding window technique is used to determine the variance value of the azimuth sequence corresponding to the sound source azimuth in each window, and the confidence weight corresponding to the time frame in each window is determined using the variance value and a preset variance experience value.
[0029] Optionally, extracting voiceprint feature vectors from the optimized speech paragraphs corresponding to the boundaries of each optimized speech paragraph, determining matching similarity between the voiceprint feature vectors using the confidence weight, and terminating the recognition operation for the current speech paragraph corresponding to the current speaker and starting the recognition operation for the new speech paragraph corresponding to the new speaker if the matching similarity meets a preset switching condition, so as to obtain a multi-speaker recognition result, including:
[0030] Extracting voiceprint feature vectors based on the optimized speech paragraphs corresponding to the boundaries of each optimized speech paragraph using a voiceprint recognition algorithm, and determining matching similarities between the voiceprint feature vectors using the confidence weights;
[0031] If the matching similarity satisfies the preset change condition and the sound source azimuth satisfies the preset jump condition, the recognition operation of the current speech segment corresponding to the current speaker is terminated, and the recognition operation of the new speech segment corresponding to the new speaker is started to obtain a multi-speaker recognition result.
[0032] In a second aspect, the present application provides a multi-speaker recognition device, comprising:
[0033] A sequence determination module is configured to determine a spatial state sequence corresponding to current sound source information based on a multi-channel microphone array and a preset sound source localization algorithm; the current sound source information includes a sound source azimuth, a sound source number, and a sound source activity;
[0034] a boundary optimization module, configured to segment the current sound source into speech segments based on the sound source azimuth and the sound source activity in the spatial state sequence and using a preset gating mechanism, determine each initial speech segment boundary based on the segmentation result, and optimize each initial speech segment boundary using a preset stable window re-detection mechanism to obtain an optimized speech segment boundary;
[0035] a weight determination module, configured to determine a stability index of each speech segment corresponding to each of the optimized speech segment boundaries using the sound source azimuth and the sound source activity, and determine a confidence weight corresponding to a time frame within each window based on the stability index and using a sliding window technique;
[0036] A similarity determination module is used to extract voiceprint feature vectors from the optimized speech paragraphs corresponding to the boundaries of each optimized speech paragraph, and use the confidence weight to determine the matching similarity between each voiceprint feature vector. If the matching similarity meets the preset switching condition, the recognition operation of the current speech paragraph corresponding to the current speaker is terminated, and the recognition operation of the new speech paragraph corresponding to the new speaker is started to obtain a multi-speaker recognition result.
[0037] In a third aspect, the present application provides an electronic device, comprising:
[0038] Memory, used to store computer programs;
[0039] The processor is configured to execute the computer program to implement the aforementioned multi-speaker recognition method.
[0040] In a fourth aspect, the present application provides a computer-readable storage medium for storing a computer program, wherein the computer program implements the aforementioned multi-speaker recognition method when executed by a processor.
[0041] The present application determines the spatial state sequence corresponding to the current sound source information based on a multi-channel microphone array and a preset sound source localization algorithm; the current sound source information includes the sound source azimuth, the sound source number and the sound source activity; based on the sound source azimuth and the sound source activity in the spatial state sequence and using a preset gating mechanism, the current sound source is segmented into speech segments, so as to determine the boundaries of each initial speech segment based on the segmentation results, and the boundaries of each initial speech segment are optimized using a preset stable window re-detection mechanism to obtain the optimized speech segment boundaries; the sound source azimuth and the sound source activity are determined The stability index of each speech paragraph corresponding to each of the optimized speech paragraph boundaries is determined based on the stability index and using the sliding window technology to determine the confidence weight corresponding to the time frame within each window; the voiceprint feature vector is extracted from the optimized speech paragraph corresponding to each of the optimized speech paragraph boundaries, and the matching similarity between each of the voiceprint feature vectors is determined using the confidence weight. If the matching similarity meets the preset switching condition, the recognition operation of the current speech paragraph corresponding to the current speaker is terminated, and the recognition operation of the new speech paragraph corresponding to the new speaker is started to obtain a multi-speaker recognition result.
[0042] As can be seen from the above, this application uses a multi-channel microphone array and a preset sound source localization algorithm to accurately obtain the azimuth, number, and activity of the sound source to form a spatial state sequence. The preset gating mechanism based on the sound source azimuth and sound source activity can effectively distinguish the start and end of the speech segment, avoiding the misjudgment of noise or silence segments as speech. The preset stable window re-detection mechanism is then used to further optimize the boundary and reduce missegmentation caused by short-term interference. The stability of the time frame is then quantified based on the confidence weight within each window. The voiceprint feature vector extracted from the optimized speech segment corresponding to each optimized speech segment boundary is combined with the confidence weight to determine the matching similarity. In this way, if the matching similarity meets the preset switching condition, the current speech segment can be terminated in time and the new speaker recognition can be started, which significantly improves the recall rate and accuracy of speaker switch detection, reduces missed detections and false detections, and maintains high recognition robustness even in complex scenarios (such as multi-person conversations and background noise). BRIEF DESCRIPTION OF THE DRAWINGS
[0043] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.
[0044] Figure 1 This is a flow chart of a multi-speaker recognition method disclosed in this application;
[0045] Figure 2 This is a schematic diagram of the structure of a multi-speaker recognition device disclosed in this application;
[0046] Figure 3 This is a structural diagram of an electronic device disclosed in this application. DETAILED DESCRIPTION
[0047] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0048] Currently, multi-speaker recognition technology often faces a series of challenges in complex acoustic environments during practical deployment. First, reverberation and noise contamination can distort the acoustic characteristics of speech signals, severely interfering with the voiceprint model's ability to extract individual speaker features, thereby reducing the overall recognition performance of the system. Second, during dynamic interactions, speakers are often in motion, resulting in frequent changes in the sound source's orientation, which significantly interferes with sound source tracking and voiceprint recognition models based on spatial consistency. Furthermore, simultaneous speech is common, and the speech signals overlap in the spectrum. Traditional recognition models based on a single acoustic modality struggle to effectively distinguish between speakers, leading to increased confusion. In real-world scenarios with high reverberation, overlapping multiple sound sources, and dynamic positional changes, recognition accuracy can still significantly decrease. Therefore, this application provides a multi-speaker recognition method that promptly terminates the current speech segment and initiates new speaker recognition when the matching similarity meets a preset switching condition. This significantly improves the recall and accuracy of speaker switch detection, reduces missed and false detections, and maintains high recognition robustness even in complex scenarios (such as multi-person conversations and background noise).
[0049] See also Figure 1 As shown, an embodiment of the present invention discloses a multi-speaker recognition method, comprising:
[0050] Step S11: determining a spatial state sequence corresponding to current sound source information based on a multi-channel microphone array and a preset sound source localization algorithm; the current sound source information includes a sound source azimuth, a sound source number, and a sound source activity.
[0051] In this embodiment, a microphone array is deployed near the current sound source to obtain a multi-channel microphone array; the multi-channel microphone array can be a circular, linear or distributed array; multi-channel audio is collected based on the multi-channel microphone array, and the collected audio signal is denoised based on spectral subtraction, Wiener filtering or a deep learning model to obtain denoised audio, and then an adaptive filter or a deep echo cancellation network is used to remove the speaker echo in the denoised audio to obtain a processed audio signal, and the sound source azimuth, sound source number and sound source activity are determined based on the processed audio signal and using a preset sound source localization algorithm, and the sound source azimuth, the sound source number and the sound source activity are used to determine the spatial state sequence corresponding to the current sound source information. Specifically, the method of determining the spatial state sequence corresponding to the current sound source information based on the multi-channel microphone array and the preset sound source localization algorithm includes: collecting multi-channel audio based on the multi-channel microphone array, and performing denoising and echo cancellation processing on the collected audio signal to obtain a processed audio signal; determining the sound source azimuth, sound source number and sound source activity based on the processed audio signal and using the preset sound source localization algorithm, and using the sound source azimuth, the sound source number and the sound source activity to determine the spatial state sequence corresponding to the current sound source information.
[0052] It can be understood that after obtaining the processed audio signal, the coordinate position of the sound source in space is determined to obtain the three-dimensional coordinates of the current sound source in each time frame, and the time difference of the sound source reaching different microphones is analyzed based on the processed audio signal and using a preset sound source localization algorithm, and the sound source azimuth is determined using the time difference; the sound source azimuth is the direction angle of the sound source relative to the multi-channel microphone array; based on the processed audio signal and using a preset sound source localization algorithm, the energy intensity of the processed audio signal is analyzed to obtain the sound source activity; the sound source activity is a numerical value that measures the strength of the processed audio signal; for example, when the speaker is speaking, the sound source activity is high, and when the speaker is not speaking, the corresponding sound source activity is low; then a number is assigned to each speaker to distinguish different sound sources to obtain the number of each sound source, and then a spatial state sequence is constructed based on the sound source azimuth, the sound source activity and the sound source number in each time frame and in chronological order. In a specific embodiment, the spatial state sequence is as follows:
[0053] ;
[0054] in, is the spatial state sequence in each time frame; numbering the sound sources; is the azimuth of the sound source; is the activity of the sound source.
[0055] Specifically, the method of determining the sound source azimuth, sound source number and sound source activity based on the processed audio signal and using a preset sound source localization algorithm, and determining the spatial state sequence corresponding to the current sound source information using the sound source azimuth, the sound source number and the sound source activity, includes: analyzing the time difference between the sound source reaching different microphones based on the processed audio signal and using a preset sound source localization algorithm, and determining the sound source azimuth using the time difference; analyzing the energy intensity of the processed audio signal based on the processed audio signal and using a preset sound source localization algorithm to obtain the sound source activity; allocating sound source numbers using the sound source azimuth and the sound source activity to obtain the numbers of each sound source; and constructing a spatial state sequence based on the sound source azimuth, the sound source activity and the sound source number in each time frame and in chronological order.
[0056] Step S12: Segment the current sound source into speech segments based on the sound source azimuth and the sound source activity in the spatial state sequence and using a preset gating mechanism, determine the boundaries of each initial speech segment based on the segmentation results, and optimize each of the initial speech segment boundaries using a preset stable window re-detection mechanism to obtain optimized speech segment boundaries.
[0057] In this embodiment, after obtaining the spatial state sequence, the current sound source is segmented into speech segments based on the sound source azimuth and the sound source activity in the spatial state sequence and using a preset gating mechanism to obtain corresponding segmentation results; the preset gating mechanism is a rule for judging the start and end of speech. In the frame, the activity of the sound source in the spatial state sequence is greater than the preset activity threshold, and the variance corresponding to the sound source azimuth is not greater than the starting gate threshold, that is, , it indicates that the corresponding current sound source is continuously speaking, that is, it is judged that this is the beginning of speech activity, and the starting speech segment boundary of this speech segment is marked; among them, is the azimuth of the sound source; is the variance corresponding to the azimuth of the sound source; is the starting gate threshold. In the frame, the activity of the sound source in the spatial state sequence is less than the preset quiet threshold, and the variance corresponding to the sound source azimuth is greater than the termination gate threshold, that is, It indicates that the corresponding current sound source has not made any sound, that is, it is judged that the speech activity ends at this time, and the end speech segment boundary of this speech segment is marked; wherein, is the termination gating threshold.
[0058] Specifically, the method of segmenting the current sound source into speech segments based on the sound source azimuth and the sound source activity in the spatial state sequence and using a preset gating mechanism to determine the boundaries of each initial speech segment based on the segmentation results includes: in a first preset number of time frames before the current sound source, if the sound source activity in the spatial state sequence is greater than a preset activity threshold, and the variance corresponding to the sound source azimuth is not greater than a starting gating threshold, then it indicates that the current sound source corresponding to the first preset number of time frames is continuously making a sound; in a second preset number of time frames after the current sound source, if the sound source activity in the spatial state sequence is less than a preset quiet threshold, and the variance corresponding to the sound source azimuth is greater than a termination gating threshold, then it indicates that the current sound source corresponding to the second preset number of time frames is not making a sound, and segmenting the current sound source into speech segments based on the second preset number of time frames to determine the boundaries of each initial speech segment; wherein the starting gating threshold is less than the termination gating threshold. It is worth mentioning that the preset activity threshold, the preset quiet threshold, the start gating threshold and the end gating threshold can be adjusted according to actual conditions and are not specifically limited here.
[0059] It is understandable that, considering that the speaker may have a short pause during the speech, or the microphone may have a slight jitter, which may lead to the misjudgment of the end of the speech segment, it is necessary to preset a stable window re-detection mechanism to optimize the initial speech segment boundaries. If the sound source activity suddenly drops and then quickly recovers within the preset time period, that is, the sound source activity drops to a preset low-frequency threshold within the preset time period and quickly recovers to normal activity; the preset low-frequency threshold can be determined according to the actual situation. Then, it is determined whether the sound source number of the preset time period has changed. If the sound source number of the preset time period has not changed, it means that the speaker is still the same person. Based on the preset time period, the mis-segmented segment is determined, and the mis-segmented segment is merged with the corresponding initial speech segment boundary to obtain the optimized speech segment boundary. The preset stable window re-detection mechanism significantly reduces the pseudo-boundary detection rate caused by the speaker's natural pause or microphone jitter.
[0060] Specifically, the preset stable window re-detection mechanism is used to optimize each of the initial speech paragraph boundaries to obtain the optimized speech paragraph boundaries, including: determining whether the sound source activity within the preset time period corresponding to each of the speech paragraph boundaries meets the preset sound source recovery condition; if so, determining whether the sound source number of the preset time period has changed; if the sound source number of the preset time period has not changed, determining the mis-segmented segment based on the preset time period, and using the mis-segmented segment to optimize the initial speech paragraph boundary to obtain the optimized speech paragraph boundary.
[0061] Step S13: using the sound source azimuth and the sound source activity to determine the stability index of each speech segment corresponding to each optimized speech segment boundary, and based on the stability index and using the sliding window technology to determine the confidence weight corresponding to the time frame in each window.
[0062] In this embodiment, after obtaining the optimized speech paragraph boundaries, the stability index of each speech paragraph corresponding to each optimized speech paragraph boundary is determined using the change degree of the sound source azimuth and the sound source activity. The formula corresponding to the stability index is as follows:
[0063] ;
[0064] in, is the stability index; is the degree of change of the azimuth angle of the sound source, for example, the azimuth angle of the sound source changes by 5° within 10 milliseconds. 0.5° / millisecond; is the degree of change in the activity of the sound source, for example, the activity of the sound source drops from 0.8 to 0.2 within 10 milliseconds. is 0.06; is a balance coefficient, for example, it is set to 1, so that the influence of the sound source azimuth and the sound source activity on the stability index is the same. The lower the stability index is, the smaller the change of the sound source azimuth and the sound source activity is, which means the speaking state is more stable; the higher the stability index is, the greater the change of the sound source azimuth and the sound source activity is, which means the speaking state is more unstable. After obtaining the stability index, set the sliding window length to W, for example, 200ms corresponds to 25 frames, the sampling interval is 8ms, and the azimuth sequence corresponding to the sound source azimuth in each frame t is obtained. , determine the variance value corresponding to the azimuth sequence , and use the variance value and the preset variance experience value to determine the confidence weight corresponding to the time frame in each window. The formula corresponding to the confidence weight is as follows:
[0065] ;
[0066] in, is the confidence weight; is the variance value corresponding to the azimuth sequence; The preset variance empirical value is a fixed benchmark value frequently preset based on history; is the sign of the natural exponential function. When the variance is small, the confidence weight is close to 1, indicating that the frame has high spatial confidence. When the variance increases significantly (i.e., the sound source direction fluctuates dramatically or a speaker switch occurs), the confidence weight decreases rapidly.
[0067] Specifically, the method of using the sound source azimuth angle and the sound source activity to determine the stability index of each speech paragraph corresponding to each speech paragraph boundary after optimization, and using the sliding window technology to determine the confidence weight corresponding to the time frame in each window based on the stability index includes: using the degree of change of the sound source azimuth angle and the sound source activity to determine the stability index of each speech paragraph corresponding to each speech paragraph boundary after optimization; based on the sound source azimuth angle and using the sliding window technology to determine the variance value of the azimuth angle sequence corresponding to the sound source azimuth angle in each window, and using the variance value and the preset variance experience value to determine the confidence weight corresponding to the time frame in each window. It is worth mentioning that the preset variance experience value can be adjusted according to actual conditions and is not specifically limited here.
[0068] Step S14: extracting voiceprint feature vectors from the optimized speech paragraphs corresponding to the boundaries of each optimized speech paragraph, and determining the matching similarity between each voiceprint feature vector using the confidence weight; if the matching similarity meets the preset switching condition, terminating the recognition operation for the current speech paragraph corresponding to the current speaker, and starting the recognition operation for the new speech paragraph corresponding to the new speaker, so as to obtain a multi-speaker recognition result.
[0069] In this embodiment, a voiceprint recognition algorithm is determined based on a deep learning model. Then, based on the optimized speech segments corresponding to the boundaries of each optimized speech segment, a voiceprint feature vector is extracted using the voiceprint recognition algorithm. The matching similarity between each voiceprint feature vector is determined using the confidence weight. The formula corresponding to the matching similarity is as follows:
[0070] ;
[0071] in, is the matching similarity; is the voiceprint feature vector; and is the confidence weight corresponding to the voiceprint feature vector; is the cosine similarity corresponding to the voiceprint feature vector. After obtaining the matching similarity, if the matching similarity suddenly decreases and the sound source azimuth suddenly decreases and then increases again, the recognition operation for the current speech segment corresponding to the current speaker is terminated, and the recognition operation for the new speech segment corresponding to the new speaker is initiated to obtain the multi-speaker recognition result. In one specific embodiment, the multi-speaker recognition result obtained is speech segment 1 (0-10s): speaker A; speech segment 2 (10-18s): speaker B; speech segment 3 (18-30s): speaker A.
[0072] Specifically, the method includes extracting voiceprint feature vectors from the optimized speech segments corresponding to the boundaries of each optimized speech segment, determining the matching similarity between the voiceprint feature vectors using the confidence weight, and if the matching similarity satisfies a preset switching condition, terminating the recognition operation for the current speech segment corresponding to the current speaker and initiating the recognition operation for the new speech segment corresponding to the new speaker to obtain a multi-speaker recognition result. The method includes: extracting voiceprint feature vectors based on the optimized speech segments corresponding to the boundaries of each optimized speech segment using a voiceprint recognition algorithm, determining the matching similarity between the voiceprint feature vectors using the confidence weight; if the matching similarity satisfies a preset change condition and the sound source azimuth satisfies a preset jump condition, terminating the recognition operation for the current speech segment corresponding to the current speaker and initiating the recognition operation for the new speech segment corresponding to the new speaker to obtain a multi-speaker recognition result. It is worth mentioning that the preset change condition and the preset jump condition can be adjusted according to actual conditions and are not specifically limited here.
[0073] As can be seen from the above, this application uses a multi-channel microphone array and a preset sound source localization algorithm to accurately obtain the azimuth, number, and activity of the sound source to form a spatial state sequence. The preset gating mechanism based on the sound source azimuth and sound source activity can effectively distinguish the start and end of the speech segment, avoiding the misjudgment of noise or silence segments as speech. The preset stable window re-detection mechanism is then used to further optimize the boundary and reduce missegmentation caused by short-term interference. The stability of the time frame is then quantified based on the confidence weight within each window. The voiceprint feature vector extracted from the optimized speech segment corresponding to each optimized speech segment boundary is combined with the confidence weight to determine the matching similarity. In this way, if the matching similarity meets the preset switching condition, the current speech segment can be terminated in time and the new speaker recognition can be started, which significantly improves the recall rate and accuracy of speaker switch detection, reduces missed detections and false detections, and maintains high recognition robustness even in complex scenarios (such as multi-person conversations and background noise).
[0074] Accordingly, see Figure 2 As shown, the present application also provides a multi-speaker recognition device, comprising:
[0075] A sequence determination module 11 is configured to determine a spatial state sequence corresponding to current sound source information based on a multi-channel microphone array and a preset sound source localization algorithm; the current sound source information includes a sound source azimuth, a sound source number, and a sound source activity;
[0076] a boundary optimization module 12, configured to segment the current sound source into speech segments based on the sound source azimuth and the sound source activity in the spatial state sequence and using a preset gating mechanism, determine each initial speech segment boundary based on the segmentation result, and optimize each initial speech segment boundary using a preset stable window re-detection mechanism to obtain an optimized speech segment boundary;
[0077] A weight determination module 13 is configured to determine a stability index of each speech segment corresponding to each of the optimized speech segment boundaries using the sound source azimuth and the sound source activity, and determine a confidence weight corresponding to a time frame within each window based on the stability index and using a sliding window technique;
[0078] The similarity determination module 14 is used to extract voiceprint feature vectors from the optimized speech paragraphs corresponding to the boundaries of each optimized speech paragraph, and use the confidence weight to determine the matching similarity between each voiceprint feature vector. If the matching similarity meets the preset switching condition, the recognition operation of the current speech paragraph corresponding to the current speaker is terminated, and the recognition operation of the new speech paragraph corresponding to the new speaker is started to obtain a multi-speaker recognition result.
[0079] As can be seen from the above, this application uses a multi-channel microphone array and a preset sound source localization algorithm to accurately obtain the azimuth, number, and activity of the sound source to form a spatial state sequence. The preset gating mechanism based on the sound source azimuth and sound source activity can effectively distinguish the start and end of the speech segment, avoiding the misjudgment of noise or silence segments as speech. The preset stable window re-detection mechanism is then used to further optimize the boundary and reduce missegmentation caused by short-term interference. The stability of the time frame is then quantified based on the confidence weight within each window. The voiceprint feature vector extracted from the optimized speech segment corresponding to each optimized speech segment boundary is combined with the confidence weight to determine the matching similarity. In this way, if the matching similarity meets the preset switching condition, the current speech segment can be terminated in time and the new speaker recognition can be started, which significantly improves the recall rate and accuracy of speaker switch detection, reduces missed detections and false detections, and maintains high recognition robustness even in complex scenarios (such as multi-person conversations and background noise).
[0080] In some specific embodiments, the sequence determination module 11 may specifically include:
[0081] An audio signal processing unit, configured to collect multi-channel audio based on a multi-channel microphone array, and perform denoising and echo cancellation on the collected audio signal to obtain a processed audio signal;
[0082] A state sequence determination unit is used to determine the sound source azimuth, sound source number and sound source activity based on the processed audio signal and using a preset sound source localization algorithm, and to determine the spatial state sequence corresponding to the current sound source information using the sound source azimuth, the sound source number and the sound source activity.
[0083] In some specific embodiments, the sequence determination module 11 may specifically include:
[0084] an azimuth angle determination unit, configured to analyze the time difference between the sound source reaching different microphones based on the processed audio signal and using a preset sound source localization algorithm, and determine the azimuth angle of the sound source using the time difference;
[0085] an activity determination unit, configured to analyze the energy intensity of the processed audio signal based on the processed audio signal and using a preset sound source localization algorithm to obtain a sound source activity;
[0086] a number determination unit, configured to assign sound source numbers using the sound source azimuth and the sound source activity to obtain numbers for each sound source;
[0087] A sequence construction unit is used to construct a spatial state sequence in time order based on the sound source azimuth, the sound source activity and the sound source number in each time frame.
[0088] In some specific implementations, the boundary optimization module 12 may specifically include:
[0089] an activity comparison unit, configured to indicate that the current sound source corresponding to the first preset number of time frames is continuously emitting sound if, in a first preset number of time frames preceding the current sound source, the activity of the sound source in the spatial state sequence is greater than a preset activity threshold, and the variance corresponding to the sound source azimuth is not greater than a starting gating threshold;
[0090] The speech segmentation unit is configured to indicate that, in the second preset number of time frames after the current sound source, if the activity of the sound source in the spatial state sequence is less than a preset quiet threshold and the variance corresponding to the sound source azimuth is greater than a termination gating threshold, then the current sound source corresponding to the second preset number of time frames has not made any sound, and to segment the current sound source into speech segments based on the second preset number of time frames to determine the boundaries of each initial speech segment.
[0091] In some specific implementations, the boundary optimization module 12 may specifically include:
[0092] an activity determination unit, configured to determine whether the sound source activity within a preset time period corresponding to each of the speech segment boundaries meets a preset sound source recovery condition;
[0093] a sound source number determination unit, configured to determine whether the sound source number in the preset time period has changed if the condition is met;
[0094] The paragraph boundary optimization unit is used to determine the mis-segmented segment based on the preset time period if the sound source number in the preset time period does not change, and optimize the initial speech paragraph boundary using the mis-segmented segment to obtain an optimized speech paragraph boundary.
[0095] In some specific implementations, the weight determination module 13 may specifically include:
[0096] An index determination unit, configured to determine a stability index of each speech segment corresponding to each optimized speech segment boundary by using the sound source azimuth angle and the degree of change of the sound source activity;
[0097] A confidence value weight determination unit is used to determine the variance value of the azimuth angle sequence corresponding to the sound source azimuth angle in each window based on the sound source azimuth angle and using a sliding window technology, and to determine the confidence weight corresponding to the time frame in each window using the variance value and a preset variance experience value.
[0098] In some specific implementations, the similarity determination module 14 may specifically include:
[0099] a matching similarity determination unit, configured to extract voiceprint feature vectors based on the optimized speech paragraphs corresponding to the boundaries of the optimized speech paragraphs using a voiceprint recognition algorithm, and determine matching similarities between the voiceprint feature vectors using the confidence weights;
[0100] The recognition result determination unit is used to terminate the recognition operation of the current speech segment corresponding to the current speaker if the matching similarity meets the preset change condition and the sound source azimuth meets the preset jump condition, and start the recognition operation of the new speech segment corresponding to the new speaker to obtain a multi-speaker recognition result.
[0101] Furthermore, the embodiment of the present application also discloses an electronic device, Figure 3 This is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content of this diagram should not be construed as limiting the scope of application of this application. The electronic device 20 may include at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 is used to store a computer program, which is loaded and executed by the processor 21 to implement the relevant steps of the multi-speaker recognition method disclosed in any of the aforementioned embodiments. Furthermore, the electronic device 20 in this embodiment may be a computer.
[0102] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and the external device. The communication protocol it follows is any communication protocol that can be applied to the technical solution of this application and is not specifically limited here; the input and output interface 25 is used to obtain external input data or output data to the outside world. Its specific interface type can be selected according to specific application needs and is not specifically limited here.
[0103] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, disk or CD, etc. The resources stored thereon can include an operating system 221, a computer program 222, etc., and the storage method can be temporary storage or permanent storage.
[0104] The operating system 221 is used to manage and control the hardware devices and computer program 222 on the electronic device 20. The operating system 221 can be Windows Server, NetWare, Unix, Linux, etc. In addition to including a computer program capable of implementing the multi-speaker recognition method performed by the electronic device 20 as disclosed in any of the aforementioned embodiments, the computer program 222 can further include computer programs capable of performing other specific tasks.
[0105] Furthermore, this application also discloses a computer-readable storage medium for storing a computer program; wherein, when executed by a processor, the computer program implements the aforementioned multi-speaker identification method. The specific steps of this method can be found in the corresponding contents disclosed in the aforementioned embodiments and will not be repeated here.
[0106] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from the other embodiments. Reference can be made to the descriptions of the identical or similar parts between the various embodiments. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple, and the relevant parts can be referred to the descriptions of the methods.
[0107] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0108] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0109] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.
[0110] The above is a detailed introduction to the technical solution provided by the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. At the same time, for those skilled in the art, according to the ideas of the present application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.
Claims
1. A multi-speaker recognition method, characterized in that: include: Determine the spatial state sequence corresponding to the current sound source information based on a multi-channel microphone array and a preset sound source localization algorithm; the current sound source information includes the sound source azimuth, sound source number, and sound source activity; Segmenting the current sound source into speech segments based on the sound source azimuth and the sound source activity in the spatial state sequence and using a preset gating mechanism, determining each initial speech segment boundary based on the segmentation result, and optimizing each initial speech segment boundary using a preset stable window re-detection mechanism to obtain optimized speech segment boundaries; Determining a stability index of each speech segment corresponding to each optimized speech segment boundary using the sound source azimuth and the sound source activity, and determining a confidence weight corresponding to a time frame within each window based on the stability index and using a sliding window technique; Voiceprint feature vectors are extracted from the optimized speech paragraphs corresponding to the boundaries of each optimized speech paragraph, and the matching similarity between each voiceprint feature vector is determined using the confidence weight. If the matching similarity meets the preset switching condition, the recognition operation of the current speech paragraph corresponding to the current speaker is terminated, and the recognition operation of the new speech paragraph corresponding to the new speaker is started to obtain a multi-speaker recognition result.
2. The multi-speaker recognition method according to claim 1, wherein: The determining of the spatial state sequence corresponding to the current sound source information based on the multi-channel microphone array and the preset sound source localization algorithm includes: The multi-channel audio is collected based on a multi-channel microphone array, and the collected audio signal is subjected to denoising and echo cancellation processing to obtain a processed audio signal; Based on the processed audio signal and using a preset sound source localization algorithm, the sound source azimuth, sound source number and sound source activity are determined, and the sound source azimuth, the sound source number and the sound source activity are used to determine the spatial state sequence corresponding to the current sound source information.
3. The multi-speaker recognition method according to claim 2, characterized in that: The method of determining a sound source azimuth, a sound source number, and a sound source activity based on the processed audio signal and using a preset sound source localization algorithm, and determining a spatial state sequence corresponding to current sound source information using the sound source azimuth, the sound source number, and the sound source activity, includes: Analyzing the time difference between the sound source arriving at different microphones based on the processed audio signal and using a preset sound source localization algorithm, and determining the azimuth of the sound source using the time difference; Analyzing the energy intensity of the processed audio signal based on the processed audio signal and using a preset sound source localization algorithm to obtain sound source activity; Allocating sound source numbers using the sound source azimuth and the sound source activity to obtain numbers for each sound source; A spatial state sequence is constructed based on the sound source azimuth, the sound source activity and the sound source number in each time frame and in chronological order.
4. The multi-speaker recognition method according to claim 3, wherein: The method of segmenting the current sound source into speech segments based on the sound source azimuth and the sound source activity in the spatial state sequence and using a preset gating mechanism to determine each initial speech segment boundary based on the segmentation result includes: If, in a first preset number of time frames preceding the current sound source, the sound source activity in the spatial state sequence is greater than a preset activity threshold, and the variance corresponding to the sound source azimuth is not greater than a starting gating threshold, then it is indicated that the current sound source corresponding to the first preset number of time frames is continuously emitting sound; In a second preset number of time frames subsequent to the current sound source, if the sound source activity in the spatial state sequence is less than a preset silence threshold, and the variance corresponding to the sound source azimuth is greater than a termination gating threshold, then it is indicated that the current sound source corresponding to the second preset number of time frames is silent, and speech segments are segmented for the current sound source based on the second preset number of time frames to determine boundaries of each initial speech segment; The starting gating threshold is smaller than the ending gating threshold.
5. The multi-speaker recognition method according to claim 1, wherein: The optimizing the boundaries of each of the initial speech paragraphs by using a preset stable window re-detection mechanism to obtain optimized speech paragraph boundaries includes: Determining whether the sound source activity within a preset time period corresponding to each of the speech segment boundaries meets a preset sound source recovery condition; If so, determining whether the sound source number in the preset time period has changed; If the sound source number in the preset time period does not change, the mis-segmented segment is determined based on the preset time period, and the initial speech paragraph boundary is optimized using the mis-segmented segment to obtain an optimized speech paragraph boundary.
6. The multi-speaker recognition method according to claim 1, wherein: The method of determining the stability index of each speech segment corresponding to each optimized speech segment boundary by using the sound source azimuth and the sound source activity, and determining the confidence weight corresponding to the time frame within each window by using a sliding window technique based on the stability index, includes: Determining the stability index of each speech segment corresponding to each optimized speech segment boundary by using the degree of change of the sound source azimuth and the sound source activity; Based on the sound source azimuth, a sliding window technique is used to determine the variance value of the azimuth sequence corresponding to the sound source azimuth in each window, and the confidence weight corresponding to the time frame in each window is determined using the variance value and a preset variance experience value.
7. The multi-speaker recognition method according to any one of claims 1 to 6, characterized in that: The extracting of voiceprint feature vectors from the optimized speech paragraphs corresponding to the boundaries of the optimized speech paragraphs, determining the matching similarity between the voiceprint feature vectors using the confidence weights, and terminating the recognition operation for the current speech paragraph corresponding to the current speaker and initiating the recognition operation for the new speech paragraph corresponding to the new speaker if the matching similarity meets a preset switching condition, so as to obtain a multi-speaker recognition result, includes: Extracting voiceprint feature vectors based on the optimized speech paragraphs corresponding to the boundaries of each optimized speech paragraph using a voiceprint recognition algorithm, and determining matching similarities between the voiceprint feature vectors using the confidence weights; If the matching similarity satisfies the preset change condition and the sound source azimuth satisfies the preset jump condition, the recognition operation of the current speech segment corresponding to the current speaker is terminated, and the recognition operation of the new speech segment corresponding to the new speaker is started to obtain a multi-speaker recognition result.
8. A multi-speaker recognition device, characterized in that: include: A sequence determination module is configured to determine a spatial state sequence corresponding to current sound source information based on a multi-channel microphone array and a preset sound source localization algorithm; the current sound source information includes a sound source azimuth, a sound source number, and a sound source activity; a boundary optimization module, configured to segment the current sound source into speech segments based on the sound source azimuth and the sound source activity in the spatial state sequence and using a preset gating mechanism, determine each initial speech segment boundary based on the segmentation result, and optimize each initial speech segment boundary using a preset stable window re-detection mechanism to obtain an optimized speech segment boundary; a weight determination module, configured to determine a stability index of each speech segment corresponding to each of the optimized speech segment boundaries using the sound source azimuth and the sound source activity, and determine a confidence weight corresponding to a time frame within each window based on the stability index and using a sliding window technique; A similarity determination module is used to extract voiceprint feature vectors from the optimized speech paragraphs corresponding to the boundaries of each optimized speech paragraph, and use the confidence weight to determine the matching similarity between each voiceprint feature vector. If the matching similarity meets the preset switching condition, the recognition operation of the current speech paragraph corresponding to the current speaker is terminated, and the recognition operation of the new speech paragraph corresponding to the new speaker is started to obtain a multi-speaker recognition result.
9. An electronic device, characterized in that: include: Memory, used to store computer programs; A processor, configured to execute the computer program to implement the multi-speaker recognition method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that Used to store a computer program, wherein when the computer program is executed by a processor, the multi-speaker recognition method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Audio signal processing method and device and electronic equipment
CN114387970A
Multi-channel array multi-speaker voice separation method, electronic equipment and medium
CN117275506A
Multi-person sound source separation method, device, equipment, medium and computer program product
CN119741938A
Meeting summary automatic generation method based on multi-source heterogeneous information fusion
CN120045701A
Information processing device and program
JP2015031729A