Audio speaker recognition method and system, storage medium and electronic equipment
Through the method of combining multi-dimensional acoustic feature analysis and deep learning model with large language model, the limitations of traditional audio speaker recognition methods in dealing with complex audio scenes are solved, the recognition accuracy and automation are improved, and various complex audio scenes are adapted to.
Patent Information
- Application Number
- CN202510697903.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-06-27
Smart Images

Figure CN120220729A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of audio processing and artificial intelligence, specifically covering speech recognition, speaker separation, audio feature analysis, and natural language processing technologies. In particular, it relates to an audio speaker recognition method, system, storage medium, and electronic device. Background Art
[0002] Traditional audio speaker recognition and content extraction methods usually use a single model for processing, and there are the following problems: 1. The processing methods for mono-channel and stereo audio lack flexibility. Often, it is necessary to manually judge the audio type and adopt different processing procedures.
[0003] 2. It is unable to effectively handle the case of pseudo-stereo (where the content of the two channels is mostly the same but there are slight differences), which easily leads to repeated recognition or information loss.
[0004] 3. Noise in the audio will significantly affect the recognition accuracy, and traditional methods have limited resistance to noise.
[0005] 4. In scenarios where speakers interact frequently, traditional methods are prone to adjacent sentence adhesion or speaker recognition errors.
[0006] 5. Traditional methods lack the ability to post-process the recognition results and cannot accurately classify speakers according to semantic content.
[0007] The existing technologies mainly rely on simple channel separation and single-model recognition, and are unable to effectively handle complex audio scenarios, resulting in low recognition accuracy and requiring a large amount of manual intervention for result correction. Summary of the Invention
[0008] The purpose of the present invention is to overcome the technical problems existing in the existing audio speaker recognition, and provides an audio speaker recognition method, system, storage medium, and electronic device. Through multi-dimensional acoustic feature analysis, the cooperation of deep learning models and large language models, the recognition accuracy and automation degree are improved, and it can adapt to various complex audio scenarios.
[0009] The purpose of the present invention is achieved through the following technical solutions: In the first aspect, an audio speaker recognition method is provided, including the following steps: S1. Preprocess the input audio, extract the channel information of the input audio, and judge whether the audio type is mono-channel or stereo; S2. For stereo audio, judge whether the audio type is pseudo-stereo or true stereo by comparing at least two acoustic feature parameters of the left and right channels; S3. Select a processing strategy according to the audio type: For monophonic audio, directly execute step S4; for pseudo-stereo audio, select the optimal channel and then execute step S4; for true stereo audio, separate the left and right channels and execute step S4 respectively, and merge the processing results of the left and right channels through a time alignment algorithm; S4. Audio content recognition: Perform noise reduction preprocessing, voice activity detection, speaker segmentation, content recognition, and punctuation restoration on the selected channel in sequence; S5. Use a large language model to perform speaker role marking on the results of audio content recognition; S6. Merge the content segments of the same speaker role to generate a structured output.
[0010] In some embodiments, the determination of the audio type as pseudo-stereo or true stereo by comparing at least two acoustic feature parameters of the left and right channels includes: Determine the audio type by combining the energy difference and cosine similarity of the left and right channels: Wherein, represents the determined audio type, represents monophonic, represents pseudo-stereo, represents true stereo; represents the number of channels of the audio; represents the energy difference; represents the cosine similarity; and are the thresholds of the energy difference and cosine similarity respectively; represents other situations.
[0011] In some embodiments, the selection of the optimal channel for the pseudo-stereo audio includes: Select the channel with higher quality through signal-to-noise ratio estimation.
[0012] In some embodiments, the use of a large language model to perform speaker role marking on the results of audio content recognition includes: According to the preset candidate role set, map the recognized speakers to the actual roles. Among them, when the number of recognized speakers is 1 and the number of candidate roles is greater than 1, use the large language model to perform secondary speaker separation according to the content; when the number of recognized speakers is greater than the number of candidate roles, use the large language model to merge and map multiple speakers to the candidate role set; Perform semantic analysis on the content through the large language model to determine whether there is a situation where the content of multiple speakers is merged.
[0013] In some embodiments, the merging of the content segments of the same speaker role to generate a structured output includes: The merging conditions are as follows: Among them, indicates whether to merge segments and is the decision result; and respectively represent the speaker roles of segments and ; and respectively represent the end and start timestamps of segments and ; represents the time threshold; Generate a structured output including speaker role, content, and timestamp.
[0014] In some embodiments, the processing results of the left and right channels are merged through a time alignment algorithm, including: Arrange the speaker segments of the left and right channels in chronological order to generate a unified timeline; Perform conflict processing on segments that overlap in time but have the same content, and select which channel's recognition result to retain based on signal strength and content integrity; When there is partial overlap in time and different content between the left and right channels, the system marks them as speaking simultaneously and retains the content of both channels.
[0015] In a second aspect, an audio speaker recognition system is provided, including: An audio processing and feature extraction module for preprocessing the input audio and extracting the channel information of the input audio; An audio type judgment and processing module for judging whether the audio type is mono or stereo. For stereo audio, by comparing at least two acoustic feature parameters of the left and right channels, judge whether the audio type is pseudo-stereo or true stereo; it is also used to select a processing strategy according to the audio type: directly perform speaker recognition and content extraction for mono audio; select the optimal channel for pseudo-stereo and then perform speaker recognition and content extraction; separate the left and right channels for true stereo and perform speaker recognition and content extraction respectively, and merge the processing results of the left and right channels through a time alignment algorithm; A speaker recognition and content extraction module for sequentially performing noise reduction preprocessing, voice activity detection, speaker segmentation, content recognition, and punctuation restoration on the selected channel; A role marking module for using a large language model to mark the speaker role of the result of audio content recognition; A result optimization and output module for merging the content segments of the same speaker role to generate a structured output.
[0016] In a third aspect, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the audio speaker recognition method described in the first aspect is implemented.
[0017] In a fourth aspect, an electronic device is provided, including a memory and a processor. A computer instruction that can run on the processor is stored on the memory. When the processor runs the computer instruction, the audio speaker recognition method described in the first aspect is executed.
[0018] It should be further noted that the corresponding technical features of the above embodiments can be combined or replaced with each other without conflict to form a new technical solution.
[0019] Compared with the prior art, the beneficial effects of the present invention are as follows: Through the calculation and analysis of multiple audio feature dimensions, the present invention automatically determines the audio type, adopts corresponding processing strategies according to different audio types, combines with a large language model for content understanding and speaker role marking, so as to improve the accuracy and automation of audio content recognition, reduce the cost of manual intervention, and can adapt to various complex audio scenarios. The specific advantages include: 1. High degree of automation: Automatically determines the audio type through multi-dimensional feature analysis without manual intervention.
[0020] 2. Strong adaptability: Can process various audio types such as mono, stereo, and pseudo-stereo.
[0021] 3. High recognition accuracy: Significantly improves the recognition accuracy through the cooperation of multiple models and the assistance of a large language model. 4. Good anti-noise performance: Adopts an advanced noise reduction algorithm to improve the recognition performance in a noisy environment.
[0022] 5. High processing efficiency: The optimized processing flow reduces the consumption of computing resources and improves the processing efficiency.
[0023] 6. Strong scalability: The modular design makes the system easy to expand and optimize. Description of the Drawings
[0024] Figure 1 It is a flowchart of an audio speaker recognition method shown in an embodiment of the present invention; Figure 2 It is the composition and flowchart of an audio speaker recognition system shown in an embodiment of the present invention; Figure 3 It is the implementation process of a multi-dimensional acoustic feature analysis module shown in an embodiment of the present invention; Figure 4 It is the implementation process of an audio type judgment and processing module shown in an embodiment of the present invention; Figure 5The implementation process of the speaker recognition and content extraction module shown in the embodiments of the present invention; Figure 6 The implementation process of the role marking module shown in the embodiments of the present invention; Figure 7 The implementation process of the result optimization and output module shown in the embodiments of the present invention. Detailed implementation manners
[0025] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of them. The components of the embodiments of the present application described and illustrated herein can be arranged and designed in various different configurations. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0026] It should be noted that the defects existing in the above prior art solutions are all the results obtained by the inventors after practice and careful research. Therefore, the discovery process of the above problems and the solutions proposed by the embodiments of the present application below for the above problems should be the contributions made by the inventors to the present application during the invention creation process, and should not be understood as the technical content known to those skilled in the art.
[0027] The present invention realizes accurate judgment of audio types through multi-dimensional acoustic feature analysis, improves audio quality and segmentation accuracy by combining advanced deep learning noise reduction models and speaker segmentation models, and uses large language models for content understanding and speaker role marking, effectively solving the limitations of traditional audio speaker recognition methods in dealing with complex audio scenarios, especially greatly improving the processing ability for pseudo-stereo audio and scenarios with frequent speaker interactions.
[0028] In view of the technical problems pointed out in the background art, the embodiments provided by the present invention are as follows: Refer to Figure 1 , in an exemplary embodiment, an audio speaker recognition method includes the following steps: S1. Preprocess the input audio, extract the channel information of the input audio, and determine whether the audio type is mono or stereo; S2. For stereo audio, judge whether the audio type is pseudo-stereo or true stereo by comparing at least two acoustic feature parameters of the left and right channels; S3. Select a processing strategy according to the audio type: directly execute step S4 for mono audio; select the optimal channel for pseudo-stereo and then execute step S4; separate the left and right channels for true stereo and execute step S4 respectively, and merge the processing results of the left and right channels through a time alignment algorithm; S4. Audio content recognition: successively perform noise reduction preprocessing, voice activity detection, speaker segmentation, content recognition, and punctuation restoration on the selected sound channels; S5. Use a large language model to perform speaker role marking on the results of audio content recognition; S6. Merge the content segments of the same speaker role to generate a structured output.
[0029] Specifically, in step S1, preprocess the input audio, including: Unify the sampling rate of the audio to facilitate subsequent processing; perform volume normalization on the audio to keep the average volume of the audio within an appropriate range.
[0030] For stereo audio, by calculating and analyzing multiple acoustic feature dimensions, achieve accurate judgment of the audio type. The common calculation methods for 10 feature dimensions are as follows: 1. Energy difference calculation: Calculate the energy difference between the left and right channels to determine whether the audio is true stereo or pseudo-stereo. The energy difference calculation formula is as follows: Among them, represents the energy difference ratio between the left and right channels; represents the energy of the left channel, which is the sum of the squares of all sampling points of the left channel ; represents the energy of the right channel, which is the sum of the squares of all sampling points of the right channel ; is the total number of audio sampling points.
[0031] 2. Bitwise similarity analysis: Calculate the bitwise similarity degree of the sampling points of the left and right channels. The bitwise similarity calculation formula is as follows: Among them, represents the bitwise similarity; is an indicator function, which takes the value of 1 when the condition is satisfied, otherwise it is 0; and are the amplitudes of the left and right channels at the th sampling point respectively; is the similarity threshold, indicating the maximum allowable difference for determining that two sampling points are similar; is the total number of audio sampling points.
[0032] 3. Cosine similarity: Calculate the vector similarity of the left and right channel signals. The cosine similarity calculation formula is as follows: Among them, represents the cosine similarity; and are the amplitudes of the left and right channels at the -th sampling point respectively; the numerator is the dot product sum of the left and right channel signals; the denominator is the product of the norms of the left and right channel signals; is the total number of audio sampling points. The value range of the cosine similarity is , and the closer the value is to 1, the more similar the two channels are.
[0033] 4. Difference method: Calculate the sample differences between the left and right channels and analyze the distribution characteristics of the differences. The difference method calculation formula is as follows: Normalized difference value: where represents the average absolute difference between the left and right channels; represents the normalized difference value; and are the amplitudes of the left and right channels at the -th sampling point respectively; and are the maximum absolute amplitudes of the left and right channels respectively; is the total number of audio sampling points.
[0034] 5. Correlation coefficient method: Calculate the degree of correlation between the left and right channels. The correlation coefficient calculation formula is as follows: where represents the Pearson correlation coefficient between the left and right channels; is the average value of the left channel; is the average value of the right channel; and are the amplitudes of the left and right channels at the -th sampling point respectively; is the total number of audio sampling points. The value range of the correlation coefficient is , and the closer the value is to 1, the stronger the positive correlation.
[0035] 6. Spectrum analysis method: Compare the spectral feature differences between the left and right channels. The spectrum similarity calculation formula is as follows: where represents the spectrum similarity; and are the complex spectra of the left and right channels at the -th frequency point respectively; and are the corresponding spectral amplitudes respectively; is the number of frequency points of the spectrum. The calculation of spectrum similarity is essentially the calculation of cosine similarity of frequency-domain signals.
[0036] 7. Short-Time Fourier Transform (STFT) method: Analyze the differences between the left and right channels in the time-frequency domain. The STFT calculation formula is as follows: where, represents the short-time Fourier transform coefficient at time frame and frequency index ; represents the time-domain signal; is the window function (such as Hamming window or Hanning window); is the number of Fourier transform points; is the time-frame index; is the frequency index; is the frame shift, that is, the difference in the number of samples between adjacent time frames.
[0037] Time-frequency domain similarity calculation: where, represents the time-frequency domain similarity; and are the STFT coefficients of the left and right channels at time frame and frequency index respectively; and are the corresponding amplitude spectra respectively; is the total number of time frames; is the total number of frequency points.
[0038] 8. Cross-correlation method: Judge the similarity by calculating the cross-correlation function of the left and right channels. The cross-correlation calculation formula is as follows: Normalized cross-correlation function: where, represents the cross-correlation function of the left and right channels; represents the normalized cross-correlation function; is the time-shift parameter, indicating the time offset of the right channel relative to the left channel; and are the amplitudes of the left and right channels at the th sampling point respectively; is the value of the autocorrelation function of the left channel at zero delay; is the value of the autocorrelation function of the right channel at zero delay; is the total number of audio sampling points.
[0039] 9. Entropy difference method: Calculate the information difference degree between the left and right channels through information entropy. The information entropy calculation formula is as follows: The entropy difference calculation formula is: where, represents the information entropy of the signal ; represents the probability distribution of the signal amplitude , usually estimated through a histogram; and are the information entropies of the left and right channels respectively; represents the entropy difference between the left and right channels. The smaller the value, the more similar the information content of the two channels.
[0040] 10. Independent Component Analysis (ICA) method: Determine the audio file type by separating the independent components of the left and right channels. The ICA model can be expressed as: where, represents the observed signal matrix, and each row represents the time-domain signal of one channel; represents the source signal matrix, and each row represents an independent sound source signal; represents the mixing matrix, which describes how the source signals are mixed to form the observed signals.
[0041] Estimate the inverse matrix of by maximizing non-Gaussianity or minimizing mutual information: : where, represents the estimated source signal matrix; represents the separation matrix, which is the estimated inverse matrix of the mixing matrix . The similarity of the separated source signals can be used to judge the audio type. If the separated independent components are highly similar, it indicates that the original channel signals may come from the same sound source.
[0042] It is found in the actual calculation process that the calculation time consumption of the cross-correlation method and the Independent Component Analysis (ICA) method is significantly higher than that of other methods. Therefore, these two methods are excluded in the present invention. Considering the threshold discrimination, the energy difference is finally selected as the main measurement index, and combined with other features for comprehensive judgment. Preferably, in some examples, judging the audio type as pseudo-stereo or true stereo by comparing at least two acoustic feature parameters of the left and right channels includes: Judging the audio type by combining the energy difference and cosine similarity of the left and right channels: Among them, represents the judged audio type, represents mono, represents pseudo-stereo, represents true stereo; represents the number of audio channels; represents the energy difference; represents the cosine similarity; and are respectively the thresholds of the energy difference and the cosine similarity, determined according to experiments , ; represents other situations.
[0043] Preferably, the pseudo-stereo selects the optimal channel, including: Selecting a channel with higher quality through signal-to-noise ratio estimation to avoid repeated recognition: Among them, represents the signal-to-noise ratio; represents the variance of the signal, reflecting the signal energy; represents the variance of the noise, reflecting the noise energy. The higher the signal-to-noise ratio, the better the signal quality.
[0044] Preferably, in step S4, the noise reduction preprocessing uses the Denoiser deep learning model, which adopts the CRNU-Net structure and consists of a causal model based on convolution and LSTM. The frame size is 40 ms and the stride is 16 ms. Specifically, the dns48 pre-trained denoising model is used, and the number of parameters is 18,867,937 (about 72 MB). Experiments show that the Denoiser method has a better noise reduction effect compared with traditional spectral subtraction, Wiener filtering, and other deep learning models (such as FRCRN). Through experimental verification, for an audio file of about 20 minutes, the Denoiser noise reduction processing takes about 10 seconds, and the processing efficiency is relatively high.
[0045] The voice activity detection uses the FSMN-Monophone VAD model for voice endpoint detection. This model is proposed by the speech team of DAMO Academy and is used to detect the start and end time point information of valid speech in the input audio. The FSMN model structure considers context information during modeling, has fast training and inference speeds, and controllable latency. At the modeling unit level, the single speech class is upgraded to Monophone, improving the model's abstract learning ability and discrimination ability.
[0046] The speaker segmentation uses the CAM++ model. CAM++ is a speaker recognition model based on a densely connected time-delay neural network. Compared with mainstream speaker recognition models such as ResNet34 and ECAPA-TDNN, it has higher accuracy and faster inference speed. The model structure includes a residual convolutional network as the front end and a time-delay neural network structure as the backbone. The front end extracts local and fine time-frequency features, and the backbone uses dense connections to reuse hierarchical features. At the same time, a lightweight context-related masking module is embedded to remove irrelevant noise in the features and retain key speaker information.
[0047] The content recognition and punctuation restoration use the CT-Punc (Controllable Time-delay Transformer) model for punctuation restoration. CT-Punc is a punctuation module in the efficient post-processing framework proposed by the Speech Team of Alibaba DAMO Academy, designed specifically for Chinese general punctuation. The model structure consists of three parts: Embedding, Encoder, and Predictor. Through an innovative controllable time-delay design, while ensuring the model performance, it effectively controls the delay of punctuation and avoids the problem of constantly changing and refreshing punctuation.
[0048] It should be noted that the specific model in step S4 is not construed as a limitation of this application, but only as a preferred implementable solution. In other examples, it can be replaced with other similar models according to the actual effect.
[0049] Preferably, in step S5, a large language model is used to perform speaker role marking on the results of audio content recognition, including: Mapping the recognized speakers to actual roles according to a preset candidate role set; Handling abnormal situations of role recognition conflicts. The most common scenario is that the number of recognized speakers is inconsistent with the number of preset roles. For example, the system only recognizes one speaker, but there are multiple preset roles; or the system recognizes multiple speakers, but the number of preset roles is less than the number of recognized speakers. This mismatch causes the system to be unable to directly and clearly assign a reasonable role to each speaking segment. In some cases, during a conversation, the model may recognize several different sentences of the same speaker as 2 - 3 speakers (usually due to some changes in tone color during the process, short sentences, etc.). Specifically, the following two abnormal situations are handled: 1. When the number of recognized speakers is 1 and the number of candidate roles is greater than 1, use the large language model to perform secondary speaker separation according to the content; 2. When the number of recognized speakers is greater than the number of candidate roles, use the large language model to merge and map multiple speakers to the candidate role set; The existing character recognition model distinguishes speakers solely based on voice information (such as timbre and other features). Introducing a large model in this application is equivalent to performing a secondary verification from the perspective of text semantics.
[0050] Adjacent sentence processing: Detect and process the problem of adjacent sentences being glued together. Use a large language model to perform semantic analysis on the content to determine whether there is a situation where the content of multiple speakers is merged.
[0051] Preferably, the optimization process for the recognition and marking results in step S6 includes: Content merging: Merge the content segments of the same speaker role. The calculation conditions for merging are as follows: Where, represents the decision result of whether to merge the segments and ; and respectively represent the speaker roles of segments and ; and respectively represent the end and start timestamps of segments and ; represents the time threshold; Time alignment: Reorganize the conversation content in chronological order to generate a coherent conversation flow.
[0052] Result formatting: Generate a structured output containing speaker roles, content, and timestamps.
[0053] Furthermore, the processing results of the left and right channels are merged through a time alignment algorithm, including: Arrange the speaker segments of the left and right channels in chronological order to generate a unified timeline; Perform conflict processing on segments that overlap in time but have the same content, and select which channel's recognition result to retain based on signal strength and content integrity; When there is partial overlap in time and different content between the left and right channels, the system marks it as speaking simultaneously and retains the content of both channels.
[0054] Through the calculation of multiple feature dimensions, the present invention accurately determines the audio type and selects the optimal processing strategy; automatically adjusts the processing flow according to the audio type to maximize the utilization of effective information in the audio. Through the organic combination of noise reduction preprocessing, speaker segmentation, content recognition, and role marking, high-quality audio speaker recognition and role marking are achieved. At the same time, the large language model is used to solve problems such as adjacent sentences and role recognition conflicts that are difficult to handle by traditional methods.
[0055] In another exemplary embodiment, referring to Figure 2 , based on the same inventive concept as the method embodiment, an audio speaker recognition system is provided, including: An audio processing and feature extraction module, configured to preprocess the input audio and extract the channel information of the input audio; An audio type judgment and processing module, configured to judge whether the audio type is mono or stereo. For stereo audio, by comparing at least two acoustic feature parameters of the left and right channels, judge whether the audio type is pseudo-stereo or true stereo; and also configured to select a processing strategy according to the audio type: directly perform speaker recognition and content extraction for mono audio; select the optimal channel for pseudo-stereo and then perform speaker recognition and content extraction; separate the left and right channels for true stereo and perform speaker recognition and content extraction respectively, and merge the processing results of the left and right channels through a time alignment algorithm; A speaker recognition and content extraction module, configured to perform noise reduction preprocessing, voice activity detection, speaker segmentation, content recognition, and punctuation restoration on the selected channel in sequence; A role marking module, configured to use a large language model to perform speaker role marking on the result of audio content recognition; A result optimization and output module, configured to merge the content segments of the same speaker role and generate a structured output.
[0056] Among them, the implementation process of the audio type judgment and processing module is as shown in Figure 4 , the implementation process of the speaker recognition and content extraction module is as shown in Figure 5 , the implementation process of the role marking module is as shown in Figure 6 , and the implementation process of the result optimization and output module is as shown in Figure 7 . It should be noted that each module in the system is implemented in the same way as the method steps to implement the corresponding functions, which will not be elaborated here.
[0057] Furthermore, the system further includes a multi-dimensional acoustic feature analysis module, which realizes accurate judgment of the audio type by calculating and analyzing multiple acoustic feature dimensions, and its implementation process is as shown in Figure 3 .
[0058] Based on the method embodiment, the present invention respectively gives the specific processing processes for mono, pseudo-stereo, and true stereo audio.
[0059] The speaker recognition and role marking for mono audio specifically include: 1. Data preparation: Select a mono audio file (such as a certain interview program), with a sampling rate of 16 kHz, a duration of about 30 minutes, and containing the conversation content between the host and the guest.
[0060] 2. Audio preprocessing: 1. Sampling rate check: Confirm that the audio sampling rate is 16 kHz. If not, resample the audio.
[0061] 2. Volume normalization: Normalize the volume of the audio so that the average volume of the audio remains within an appropriate range.
[0062] 3. Audio type judgment: The system detects that the audio is of mono type and directly enters the mono processing flow.
[0063] 3. Noise reduction processing: 1. Use the Denoiser model (dns48 pre-trained model) to perform noise reduction on the audio.
[0064] 2. Calculate the signal-to-noise ratio before and after noise reduction. The signal-to-noise ratio of the original audio is about 15.8 dB, and it is increased to 21.3 dB after noise reduction.
[0065] 3. The time taken for noise reduction processing is about 1 / 60 of the audio duration. For a 30-minute audio, the processing time is about 30 seconds. 4. Voice activity detection: 1. Use the FSMN-Monophone VAD model to perform voice endpoint detection.
[0066] 2. The system detects a total of 87 valid voice segments, with a total duration of about 26 minutes, accounting for 86.7% of the original audio.
[0067] 3. Each voice segment contains start time and end time information, such as: [0.5s - 15.2s], [16.8s - 42.3s], etc.
[0068] 5. Speaker diarization: 1. Use the CAM++ model to perform speaker diarization on each voice segment.
[0069] 2. The system identifies two main speakers (denoted as Speaker 1 and Speaker 2).
[0070] 3. The speaker diarization results show that the appearance frequency of Speaker 1 is about 45% and that of Speaker 2 is about 55%.
[0071] 4. After speaker diarization, a total of 124 speaker segments are obtained, and each segment contains speaker ID, start time, and end time information.
[0072] 6. Content recognition and punctuation restoration: 1. Use the Paraformer model to perform speech recognition on each speaker segment to obtain the text content.
[0073] 2. Use the CT-Punc model to restore punctuation for the recognized text and improve text readability.
[0074] 3. Example of recognition results: - Speaker 1 [0.5s - 15.2s]: "Welcome to today's interview program. I'm the host, Xiao Li. Today we have invited the famous entrepreneur, Mr. Wang, to share his entrepreneurial experience with us." - Speaker 2 [16.8s - 42.3s]: "Thank you, host. I'm very glad to be on this program. My entrepreneurial experience has not been smooth. I've experienced many setbacks to achieve today's results." 7. Role marking: 1. Based on the preset candidate role set ["host", "guest"], use the large language model to mark the roles of the recognition results.
[0075] 2. The large language model analyzes the text content and identifies Speaker 1 as the "host" and Speaker 2 as the "guest".
[0076] 3. Example of role marking: - Host [0.5s - 15.2s]: "Welcome to today's interview program. I'm the host, Xiao Li. Today we have invited the famous entrepreneur, Mr. Wang, to share his entrepreneurial experience with us." - Guest [16.8s - 42.3s]: "Thank you, host. I'm very glad to be on this program. My entrepreneurial experience has not been smooth. I've experienced many setbacks to achieve today's results." 8. Result optimization: 1. Optimize the marked results and merge adjacent content of the same role.
[0077] 2. Adjacent segments with a time interval less than 0.8 seconds are merged into a complete segment.
[0078] 3. Finally, output the structured dialogue content, including the role, timestamp, and text content.
[0079] 9. Evaluation results: 1. The accuracy rate of role marking reaches 98.2%, and only a small number of segments need manual correction.
[0080] 2. The processing time of the entire processing flow for 30 minutes of audio is about 2 minutes, and the processing efficiency is relatively high.
[0081] 3. The finally output dialogue content has a clear structure and can be directly used for subsequent text analysis or content presentation.
[0082] Speaker recognition and role marking for pseudo-stereo audio specifically include: 1. Data preparation: Select a pseudo-stereo audio file (such as a movie dialogue), with a sampling rate of 44.1 kHz, a duration of about 45 minutes, and containing dialogue content of multiple characters.
[0083] 2. Audio preprocessing: 1. Sampling rate unification: Resample the audio to 16 kHz for subsequent processing.
[0084] 2. Volume normalization: Perform volume normalization on the audio to make the average volume of the left and right channels consistent.
[0085] 3. Multidimensional feature calculation and audio type judgment: 1. Energy difference calculation: The energy difference ratio between the left and right channels .
[0086] 2. Cosine similarity calculation: The cosine similarity between the left and right channels .
[0087] 3. Correlation coefficient calculation: The correlation coefficient between the left and right channels .
[0088] 4. Spectral similarity calculation: The spectral similarity between the left and right channels = 0.968.
[0089] 5. According to the judgment algorithm, the system determines that this audio is pseudo-stereo ( and ).
[0090] 4. Optimal channel selection: 1. Calculate the signal-to-noise ratio of the left and right channels: Left channel SNR = 19.2 dB, right channel SNR = 18.6 dB.
[0091] 2. The system selects the left channel with a higher signal-to-noise ratio for subsequent processing.
[0092] 3. By selecting the optimal channel, duplicate processing of the same content is avoided, improving the processing efficiency.
[0093] 5. Noise reduction processing: 1. Use the Denoiser model (dns48 pre-trained model) to perform noise reduction processing on the selected left channel.
[0094] 2. Calculate the signal-to-noise ratio before and after noise reduction. After noise reduction, the signal-to-noise ratio is increased to 24.5 dB, an increase of 5.3 dB.
[0095] 3. The noise reduction process for a 45-minute audio takes approximately 40 seconds.
[0096] 6. Voice Activity Detection: 1. Use the FSMN-Monophone VAD model to detect the endpoints of speech.
[0097] 2. The system detected a total of 156 valid speech segments, with a total duration of approximately 38 minutes, accounting for 84.4% of the original audio.
[0098] 3. Effectively filtered out the background music and ambient sound effects in the movie.
[0099] 7. Speaker Diarization: 1. Use the CAM++ model to perform speaker diarization on each speech segment.
[0100] 2. The system identified 4 main speakers (denoted as Speaker 1, 2, 3, 4).
[0101] 3. After speaker diarization, a total of 210 speaker segments were obtained.
[0102] 8. Content Recognition and Punctuation Restoration: 1. Use the Paraformer model to perform speech recognition on each speaker segment to obtain the text content.
[0103] 2. Use the CT-Punc model to restore punctuation to the recognized text.
[0104] 3. Example of recognition results: - Speaker 1 [125.3s - 132.6s]: "Do you know the secrets of this city? It hides many unknown histories." - Speaker 2 [134.0s - 142.5s]: "I only know that there was a big fire here that almost burned down the whole city." 9. Character Tagging: 1. According to the preset candidate character set ["Protagonist", "Supporting Role A", "Supporting Role B", "Supporting Role C"], use a large language model to perform character tagging on the recognition results.
[0105] 2. The large language model analyzes the text content and maps Speakers 1 - 4 to specific characters respectively.
[0106] 3. Character mapping results: Speaker 1 → "Protagonist", Speaker 2 → "Supporting Role A", Speaker 3 → "Supporting Role B", Speaker 4 → "Supporting Role C".
[0107] 4. Example of character tagging: - Protagonist [125.3s - 132.6s]: "Do you know the secret of this city? It hides a lot of unknown history." - Supporting Role A [134.0s - 142.5s]: "I only know that there was a big fire here before, which almost burned down the whole city." 10. Result Optimization: 1. Handling the problem of sticky sentences: The system detected 3 cases where the content of multiple speakers might be merged, and used a large language model for semantic analysis and separation.
[0108] 2. Merging adjacent content of the same character: Adjacent segments with a time interval less than 1.0 second were merged.
[0109] 3. Finally, output the structured dialogue content, including the character, timestamp, and text content.
[0110] 11. Evaluation Results: 1. The character marking accuracy reached 96.5%.
[0111] 2. The processing time of the entire processing flow for 45 minutes of pseudo - stereo audio was about 2.5 minutes.
[0112] 3. By identifying the pseudo - stereo features and selecting the optimal channel for processing, the computing resource usage was reduced by about 45%, and the efficiency was greatly improved compared with the method of directly processing stereo audio.
[0113] Speaker recognition and character marking for true stereo audio specifically include: 1. Data preparation: Select a true stereo audio file (such as the stereo recording of an interview program), with a sampling rate of 48 kHz, a duration of about 60 minutes, and different speakers' voices recorded in the left and right channels respectively.
[0114] 2. Audio pre - processing: 1. Sampling rate unification: Resample the audio to 16 kHz for subsequent processing.
[0115] 2. Volume normalization: Perform volume normalization on the left and right channels of the audio respectively.
[0116] 3. Multidimensional feature calculation and audio type judgment: 1. Energy difference calculation: The energy difference ratio between the left and right channels 。
[0117] 2. Cosine similarity calculation: The cosine similarity between the left and right channels 。
[0118] 3. Correlation coefficient calculation: The correlation coefficient between the left and right channels 。
[0119] 4. Spectrum similarity calculation: The spectrum similarity between the left and right channels 。
[0120] 5. Differential method calculation: Normalized difference value 。
[0121] 6. According to the judgment algorithm, the system determines that this audio is a true stereo ( or ).
[0122] 4. Left and right channel separation processing: 1. The system extracts the left and right channels respectively as two independent mono audio streams for processing.
[0123] 2. The left channel mainly contains the host's voice, and the right channel mainly contains the guest's voice, but there is a certain degree of voice overlap.
[0124] 5. Noise reduction processing: 1. Use the Denoiser model to perform noise reduction processing on the left and right channels respectively.
[0125] 2. Left channel: SNR before noise reduction = 17.8dB, SNR after noise reduction = 22.9dB.
[0126] 3. Right channel: SNR before noise reduction = 16.5dB, SNR after noise reduction = 21.7dB.
[0127] 4. The total time consumed for noise reduction processing of a 60-minute stereo audio is about 100 seconds.
[0128] 6. Voice activity detection: 1. Use the FSMN-Monophone VAD model to perform voice endpoint detection on the left and right channels respectively.
[0129] 2. 98 effective voice segments are detected in the left channel, with a total duration of about 28 minutes.
[0130] 3. 112 effective voice segments are detected in the right channel, with a total duration of about 32 minutes.
[0131] 4. The system records the channel information, start time and end time of each voice segment.
[0132] 7. Speaker segmentation: 1. Use the CAM++ model to perform speaker segmentation on each voice segment of the left and right channels respectively.
[0133] 2. One speaker (denoted as speaker L) is mainly identified in the left channel.
[0134] 3. The right channel mainly identifies 2 speakers (denoted as Speaker R1 and Speaker R2).
[0135] 4. After speaker segmentation, a total of about 250 speaker segments are obtained.
[0136] 8. Content recognition and punctuation restoration: 1. Use the Paraformer model to perform speech recognition on each speaker segment to obtain the text content.
[0137] 2. Use the CT-Punc model to restore punctuation to the recognized text.
[0138] 3. Example of recognition results: - Speaker L [left channel, 215.6s - 225.8s]: "Could you share what the biggest challenge your company faced in the past year was?" - Speaker R1 [right channel, 226.3s - 248.9s]: "The biggest challenge was the supply chain disruption caused by the pandemic. We had to re-plan the entire production process and find alternative suppliers." 9. Time alignment and merging: 1. The system arranges the speaker segments of the left and right channels in chronological order to generate a unified timeline.
[0139] 2. Handle conflicts for overlapping segments in time, and select which channel's recognition result to retain based on signal strength and content integrity.
[0140] 3. Handle cross-talking situations: When there is partial overlap in time and different content between the left and right channels, the system marks it as simultaneous speech and retains the content of both channels.
[0141] 4. After merging, a total of about 220 non-overlapping speaker segments in time are obtained.
[0142] 10. Role marking: 1. According to the preset candidate role set ["Host", "Guest A", "Guest B"], use a large language model to perform role marking on the recognition results.
[0143] 2. The large language model analyzes the text content and maps Speaker L to "Host", Speaker R1 to "Guest A", and Speaker R2 to "Guest B".
[0144] 3. Example of role marking: - Host [215.6s - 225.8s]: "Could you share what the biggest challenge your company faced in the past year was?" - Guest A [226.3s - 248.9s]: "The biggest challenge is the supply chain disruption caused by the pandemic. We had to re-plan the entire production process and find alternative suppliers." - Guest B [250.1s - 265.4s]: "To add, besides supply chain issues, talent mobility is also an important challenge we face. The remote work mode has changed many work habits." 11. Abnormal situation handling: 1. Handling voice aliasing problems: In certain time periods (such as [312.5s - 318.2s]), both the left and right channels contain the voices of multiple speakers simultaneously. The system analyzes the content through the large language model to distinguish the roles.
[0145] 2. Handling role recognition conflicts: The system detected 3 conflicts in the role recognition results. By analyzing the context content, the large language model is used for the final determination.
[0146] 3. Handling sticky sentence problems: The system detected 5 situations where the content of multiple speakers might be merged. The large language model is used for semantic analysis and separation.
[0147] 12. Result optimization: 1. Merging adjacent content of the same role: Adjacent segments with a time interval less than 1.2 seconds are merged into a complete segment.
[0148] 2. Reorganizing the dialogue content in chronological order to generate a coherent dialogue flow.
[0149] 3. Finally, outputting structured dialogue content, including the role, timestamp, and text content.
[0150] 13. Evaluation of results: 1. The accuracy rate of role marking reached 94.8%, which is higher than the processing effects of mono-channel and pseudo-stereo audio.
[0151] 2. The processing time of the entire processing flow for 60 minutes of true stereo audio is approximately 5 minutes, and the processing efficiency is relatively high.
[0152] 3. By processing the left and right channels separately, the system can more accurately separate the content of different speakers, especially performing excellently in scenarios where multiple people are speaking simultaneously.
[0153] 4. The main advantage of true stereo audio processing is that it can better utilize spatial information to distinguish different speakers, reducing the difficulty of speaker segmentation and improving the recognition accuracy.
[0154] Through the comparison of the above three types of audio, the advantages of the present invention in processing different types of audio can be seen: 1. Comparison of processing efficiency: - Mono audio: The processing time is approximately 1 / 15 of the audio duration.
[0155] - Pseudo-stereo audio: By identifying pseudo-stereo features and selecting the optimal channel for processing, the processing time is approximately 1 / 18 of the audio duration, and the computing resource usage is reduced by approximately 45%.
[0156] - True stereo audio: It is necessary to process the left and right channels separately. The processing time is approximately 1 / 12 of the audio duration, but a higher speaker separation accuracy can be obtained.
[0157] 2. Comparison of character marking accuracy: - Mono audio: The character marking accuracy is approximately 98.2%, which is suitable for scenarios with a small number of speakers and obvious voice feature differences.
[0158] - Pseudo-stereo audio: The character marking accuracy is approximately 96.5%, slightly lower than the mono processing effect.
[0159] - True stereo audio: The character marking accuracy is approximately 94.8%. Although the overall accuracy is slightly lower, it performs better in complex scenarios where multiple people are speaking simultaneously.
[0160] 3. Comparison of applicable scenarios: - Mono audio processing is applicable to scenarios with high recording quality and obvious speaker alternation, such as interview programs recorded with a single microphone.
[0161] - Pseudo-stereo audio processing is applicable to scenarios where the source file is stereo but the actual content is the same, such as some movie and TV drama dialogues or mono recordings transcribed into stereo.
[0162] - True stereo audio processing is applicable to scenarios where different contents are recorded in the left and right channels, such as meetings or interview programs recorded in multi-microphone stereo In another exemplary embodiment, based on the same inventive concept as the method embodiment, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the audio speaker recognition method provided in the method embodiment of the present invention. Based on such an understanding, the technical solution of this embodiment, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present invention. The foregoing storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc.
[0163] In another exemplary embodiment, based on the same inventive concept as the method embodiment, an electronic device is provided, including a memory and a processor. The memory stores computer instructions that can run on the processor, and when the processor runs the computer instructions, it executes the audio speaker recognition method provided in the method embodiment of the present invention.
[0164] The processor may be a single-core or multi-core central processing unit or a specific integrated circuit, or one or more integrated circuits configured to implement the present invention.
[0165] The embodiments of the subject matter and the functional operations described in this specification can be implemented in: tangibly embodied computer software or firmware, computer hardware including the structures disclosed in this specification and their structural equivalents, or a combination of one or more of them. The embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non-transitory program carrier to be executed by a data processing apparatus or to control the operation of a data processing apparatus. Alternatively or additionally, the program instructions can be encoded on an artificially generated propagated signal, such as a machine-generated electrical, optical, or electromagnetic signal, which is generated to encode and transmit information to a suitable receiver device for execution by a data processing apparatus.
[0166] The processes and logical flows described in this specification can be executed by one or more programmable computers executing one or more computer programs to perform corresponding functions by operating on input data and generating output. The processes and logical flows can also be executed by dedicated logic circuits, such as FPGAs (Field Programmable Gate Arrays) or ASICs (Application Specific Integrated Circuits), and the apparatus can also be implemented as dedicated logic circuits.
[0167] Processors suitable for executing computer programs include, for example, general and / or special-purpose microprocessors, or any other type of central processing unit. Generally, the central processing unit will receive instructions and data from a read-only memory and / or a random access memory. The basic components of a computer include a central processing unit for implementing or executing instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, etc., or the computer will be operatively coupled to such a mass storage device to receive data therefrom or transfer data thereto, or both. However, a computer is not necessarily required to have such devices. In addition, a computer may be embedded in another device, such as a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device such as a universal serial bus (USB) flash drive, just to name a few.
[0168] It should be understood that each block in a flowchart or block diagram can represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the drawings. For example, two consecutive blocks can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0169] The above specific embodiments are detailed descriptions of the present invention. It cannot be determined that the specific embodiments of the present invention are only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions and substitutions can still be made, and all should be regarded as belonging to the protection scope of the present invention.
Claims
1. An audio speaker recognition method, characterized in that, It includes the following steps: S1. Preprocess the input audio, extract the channel information of the input audio, and determine whether the audio type is mono or stereo; S2. For stereo audio, by comparing at least two acoustic feature parameters of the left and right channels, determine whether the audio type is pseudo-stereo or true stereo; S3. Select a processing strategy according to the audio type: directly execute step S4 for mono audio; select the optimal channel for pseudo-stereo and then execute step S4; separate the left and right channels for true stereo and execute step S4 respectively, and merge the processing results of the left and right channels through a time alignment algorithm; S4. Audio content recognition: perform noise reduction preprocessing, voice activity detection, speaker segmentation, content recognition, and punctuation restoration on the selected channel in sequence; S5. Use a large language model to perform speaker role marking on the results of audio content recognition; S6. Merge the content segments of the same speaker role to generate a structured output.
2. The audio speaker recognition method according to claim 1, wherein The determining whether the audio type is pseudo-stereo or true stereo by comparing at least two acoustic feature parameters of the left and right channels includes: Combining the energy difference and cosine similarity of the left and right channels to determine the audio type: ; Among them, represents the judged audio type, represents mono, represents pseudo stereo, represents true stereo; represents the number of channels of the audio; represents the energy difference; represents the cosine similarity; is the threshold of the energy difference, is the threshold of the cosine similarity; represents other situations.
3. An audio speaker recognition method according to claim 1, characterized in that The selecting the optimal channel for pseudo-stereo includes: Selecting the channel with higher quality through signal-to-noise ratio estimation.
4. An audio speaker recognition method according to claim 1, characterized in that, The using a large language model to perform speaker role marking on the results of audio content recognition includes: According to the preset candidate role set, mapping the recognized speakers to actual roles. Among them, when the number of recognized speakers is 1 and the number of candidate roles is greater than 1, use the large language model to perform secondary speaker separation according to the content; when the number of recognized speakers is greater than the number of candidate roles, use the large language model to merge and map multiple speakers to the candidate role set; Performing semantic analysis on the content through the large language model to determine whether there is a situation where the content of multiple speakers is merged.
5. An audio speaker recognition method according to claim 1, characterized in that, The merging the content segments of the same speaker role to generate a structured output includes: The merging conditions are as follows: ; Among them, indicates whether to merge segments and the decision result; and respectively represent the speaker roles of segments and ; and respectively represent the end and start timestamps of segments and ; represents the time threshold; Generating a structured output including speaker role, content, and timestamp.
6. An audio speaker recognition method according to claim 1, characterized in that, The merging the processing results of the left and right channels through a time alignment algorithm includes: Arranging the speaker segments of the left and right channels in chronological order to generate a unified timeline; Performing conflict processing on the segments that overlap in time but have the same content, and selecting which channel's recognition result to retain according to the signal strength and content integrity; When there is partial overlap in time and different content between the left and right channels, the system marks it as speaking simultaneously and retains the content of both channels.
7. An audio speaker recognition system, characterized in that, It includes: An audio processing and feature extraction module, which is used to preprocess the input audio and extract the channel information of the input audio; An audio type judgment and processing module, which is used to judge whether the audio type is mono or stereo. For stereo audio, by comparing at least two acoustic feature parameters of the left and right channels, it judges whether the audio type is pseudo-stereo or true stereo; it is also used to select a processing strategy according to the audio type: directly perform speaker recognition and content extraction for mono audio; select the optimal channel for pseudo-stereo and then perform speaker recognition and content extraction; separate the left and right channels of true stereo and perform speaker recognition and content extraction respectively, and merge the processing results of the left and right channels through a time alignment algorithm; A speaker recognition and content extraction module, which is used to perform noise reduction preprocessing, voice activity detection, speaker segmentation, content recognition and punctuation restoration on the selected channel in sequence; A role marking module, which is used to perform speaker role marking on the result of audio content recognition by using a large language model; A result optimization and output module, which is used to merge the content segments of the same speaker role and generate a structured output.
8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the audio speaker recognition method described in any one of claims 1-6.
9. An electronic device, comprising a memory and a processor, wherein computer instructions capable of running on the processor are stored on the memory, characterized in that When the processor runs computer instructions, it executes the audio speaker recognition method described in any one of claims 1-6.
Citation Information
Patent Citations
Method and device for recognizing pseudo stereo audio and storage medium
CN107659888A
Single track audio frequency determination method and device
CN108962268A
Audio file processing method and device, electronic equipment and storage medium
CN110941415A
Speech recognition method, device and system and storage medium
CN111883132A
Speech recognition method and device and storage medium
CN116153295A
Cited By
LLM enhancement-based multi-speaker voice recognition and voiceprint matching system
CN121306145A