Speech recognition method and device for multi-speaker environment and electronic equipment

Through the combination of voice activity detection, automatic speech recognition, speaker separation and time alignment algorithms, the problem of low speech recognition accuracy in multiple people's speech scenes is solved, and efficient and accurate speech recognition for multi-speakers is achieved.

CN120126480APending Publication Date: 2025-06-10CHINA TELECOM ARTIFICIAL INTELLIGENCE TECHNOLOGY (BEIJING) CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510287485.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The existing voice recognition technology has low recognition accuracy in scenarios where multiple people speak at the same time, making it difficult to accurately locate the spokesperson corresponding to each voice segment.

Method used

The audio data is calibrated by voice activity detection technology, combined with automatic speech recognition technology for transcription processing, and cluster analysis is performed using speaker separation technology, and the transcription text set and fragment data set are fused through a time alignment algorithm to obtain the final recognition results of the audio data.

Benefits of technology

It significantly improves the accuracy and efficiency of speech recognition in the environment of multiple speakers, can effectively distinguish the speech of multiple speakers, and improves the practical value and user experience of speech recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120126480A_ABST
    Figure CN120126480A_ABST
Patent Text Reader

Abstract

The invention provides a speech recognition method and device for a multi-speaker environment and electronic equipment. Comprises: acquiring audio data; the voice activity detection technology is adopted to calibrate the starting and ending time of each voice in the audio data to obtain an audio calibration result, then the automatic voice recognition technology is adopted to transcribe the audio calibration result to obtain a transcriptional text set corresponding to the audio data, and the transcriptional text set comprises a plurality of audio text segments. The audio text segment is marked with starting and ending time; a speaker separation technology is adopted to perform clustering analysis processing on the audio data to obtain a fragment data set grouped by speakers, and the fragment data set comprises multiple pieces of fragment data for recording starting and ending time of fragments and serial numbers of the speakers; and carrying out fusion processing on the transcriptional text set and the fragment data set by adopting a time alignment algorithm to obtain a final recognition result of the audio data. The problem that the existing speech recognition technology is low in recognition accuracy under the scene that multiple persons speak at the same time is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of speech recognition for multi-speaker environments. Specifically, it relates to a speech recognition method, apparatus, computer-readable storage medium, and electronic device for multi-speaker environments. Background Art

[0002] With the rapid development of speech recognition technology, speech transcription systems have been widely used in scenarios such as real-time meetings, remote communication, and high-precision post-meeting transcription. In these applications, it is common to encounter situations where multiple speakers take turns speaking. How to accurately identify the speech content and corresponding speaking duration of each speaker has become a key challenge for speech recognition systems. However, in multi-speaker environments, traditional ASR methods often show limitations when dealing with speaker alternation, short pauses, and background noise, which may lead to recognition errors and it is difficult to accurately locate the speaker corresponding to each speech segment.

[0003] To solve this problem, in recent years, speaker diarization (SD) technology has gradually been introduced into speech recognition systems to improve the recognition accuracy in multi-speaker scenarios. Although SD technology has powerful functions, speech recognition that only uses SD technology still has problems with inaccurate recognition in scenarios where multiple people speak simultaneously. Summary of the Invention

[0004] The main objective of this application is to provide a speech recognition method, apparatus, computer-readable storage medium, and electronic device for multi-speaker environments, so as to at least solve the problem of low recognition accuracy of existing speech recognition technology in scenarios where multiple people speak simultaneously.

[0005] To achieve the above objective, according to one aspect of this application, a speech recognition method for multi-speaker environments is provided, including: an acquisition step of acquiring audio data; a first processing step of using speech activity detection technology to calibrate the start and end times of each speech in the audio data to obtain an audio calibration result, and then using automatic speech recognition technology to transcribe the audio calibration result to obtain a transcription text set corresponding to the audio data, where the transcription text set includes multiple audio text segments, and the audio text segments are marked with start and end times; a second processing step of using speaker diarization technology to perform clustering analysis on the audio data to obtain a segment data set grouped by speakers, where the segment data set includes multiple segment data records of the start and end times and speaker numbers; a third processing step of using a time alignment algorithm to fuse the transcription text set and the segment data set to obtain the final recognition result of the audio data.

[0006] Optionally, a time alignment algorithm is used to fuse the transcribed text set and the segment data set to obtain the final recognition result of the audio data, including: traversal step: obtaining the audio text segments in the transcribed text set, traversing the segment data set, and determining the segment data corresponding to the audio text segment based on the start and end times of the audio text segment; numbering process step: numbering the audio text segments according to the speaker numbers of the segment data to obtain speech recognition segments; the traversal step and the numbering process step are executed multiple times until all the audio text segments in the transcribed text set are traversed to obtain the final recognition result.

[0007] Optionally, determining the segment data corresponding to the audio text segment includes: determining at least one segment data that overlaps in time with the audio text segment according to the start and end times of the audio text segment and the start and end times of the segment data; determining whether the number of segment data that overlaps in time with the audio text segment is 1; in the case where the number of segment data is 1, determining that this segment data corresponds to the audio text segment.

[0008] Optionally, determining the segment data corresponding to the audio text segment further includes: in the case where the number of segment data that overlaps in time with the audio text segment is greater than 1, calculating the weights of the segment data that overlap in time with the audio text segment according to the overlap duration between the segment data and each audio text segment; using the segment data with the largest weight as the segment data corresponding to the audio text segment.

[0009] Optionally, calculating the weights of the segment data that overlap in time with the audio text segment according to the overlap duration between the segment data and the audio text segment includes: according to the formula: determining the weights of the segment data, where ω i is the weight of the segment data, slice_duration is the overlap duration between the segment data and the audio text segment, and seg_duration m is the total duration of the audio text segment.

[0010] Optionally, after obtaining the segment data set grouped by speakers, the method further includes: determining whether the start and end times of the segment data set cover the start and end times of the transcribed text set; in the case where the start and end times of the transcribed text set are not covered, continuing to use the speaker separation technology to obtain segment data.

[0011] Optionally, the first processing step and the second processing step are executed in parallel.

[0012] According to another aspect of the present application, there is provided a speech recognition device for a multi-speaker environment, including: a first acquisition unit configured to perform an acquisition step to acquire audio data; a first processing unit configured to perform a first processing step, using voice activity detection technology to calibrate the start and end times of each speech in the audio data to obtain an audio calibration result, and then using automatic speech recognition technology to transcribe the audio calibration result to obtain a transcription text set corresponding to the audio data, wherein the transcription text set includes multiple audio text segments, and the audio text segments are marked with start and end times; a second processing unit configured to perform a second processing step, using speaker separation technology to perform clustering analysis on the audio data to obtain a segment data set grouped by speakers, wherein the segment data set includes multiple segment data records of segment start and end times and speaker numbers; a third processing unit configured to perform a third processing step, using a time alignment algorithm to fuse the transcription text set and the segment data set to obtain a final recognition result of the audio data.

[0013] According to still another aspect of the present application, there is provided a computer-readable storage medium, which includes a stored program, wherein when the program runs, it controls the device where the computer-readable storage medium is located to execute any one of the speech recognition methods for a multi-speaker environment.

[0014] According to yet another aspect of the present application, there is provided an electronic device, including: one or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the one or more programs include those for executing any one of the speech recognition methods for a multi-speaker environment.

[0015] Applying the technical solution of the present application, the acquisition step is to obtain audio data; the first processing step is to use voice activity detection technology to calibrate the start and end times of each voice in the audio data to obtain an audio calibration result, and then use automatic speech recognition technology to transcribe the audio calibration result to obtain a transcription text set corresponding to the audio data. The transcription text set includes multiple audio text segments, and the audio text segments are marked with start and end times; the second processing step is to use speaker separation technology to perform clustering analysis on the audio data to obtain a segment data set grouped by speakers, where the segment data set includes multiple segment data records of the start and end times and speaker numbers; the third processing step is to use a time alignment algorithm to fuse the transcription text set and the segment data set to obtain the final recognition result of the audio data. By combining voice activity detection, automatic speech recognition, speaker separation, and time alignment algorithms, the accuracy and efficiency of recognition are significantly improved. Voice activity detection technology ensures the accurate calibration of voice segments, automatic speech recognition technology provides fast text transcription, speaker separation technology realizes the separation of voices of different speakers, and the time alignment algorithm ensures the accuracy of the correspondence between the transcribed text and the speaker. This method is particularly suitable for scenarios such as meeting records, court records, and large-scale multi-person discussions, and can effectively distinguish the speeches of multiple speakers, enhancing the practical value and user experience of speech recognition. Brief Description of the Drawings

[0016] The specification drawings forming a part of the present application are used to provide a further understanding of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:

[0017] Figure 1 Shows a hardware structure block diagram of a mobile terminal for implementing a speech recognition method for a multi-speaker environment provided in an embodiment of the present application;

[0018] Figure 2 Shows a schematic flowchart of a speech recognition method for a multi-speaker environment provided in an embodiment of the present application;

[0019] Figure 3 Shows a schematic flowchart of a specific speech recognition method for a multi-speaker environment provided in an embodiment of the present application;

[0020] Figure 4 Shows a schematic diagram of a specific speech recognition system for a multi-speaker environment provided in an embodiment of the present application;

[0021] Figure 5 Shows a structure block diagram of a speech recognition device for a multi-speaker environment provided in an embodiment of the present application.

[0022] Among them, the above-mentioned drawings include the following reference numerals:

[0023] 102, processor; 104, memory; 106, transmission device; 108, input / output device. Detailed implementation manners

[0024] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other. The present application will be described in detail below with reference to the drawings and in combination with the embodiments.

[0025] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0026] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so as to implement the embodiments of the present application described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0027] For the convenience of description, some nouns or terms related to the embodiments of the present application are described below:

[0028] Automatic Speech Recognition, ASR, is a technology that converts speech signals into text, aiming to enable a computer to understand and process human spoken language. ASR systems usually rely on complex models and algorithms to analyze speech signals, extract features, and match them with pre-trained speech models to generate corresponding text. Common ASR applications include intelligent assistants, speech-to-text systems, voice-controlled devices, etc.

[0029] Speaker Diarization, SD, is a technology used in multi-speaker environments, aiming to label and distinguish the speech segments of different speakers in an audio.

[0030] Voice Activity Detection, VAD. Voice Activity Detection is a technology used to detect the presence of voice activity in an audio signal, capable of distinguishing voice segments from non-voice segments (such as silence, noise, etc.).

[0031] As introduced in the background art, existing speech recognition technologies have relatively low recognition accuracy in scenarios where multiple people speak simultaneously. To address the problem of low recognition accuracy of existing speech recognition technologies in scenarios where multiple people speak simultaneously, embodiments of the present application provide a speech recognition method, apparatus, computer-readable storage medium, and electronic device for a multi-speaker environment.

[0032] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention.

[0033] The method embodiments provided in the embodiments of the present application can be executed on a mobile terminal, a computer terminal, or a similar computing device. Taking running on a mobile terminal as an example,[[]] Figure 1 is a hardware structure block diagram of a mobile terminal of a speech recognition method for a multi-speaker environment according to an embodiment of the present invention. As Figure 1 shown, the mobile terminal may include one or more ( Figure 1 only one is shown in the figure) processors 102 (the processor 102 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data. Among them, the above-mentioned mobile terminal may further include a transmission device 106 for communication functions and an input / output device 108. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only schematic and does not limit the structure of the above-mentioned mobile terminal. For example, the mobile terminal may further include more or fewer components than Figure 1 shown in the figure, or have a different configuration from Figure 1 shown in the figure.

[0034] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to the voice recognition method for a multi-speaker environment in the embodiments of the present invention. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, implements the above-mentioned method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely disposed relative to the processor 102, and these remote memories can be connected to the mobile terminal through a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof. The transmission device 106 is used to receive or send data via a network. Specific examples of the above-mentioned network may include a wireless network provided by a communication provider of the mobile terminal. In one instance, the transmission device 106 includes a network adapter (Network Interface Controller, abbreviated as NIC), which can be connected to other network devices through a base station and thus can communicate with the Internet. In one instance, the transmission device 106 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0035] In this embodiment, a voice recognition method for a multi-speaker environment running on a mobile terminal, a computer terminal, or a similar computing device is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0036] Figure 2 It is a flowchart of the voice recognition method for a multi-speaker environment according to an embodiment of the present application. As Figure 2 shown, the method includes the following steps:

[0037] Step S201, an acquisition step, to acquire audio data;

[0038] Step S202, a first processing step, using voice activity detection technology to calibrate the start and end times of each voice in the above audio data to obtain an audio calibration result, and then using automatic speech recognition technology to transcribe the above audio calibration result to obtain a transcription text set corresponding to the above audio data, where the above transcription text set includes multiple audio text segments, and the above audio text segments are marked with start and end times;

[0039] Specifically, Voice Activity Detection (VAD) technology is mainly used to identify the speech and non-speech parts in audio. By analyzing the features of the audio signal, such as energy, zero-crossing rate, etc., it determines the start and end times of the speech. Automatic Speech Recognition (ASR) technology then converts these speech segments into text.

[0040] Step S203, the second processing step, uses speaker separation technology to perform clustering analysis on the above audio data to obtain a segment data set grouped by speakers. Among them, the above segment data set includes multiple segment data records of the start and end times and speaker numbers.

[0041] Specifically, speaker separation technology, such as the speaker recognition model in deep learning, can distinguish the voices of different speakers and can effectively separate them even in the case of overlapping speech.

[0042] Step S204, the third processing step, uses a time alignment algorithm to fuse the above transcription text set and the above segment data set to obtain the final recognition result of the above audio data.

[0043] Specifically, time alignment algorithms, such as Dynamic Time Warping (DTW) or Hidden Markov Model (HMM), are used to match and align the timestamps of the transcription text and the speaker segments to ensure the accuracy of the recognition result.

[0044] Through this embodiment, in the acquisition step, audio data is obtained; in the first processing step, the start and end times of each speech in the audio data are calibrated using voice activity detection technology to obtain an audio calibration result, and then the audio calibration result is transcribed using automatic speech recognition technology to obtain a transcription text set corresponding to the audio data. The transcription text set includes multiple audio text segments, and the audio text segments are marked with start and end times; in the second processing step, the audio data is subjected to clustering analysis using speaker separation technology to obtain a segment data set grouped by speakers. The segment data set includes multiple segment data records of the start and end times and speaker numbers; in the third processing step, a time alignment algorithm is used to fuse the transcription text set and the segment data set to obtain the final recognition result of the audio data. By combining the use of voice activity detection, automatic speech recognition, speaker separation, and time alignment algorithms, the accuracy and efficiency of recognition are significantly improved. The voice activity detection technology ensures the precise calibration of speech segments, the automatic speech recognition technology provides fast text transcription, the speaker separation technology realizes the separation of speech of different speakers, and the time alignment algorithm ensures the accuracy of the correspondence between the transcribed text and the speaker. This method is particularly applicable to scenarios such as meeting records, court records, and large-scale multi-person discussions, and can effectively distinguish the speeches of multiple speakers, enhancing the practical value and user experience of speech recognition.

[0045] This method is not limited to the combination of the above technical features, and can further integrate preprocessing technologies such as noise suppression and echo cancellation, as well as post-processing technologies such as text correction and semantic understanding, to enhance the adaptability to complex environments and the usability of recognition results.

[0046] In the specific implementation process, a time alignment algorithm is used to fuse the above transcription text set and the above segment data set to obtain the final recognition result of the above audio data, including: traversal step: obtaining the above audio text segments in the above transcription text set, traversing the above segment data set, and determining the segment data corresponding to the above audio text segments based on the start and end times of the above audio text segments; numbering processing step: numbering the above audio text segments according to the speaker numbers of the above segment data to obtain speech recognition segments; the above traversal step and the above numbering processing step are executed multiple times until all the above audio text segments in the above transcription text set are traversed to obtain the above final recognition result.

[0047] The time alignment algorithm of this method is the key to ensuring the correct matching of the transcribed text with the speaker segments. During the traversal process, the algorithm compares the start and end times of the audio text segments with those of each segment in the segment dataset to find the speaker segment that best matches in terms of time. Through the correspondence of the speaker numbers, the recognized text segments can be classified under the correct speaker, thus constructing a complete recognition result. This method can effectively handle the time alignment problem in a multi-speaker environment and maintain a high recognition accuracy even in cases of frequent speaker alternation and overlapping speech.

[0048] Among them, the implementation of the time alignment algorithm can include but is not limited to dynamic time warping (DTW), hidden Markov model (HMM), deep learning-based time series alignment models, etc. These algorithms can flexibly adjust the alignment strategy according to the needs of different scenarios, further improving the accuracy and robustness of recognition.

[0049] Specifically, determining the segment data corresponding to the above audio text segment includes: determining at least one of the above segment data that overlaps in time with the above audio text segment according to the start and end times of the above audio text segment and the start and end times of the above segment data; determining whether the number of the above segment data that overlaps in time with the above audio text segment is 1; in the case where the number of the above segment data is 1, determining that this one of the above segment data corresponds to the above audio text segment.

[0050] When this method determines the correspondence between the segment data and the audio text segment, it first filters out the possible matching segments through time overlap. If only one segment data overlaps in time with the audio text segment, then this segment data is the correct match. This method is simple and efficient, can quickly locate the correct speaker segment, reduces the computational amount, and improves the processing speed. In addition, the determination of the segment data can also be achieved by calculating the similarity between the segment data and the audio text segment, such as similarity calculation based on deep learning, or intelligent matching in combination with the context, to improve the recognition accuracy in complex environments.

[0051] More specifically, determining the segment data corresponding to the above audio text segment further includes: in the case where the number of the above segment data that overlaps in time with the above audio text segment is greater than 1, calculating the weights of the above segment data that overlap in time with the above audio text segment according to the overlap duration between the above segment data and each of the above audio text segments; taking the segment data with the largest weight as the segment data corresponding to the above audio text segment.

[0052] When there are multiple fragment data overlapping with the audio text fragment in terms of time, this method determines the most matching fragment data by calculating the weights of the overlapping durations. This method can effectively handle the situations of speaker alternation or overlapping speech. By quantifying the degree of time overlap, the accuracy of the recognition result is ensured. The weight calculation is based on the ratio of the overlapping duration between the fragment data and the audio text fragment to the total duration of the audio text fragment. This not only considers the time matching degree but also takes into account the integrity of the audio text fragment, thereby improving the reliability of the recognition. Additionally, in addition to the weight calculation based on time overlap, audio features (such as frequency, pitch, etc.) and semantic analysis of the text content can be combined to further optimize the matching process of the fragment data, making it applicable to a wider range of scenarios, such as meeting records in noisy environments, discussions in multilingual environments, etc.

[0053] Further, according to the overlapping duration between the above-mentioned fragment data and the above-mentioned audio text fragment, calculate the weights of each of the above-mentioned fragment data that overlaps with the above-mentioned audio text fragment in terms of time, including: According to the formula: Determine the weights of each of the above-mentioned fragment data, where ω i is the weight of the above-mentioned fragment data, slice_duration is the overlapping duration between the above-mentioned fragment data and the above-mentioned audio text fragment, and seg_duration m is the total duration of the above-mentioned audio text fragment.

[0054] This method calculates the weights of the fragment data through the above formula, which can quantify the time matching degree between the fragment data and the audio text fragment. The higher the weight value, the higher the matching degree between the fragment data and the audio text fragment. This method is simple and intuitive, and can quickly evaluate the matching situations of multiple fragment data, providing a reliable basis for the final recognition result. Additionally, the weight calculation formula can be further optimized. For example, the similarity calculation of speaker features can be added, or the context of the audio text fragment can be considered to adapt to more complex and changeable recognition environments and improve the accuracy and efficiency of the recognition.

[0055] Even further, after obtaining the fragment data set grouped by speakers, the above method further includes: determining whether the start and end times of the above-mentioned fragment data set cover the start and end times of the above-mentioned transcription text set; in the case where the start and end times of the transcription text set are not covered, continue to use the above-mentioned speaker separation technology to obtain fragment data.

[0056] After obtaining the ASR segment, the system checks whether the time of the current SD result is sufficient to cover the ASR segment. If the SD segment is insufficient, more SD segmentation results are obtained through the function of obtaining the SD result. If sufficient SD results cannot be obtained, it is considered that there is a problem with the speaker segmentation task, and the processing is terminated. Among them, to solve the problem of insufficient SD segments, the system will call the "function of obtaining the SD result", which is usually an interaction with the SD service, aiming to request more SD processing results to cover the time range of the ASR text segment for which the speaker has not been assigned. For example, if a 4-second text is recognized by ASR, but the SD currently only has 2 seconds of speaker information, then the system will request the SD service to continue processing the remaining 2 seconds of audio until sufficient SD results are returned to cover the entire time range of the ASR text segment.

[0057] In addition, if the system attempts to obtain more SD results but fails, or the SD service cannot provide a sufficient time range to cover all ASR segments, then the system will consider that there is a problem with the speaker segmentation task. Such problems may stem from reasons such as delays in SD processing, limitations of the SD algorithm itself, or poor input audio quality, resulting in the system being unable to accurately or completely identify the speaker. In such a case, to avoid outputting inaccurate or incomplete results, the system will terminate the current speaker ID assignment process, which may mean temporarily stopping the fusion of ASR and SD results until the problem is solved or other remedial measures are taken.

[0058] By checking the integrity of the segment dataset, it is ensured that all audio text segments have corresponding speaker segments, avoiding recognition omissions caused by incomplete segment datasets. This method can effectively improve the integrity of the recognition results. Even in long recordings, it can ensure that the speeches of all speakers are accurately recognized and recorded. The integrity check can include, but is not limited to, time coverage check, speaker feature check, etc., to ensure the comprehensiveness and accuracy of the recognition process, and is applicable to scenarios with long recordings and frequent alternation of multiple speakers, such as long meeting records, court trial records, etc.

[0059] Specifically, the above first processing step and the above second processing step are executed in parallel.

[0060] Through the parallelized ASR and SD result acquisition and the processing mechanism of the ASR and SD merged results, this method can greatly improve the efficiency of speech processing in real-time meeting scenarios, is particularly suitable for large-scale real-time audio data processing, and supports high-concurrency real-time meeting transcription applications.

[0061] To enable those skilled in the art to more clearly understand the technical solution of this application, the following will detail the implementation process of the speech recognition method for a multi-speaker environment of this application in combination with specific embodiments.

[0062] This embodiment relates to a specific speech recognition method for a multi-speaker environment, aiming to achieve efficient and accurate multi-speaker speech recognition. By first using VAD technology to real-time judge the start and end times of speech in the audio, and applying speech recognition technology to obtain the transcribed text of the speech, at the same time separating the segments of different speakers through SD technology, and finally using the method in this system to perform speaker tagging on each ASR recognition result segment, realizing the result fusion of ASR and SD, and obtaining clear and accurate recognition results for each speaker. This method can be applied to real-time meeting transcription or high-precision transcription scenarios after the meeting, especially suitable for multi-speaker meeting scenarios. The specific contents are as follows:

[0063] The merging module in this embodiment is responsible for merging the output results of the speech recognition module (ASR module) and the speaker segmentation module (SD module). The following is its specific logical process, and the entire technical solution process is as Figure 3 shown:

[0064] Step 1, Streaming Speech Recognition (ASR) Request: The service obtains the audio sent by the client and sends it to the VAD and ASR services, and obtains the recognition result of the streaming ASR. The result includes the text, the start time and end time of the current speech segment, and is put into the ASR result queue.

[0065] Step 2, Obtain Speech Recognition (ASR) Segment: The system obtains the recognition result of a speech segment from the ASR result queue. The result includes the recognized text, meta-information such as the start time and end time of this segment. At the same time, the system will judge whether this segment is the last segment (is_asr_eof). If it is the last segment, it indicates that the speech recognition task is completed.

[0066] Step 3, Streaming Speaker Recognition (SD) Request: While performing Step 1, the audio is sent to the SD service. The result includes the speaker ID, the start time and end time when this speaker speaks in this SD speech segment, and the SD service result is put into the SD result queue.

[0067] Step 4, Check Speaker Segmentation (SD) Segment: After obtaining the ASR segment, the system will check whether the time of the current SD result is sufficient to cover this ASR segment. If the SD segment is insufficient, more SD segmentation results are obtained through the function of obtaining SD results. If insufficient SD results cannot be obtained, it is considered that there is a problem with this speaker segmentation task, and the processing is terminated.

[0068] Step 5: Determine the speaker ID: On the premise of ensuring sufficient SD results, the system calculates the speaker ID corresponding to the start time and end time of the ASR segment through the time alignment algorithm. In this implementation plan, the system uses the "Speaker ID Discrimination" function to accurately judge the speaker identity in a multi-speaker scenario. The key steps of this function include the following aspects:

[0069] 1. Pass in the start time and end time of the current ASR recognition segment and the start and end times of the SD segment;

[0070] 2. Calculate the segment overlap:

[0071] 1). The code first traverses each speaker segment obtained by speaker separation according to the time range of the current ASR segment to find overlapping time segments.

[0072] 2). For the time segment (sd_begin_time, sd_end_time) of each speaker i and the m-th segment (begin_time, end_time) of the ASR, calculate the start time and end time of the overlapping part:

[0073] slice_begin_time = max(sd_begin_time, begin_time);

[0074] slice_end_time = min(sd_end_time, end_time);

[0075] When slice_begin_time < slice_end_time, the duration of the overlapping segment is:

[0076] slice_duration = slice_end_time - slice_begin_time;

[0077] 3). Through the time segment of SD, find the speaker time period overlapping with the ASR segment, and calculate the time proportion of this speaker i in the ASR segment: seg_duration m is the total duration of the current ASR segment.

[0078] 3. Record the weight:

[0079] For each overlapping speaker segment, its weight is accumulated into a hash speaker_weight_map. This hash table records each speaker ID and its corresponding weight value.

[0080] Select the speaker with the largest weight:

[0081] 1) After traversing all speaker segments, the system selects the speaker with the highest weight as the speaker for the current ASR segment by comparing the weight values.

[0082] 2) Find the element with the highest weight in speaker_weight_map to obtain the corresponding speaker ID.

[0083] Establish speaker mapping:

[0084] 1) The ID of this speaker will be mapped to a unique identification ID estimated , and the mapping method is that the speaker ID returned by SD SD is used as the key and the ID in this system is used as the value, corresponding one by one for subsequent use. If this speaker ID has not appeared before, a new ID will be assigned to it estimated , starting from 0 for numbering.

[0085] 2) This ID estimated will be stored in session→speaker_id_map for subsequent use of the same ID to identify the same speaker.

[0086] 4. Result return: Finally, the ID estimated will be output as the speaker ID of this ASR segment, and the ID, text, start time and end time of this segment will be used as the result and returned to the requesting client.

[0087] In order to better perform speech recognition and speaker identity discrimination in a real-time conference system, a service based on the GRPC standard protocol can be encapsulated through the streaming VAD and speaker segmentation and merging system proposed in this embodiment. The service based on this system is linked to the streaming ASR service with VAD and the streaming SD service through the GRPC protocol in two threads, obtaining ASR results and SD result requests in real time, and enqueuing the results respectively. At the same time, the results are merged in another thread:

[0088] Example of ASR result queue:

[0089] {"text":"Welcome everyone to today's meeting","start_time":0,"end_time":4000,"is_asr_eof":false};

[0090] {"text":"Let's first discuss the sales performance of this quarter","start_time":4500,"end_time":10000,"is_asr_eof":false};

[0091] {"text":"Next is the progress report of the R & D department","start_time":10500,"end_time":14000,"is_asr_eof":false};

[0092] {"text":"Finally, we summarize today's decisions","start_time":14500,"end_time":18000,"is_asr_eof":true};

[0093] SD result queue example:

[0094] {"speaker_id":"0","start_time":0,"end_time":3000};

[0095] {"speaker_id":"2","start_time":3000,"end_time":8000};

[0096] {"speaker_id":"0","start_time":8000,"end_time":11000};

[0097] {"speaker_id":"1","start_time":11000,"end_time":16000};

[0098] {"speaker_id":"0","start_time":16000,"end_time":18000};

[0099] Example of processing and merging results:

[0100] 1). For ASR segment 1 ("Welcome everyone to today's meeting"), within the time range 0 to 4000, find overlapping SD results:

[0101] Overlapping SD segments:

[0102] {"speaker_id":"0","start_time":0,"end_time":3000}: Overlap of 3000 milliseconds, weight 0.75;

[0103] {"speaker_id":"2","start_time":3000,"end_time":4000}: Overlap of 1000 milliseconds, weight 0.25;

[0104] Select speaker_id: 0 as the speaker because the overlapping time of 0 is longer and the weight is greater, and obtain the speaker ID in this system. Since there is no speaker yet, the result ID_sd = 0 returned by SD is mapped to the speaker ID in this system, which is also 0, ID = 0.

[0105] 2) For the ASR segment 2 ("Let's first discuss the sales performance this quarter"), the time range is 4500 to 10000. Compare with the SD result:

[0106] Overlapping SD segments:

[0107] {"speaker_id":"2","start_time":4500,"end_time":8000}: Overlap for 3500 milliseconds, weight is approximately 0.64;

[0108] {"speaker_id":"0","start_time":8000,"end_time":10000}: Overlap for 2000 milliseconds, weight is approximately 0.37;

[0109] Select speaker_id: 2 because its overlapping time is longer and the weight is greater, and obtain the speaker ID in this system. The maximum system ID is 0, and ID_sd = 2 cannot be found in the system map. So the result ID_sd = 2 returned by SD is mapped to the speaker ID in this system, which is also 1, ID = 1.

[0110] 3) For the ASR segment 3 ("Next is the progress report of the R & D department"), the time range is 10500 to 14000. Compare with the SD result:

[0111] Overlapping SD segments:

[0112] {"speaker_id":"1","start_time":11000,"end_time":14000}: Overlap for 3000 milliseconds;

[0113] Select speaker_id: 1, and the corresponding ID in this system is 2.

[0114] 4) For the ASR segment 4 ("Finally, we summarize today's decisions"), the time range is 14500 to 18000. Compare with the SD result:

[0115] Overlapping SD segments:

[0116] {"speaker_id":"2","start_time":14500,"end_time":16000}: Overlap for 1500 milliseconds;

[0117] {"speaker_id":"0","start_time":16000,"end_time":18000}: Overlap for 2000 milliseconds;

[0118] Select speaker_id: 0.

[0119] 5), The final result of merging:

[0120] Based on the above analysis, the merged result is as follows:

[0121] {"speaker_id":"0","text":"Welcome everyone to today's meeting","start_time":0,"end_time":4000};

[0122] {"speaker_id":"1","text":"Let's first discuss the sales performance of this quarter","start_time":4500,"end_time":10000};

[0123] {"speaker_id":"2","text":"Next is the progress report of the R & D department","start_time":10500,"end_time":14000};

[0124] {"speaker_id":"0","text":"Finally, we summarize today's decisions","start_time":14500,"end_time":18000};

[0125] This process combines the information of ASR and SD, ensures that the text segments are correctly matched with the corresponding speakers, and immediately returns the speaker ID of each determined ASR segment to the client to achieve low latency. The process of the entire embodiment is as Figure 4 shown.

[0126] This embodiment has the following advantages:

[0127] 1. Efficient parallel processing: Through the parallelized ASR and SD result acquisition and the processing mechanism of the ASR and SD merged results, this system can greatly improve the efficiency of voice processing in real-time meeting scenarios, especially suitable for large-scale real-time audio data processing, and supports high-concurrency real-time meeting transcription applications.

[0128] 2. Intelligent merging of ASR and SD segments: By merging the ASR and SD results, the problem of segment inconsistency is avoided, ensuring the integrity of the ASR recognition results.

[0129] 3. Time Alignment and Speaker Prediction: By adopting a combined algorithm of time alignment and speaker ID prediction and through weights, the system can ensure the accuracy of the recognized text and speaker information, reduce the situation of incorrect matching, enable the speakers to be accurately distinguished to the greatest extent in the application scenario, and greatly improve the user experience.

[0130] 4. Speaker ID Mapping: By remapping the speaker ID, the continuity of the speaker ID in the output result is ensured, and the user experience in the multi-speaker recognition scenario is improved.

[0131] 5. Precise Flow Termination Mechanism: Through reasonable flow termination conditions, the system avoids the risk of task termination in the middle or result loss and ensures a complete processing flow.

[0132] The embodiment of the present application also provides a speech recognition device for a multi-speaker environment. It should be noted that the speech recognition device for a multi-speaker environment in the embodiment of the present application can be used to execute the speech recognition method for a multi-speaker environment provided by the embodiment of the present application. This device is used to implement the above embodiments and preferred implementation manners, and those that have been described will not be repeated. As used hereinafter, the term "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.

[0133] The following introduces the speech recognition device for a multi-speaker environment provided by the embodiment of the present application.

[0134] Figure 5 is a schematic diagram of the speech recognition device for a multi-speaker environment according to the embodiment of the present application. As Figure 5 shown, the device includes:

[0135] A first acquisition unit 51, configured to execute an acquisition step to acquire audio data;

[0136] A first processing unit 52, configured to execute a first processing step, use voice activity detection technology to calibrate the start and end times of each speech in the above audio data to obtain an audio calibration result, and then use automatic speech recognition technology to transcribe the above audio calibration result to obtain a transcription text set corresponding to the above audio data, where the above transcription text set includes multiple audio text segments, and the above audio text segments are marked with start and end times;

[0137] A second processing unit 53, configured to execute a second processing step, perform clustering analysis processing on the above audio data by using a speaker separation technique, and obtain a segment data set grouped by speakers, where the segment data set includes a plurality of segment data records of start and end times and speaker numbers;

[0138] A third processing unit 54, configured to execute a third processing step, perform fusion processing on the above transcription text set and the above segment data set by using a time alignment algorithm, and obtain a final recognition result of the above audio data.

[0139] In this embodiment, a first acquisition unit is configured to execute an acquisition step to acquire audio data; a first processing unit is configured to execute a first processing step, perform calibration processing on start and end times of each voice in the audio data by using a voice activity detection technique, obtain an audio calibration result, and then perform transcription processing on the audio calibration result by using an automatic speech recognition technique to obtain a transcription text set corresponding to the audio data, where the transcription text set includes a plurality of audio text segments, and the audio text segments are marked with start and end times; a second processing unit is configured to execute a second processing step, perform clustering analysis processing on the audio data by using a speaker separation technique, and obtain a segment data set grouped by speakers, where the segment data set includes a plurality of segment data records of start and end times and speaker numbers; a third processing unit is configured to execute a third processing step, perform fusion processing on the transcription text set and the segment data set by using a time alignment algorithm, and obtain a final recognition result of the audio data. By combining the use of voice activity detection, automatic speech recognition, speaker separation, and time alignment algorithms, the accuracy and efficiency of recognition are significantly improved. The voice activity detection technique ensures the accurate calibration of voice segments, the automatic speech recognition technique provides fast text transcription, the speaker separation technique realizes the separation of voices of different speakers, and the time alignment algorithm ensures the accuracy of the correspondence between the transcription text and the speakers. This method is particularly applicable to scenarios such as meeting records, court records, and large-scale multi-person discussions, and can effectively distinguish the speeches of multiple speakers, improving the practical value and user experience of speech recognition.

[0140] As an optional solution, the third processing unit includes a traversal module, a number processing module, and an execution module; the traversal module is configured to execute a traversal step: acquire the above audio text segments in the above transcription text set, traverse the above segment data set, and determine segment data corresponding to the above audio text segments based on the start and end times of the above audio text segments; the number processing module is configured to execute a number processing step: number the above audio text segments according to the speaker numbers of the above segment data to obtain speech recognition segments; the execution module is configured to execute the above traversal step and the above number processing step multiple times until all the above audio text segments in the above transcription text set are traversed to obtain the above final recognition result.

[0141] An alternative solution is that the traversal module includes a first determination sub-module, a second determination sub-module, and a third determination sub-module. The first determination sub-module is configured to determine at least one of the above segment data that overlaps with the above audio text segment according to the start and end times of the above audio text segment and the start and end times of the above segment data. The second determination sub-module is configured to determine whether the number of the above segment data that overlaps with the above audio text segment is 1. The third determination sub-module is configured to determine that this one of the above segment data corresponds to the above audio text segment when the number of the above segment data is 1.

[0142] An alternative solution is that the traversal module further includes a calculation sub-module and a fourth determination sub-module. The calculation sub-module is configured to calculate the weights of the above segment data that overlap with the above audio text segment according to the overlapping duration between the above segment data and each of the above audio text segments when it is determined that the number of the above segment data that overlaps with the above audio text segment is greater than 1. The fourth determination sub-module is configured to determine the segment data with the largest weight as the segment data corresponding to the above audio text segment.

[0143] An alternative solution is that the calculation sub-module includes a fifth determination sub-module, which is configured to determine the weights of the above segment data according to the formula: where ω i is the weight of the above segment data, slice_duration is the overlapping duration between the above segment data and the above audio text segment, and seg_duration m is the total duration of the above audio text segment.

[0144] An alternative solution is that the device further includes a determination unit and a second acquisition unit. The determination unit is configured to determine whether the start and end times of the above segment data set cover the start and end times of the above transcription text set after obtaining the segment data set grouped by speakers. The second acquisition unit is configured to continue to obtain segment data by using the above speaker separation technology when the start and end times of the transcription text set are not covered.

[0145] An alternative solution is that the above first processing step and the above second processing step are executed in parallel.

[0146] The above speech recognition device for a multi-speaker environment includes a processor and a memory. The above first acquisition unit, first processing unit, second processing unit, third processing unit, etc. are all stored in the memory as program units, and the processor executes the above program units stored in the memory to implement corresponding functions. The above modules are all located in the same processor; or, the above modules are respectively located in different processors in any combination.

[0147] The processor contains a kernel, which retrieves the corresponding program units from the memory. One or more kernels can be set, and by adjusting the kernel parameters, the problem of low recognition accuracy in the existing speech recognition technology in the scenario of multiple people speaking simultaneously can be solved.

[0148] The memory may include non-permanent memory in a computer-readable medium, in the form of random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM), and the memory includes at least one storage chip.

[0149] An embodiment of the present invention provides a computer-readable storage medium, and the above computer-readable storage medium includes a stored program. Wherein, when the above program runs, it controls the device where the above computer-readable storage medium is located to execute the above speech recognition method for a multi-speaker environment.

[0150] Specifically, the speech recognition method for a multi-speaker environment includes:

[0151] Step S201, an acquisition step, to acquire audio data;

[0152] Step S202, a first processing step, using voice activity detection technology to calibrate the start and end times of each voice in the above audio data to obtain an audio calibration result, and then using automatic speech recognition technology to transcribe the above audio calibration result to obtain a transcription text set corresponding to the above audio data. Wherein, the above transcription text set includes multiple audio text segments, and the above audio text segments are marked with start and end times;

[0153] Step S203, a second processing step, using speaker separation technology to perform clustering analysis on the above audio data to obtain a segment data set grouped by speakers. Wherein, the above segment data set includes multiple segment data records of the start and end times and speaker numbers;

[0154] Step S204, a third processing step, using a time alignment algorithm to fuse the above transcription text set and the above segment data set to obtain the final recognition result of the above audio data.

[0155] An embodiment of the present invention provides a processor, and the above processor is used to run a program. Wherein, when the above program runs, it executes the above speech recognition method for a multi-speaker environment.

[0156] Specifically, the speech recognition method for a multi-speaker environment includes:

[0157] Step S201, an acquisition step, to acquire audio data;

[0158] Step S202, the first processing step: Use voice activity detection technology to calibrate the start and end times of each voice in the above audio data to obtain an audio calibration result, and then use automatic speech recognition technology to transcribe the above audio calibration result to obtain a transcription text set corresponding to the above audio data. Among them, the above transcription text set includes multiple audio text segments, and the above audio text segments are marked with start and end times;

[0159] Step S203, the second processing step: Use speaker separation technology to perform clustering analysis on the above audio data to obtain a segment data set grouped by speakers. Among them, the above segment data set includes multiple segment data records of segment start and end times and speaker numbers;

[0160] Step S204, the third processing step: Use a time alignment algorithm to fuse the above transcription text set and the above segment data set to obtain the final recognition result of the above audio data.

[0161] An embodiment of the present invention provides an electronic device, which includes a processor, a memory, and a program stored on the memory and executable on the processor. When the processor executes the program, it implements at least the following steps:

[0162] Step S201, the acquisition step: Acquire audio data;

[0163] Step S202, the first processing step: Use voice activity detection technology to calibrate the start and end times of each voice in the above audio data to obtain an audio calibration result, and then use automatic speech recognition technology to transcribe the above audio calibration result to obtain a transcription text set corresponding to the above audio data. Among them, the above transcription text set includes multiple audio text segments, and the above audio text segments are marked with start and end times;

[0164] Step S203, the second processing step: Use speaker separation technology to perform clustering analysis on the above audio data to obtain a segment data set grouped by speakers. Among them, the above segment data set includes multiple segment data records of segment start and end times and speaker numbers;

[0165] Step S204, the third processing step: Use a time alignment algorithm to fuse the above transcription text set and the above segment data set to obtain the final recognition result of the above audio data.

[0166] The device in this article can be a server, a PC, a PAD, a mobile phone, etc.

[0167] The present application also provides a computer program product, which is suitable for executing a program initialized with at least the following method steps when executed on a data processing device:

[0168] Step S201, the acquisition step, to acquire audio data;

[0169] Step S202, the first processing step, using voice activity detection technology to calibrate the start and end times of each voice in the above audio data to obtain an audio calibration result, and then using automatic speech recognition technology to transcribe the above audio calibration result to obtain a transcription text set corresponding to the above audio data. Among them, the above transcription text set includes multiple audio text segments, and the above audio text segments are marked with start and end times;

[0170] Step S203, the second processing step, using speaker separation technology to perform clustering analysis on the above audio data to obtain a segment data set grouped by speakers. Among them, the above segment data set includes multiple segment data records of the start and end times and speaker numbers;

[0171] Step S204, the third processing step, using a time alignment algorithm to fuse the above transcription text set and the above segment data set to obtain the final recognition result of the above audio data.

[0172] Obviously, those skilled in the art should understand that the above modules or steps of the present invention can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. They can be implemented by program codes executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order than here, or they can be separately made into individual integrated circuit modules, or multiple modules or steps among them can be made into a single integrated circuit module to implement. In this way, the present invention is not limited to any specific combination of hardware and software.

[0173] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.

[0174] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to generate a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices produce means for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.

[0175] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufacture including instruction means that implement the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.

[0176] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.

[0177] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0178] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer-readable media.

[0179] A computer-readable medium includes both permanent and non-permanent, removable and non-removable media and can implement information storage by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, tape disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information that can be accessed by a computing device. As defined herein, a computer-readable medium does not include transitory computer-readable media such as modulated data signals and carrier waves.

[0180] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but also other elements not expressly listed, or elements that are inherent to such process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising an..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that comprises the element.

[0181] From the above description, it can be seen that the above embodiments of the present application achieve the following technical effects:

[0182] 1), A speech recognition method for a multi-speaker environment of the present application includes: an acquisition step of acquiring audio data; a first processing step of using a voice activity detection technique to calibrate the start and end times of each speech in the audio data to obtain an audio calibration result, and then using an automatic speech recognition technique to transcribe the audio calibration result to obtain a transcription text set corresponding to the audio data, where the transcription text set includes multiple audio text segments, and the audio text segments are marked with start and end times; a second processing step of using a speaker separation technique to perform clustering analysis on the audio data to obtain a segment data set grouped by speakers, where the segment data set includes multiple segment data records of the start and end times and speaker numbers; a third processing step of using a time alignment algorithm to fuse the transcription text set and the segment data set to obtain a final recognition result of the audio data. By combining the use of voice activity detection, automatic speech recognition, speaker separation, and time alignment algorithms, the accuracy and efficiency of recognition are significantly improved. The voice activity detection technique ensures the precise calibration of speech segments, the automatic speech recognition technique provides fast text transcription, the speaker separation technique realizes the separation of speech of different speakers, and the time alignment algorithm ensures the accuracy of the correspondence between the transcription text and the speaker. This method is particularly suitable for scenarios such as meeting records, court records, and large-scale multi-person discussions, and can effectively distinguish the speeches of multiple speakers, improving the practical value and user experience of speech recognition.

[0183] 2) A speech recognition device for a multi-speaker environment according to the present application includes: a first acquisition unit configured to perform an acquisition step to acquire audio data; a first processing unit configured to perform a first processing step, using voice activity detection technology to calibrate the start and end times of each speech in the audio data to obtain an audio calibration result, and then using automatic speech recognition technology to transcribe the audio calibration result to obtain a transcription text set corresponding to the audio data, wherein the transcription text set includes multiple audio text segments, and the audio text segments are marked with start and end times; a second processing unit configured to perform a second processing step, using speaker separation technology to perform clustering analysis on the audio data to obtain a segment data set grouped by speakers, wherein the segment data set includes multiple segment data records of the start and end times and speaker numbers; a third processing unit configured to perform a third processing step, using a time alignment algorithm to fuse the transcription text set and the segment data set to obtain a final recognition result of the audio data. By combining the use of voice activity detection, automatic speech recognition, speaker separation, and time alignment algorithms, the accuracy and efficiency of recognition are significantly improved. The voice activity detection technology ensures the accurate calibration of speech segments, the automatic speech recognition technology provides fast text transcription, the speaker separation technology realizes the separation of speech of different speakers, and the time alignment algorithm ensures the accuracy of the correspondence between the transcribed text and the speaker. This method is particularly applicable to scenarios such as meeting records, court records, and large-scale multi-person discussions, and can effectively distinguish the speeches of multiple speakers, improving the practical value and user experience of speech recognition.

[0184] The foregoing are only the preferred embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various modifications and changes can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.

Claims

1. A speech recognition method for a multi-speaker environment, characterized in that: include: Acquisition step, acquiring audio data; The first processing step is to use voice activity detection technology to calibrate the start and end time of each voice in the audio data to obtain an audio calibration result, and then use automatic speech recognition technology to transcribe the audio calibration result to obtain a transcribed text set corresponding to the audio data, wherein the transcribed text set includes multiple audio text segments, and the audio text segments are marked with start and end times; The second processing step is to perform cluster analysis on the audio data using speaker separation technology to obtain a segment data set grouped by speaker, wherein the segment data set includes multiple segment data recording segment start and end times and speaker numbers; The third processing step is to use a time alignment algorithm to fuse the transcribed text set and the segment data set to obtain a final recognition result of the audio data.

2. The method according to claim 1, characterized in that The transcribed text set and the segment data set are fused using a time alignment algorithm to obtain a final recognition result of the audio data, including: Traversal step: obtaining the audio text segment in the transcribed text set, traversing the segment data set, and determining the segment data corresponding to the audio text segment based on the start and end time of the audio text segment; Numbering processing step: numbering the audio text segment according to the speaker number of the segment data to obtain a speech recognition segment; The traversal step and the numbering step are performed multiple times until all the audio text segments in the transcribed text set are traversed to obtain the final recognition result.

3. The method according to claim 2, characterized in that Determining segment data corresponding to the audio text segment includes: Determine at least one of the segment data overlapping with the audio text segment time according to the start and end time of the audio text segment and the start and end time of the segment data; Determine whether the number of the segment data overlapping with the audio text segment time is 1; When the number of the segment data is 1, it is determined that the segment data corresponds to the audio text segment.

4. The method according to claim 2, characterized in that: Determining the segment data corresponding to the audio text segment also includes: When it is determined that the number of the segment data overlapping with the audio text segment time is greater than 1, calculating the weight of each segment data overlapping with the audio text segment time according to the overlapping duration of the segment data and each audio text segment; The segment data with the largest weight is determined as the segment data corresponding to the audio text segment.

5. The method according to claim 4, characterized in that Calculating the weight of each of the segment data overlapping with the audio text segment in time according to the overlapping duration of the segment data and each of the audio text segments, including: According to the formula: Determine the weight of each of the fragment data, where ω i is the weight of the segment data, slice_duration is the overlapping duration between the segment data and the audio text segment, and seg_duration is m is the total duration of the audio text segment.

6. The method according to claim 1, characterized in that After obtaining the segment data set grouped by speakers, the method further includes: Determining whether the start and end times of the segment data set overlap the start and end times of the transcribed text set; In the case where the start and end times of the transcribed text set are not covered, the speaker separation technology is continued to be used to obtain segment data.

7. The method according to claim 1, characterized in that The first processing step and the second processing step are performed in parallel.

8. A speech recognition device for a multi-speaker environment, characterized in that: include: A first acquisition unit, used to execute an acquisition step to acquire audio data; A first processing unit is used to perform a first processing step, using a voice activity detection technology to calibrate the start and end time of each voice in the audio data to obtain an audio calibration result, and then using an automatic speech recognition technology to transcribe the audio calibration result to obtain a transcribed text set corresponding to the audio data, wherein the transcribed text set includes multiple audio text segments, and the audio text segments are marked with start and end times; A second processing unit is used to perform a second processing step, using a speaker separation technology to perform cluster analysis on the audio data to obtain a segment data set grouped by speakers, wherein the segment data set includes multiple segment data recording segment start and end times and speaker numbers; The third processing unit is used to execute the third processing step, and adopt a time alignment algorithm to fuse the transcribed text set and the segment data set to obtain a final recognition result of the audio data.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored program, wherein when the program is executed, the device where the computer-readable storage medium is located is controlled to execute the speech recognition method for a multi-speaker environment according to any one of claims 1 to 7.

10. An electronic device, characterized in that: include: One or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the one or more programs include a method for executing the speech recognition method for a multi-speaker environment as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Speaker recognition method and device and electronic equipment

    CN112653902A

  • Evaluation method and device of speaker separation algorithm, electronic equipment and storage medium

    CN113593529A

  • Sentence segmentation method and device, storage medium and electronic equipment

    CN113889113A

  • Speech recognition and audio and video processing method, device and system and storage medium

    CN114360545A

  • Speaker segmentation clustering method and device, storage medium and electronic device

    CN114999465A