Audio content extraction method and system for audio feature recognition

By extracting the target voice from multi-user conversation audio, dividing the voice signal segments and performing cluster analysis, and setting overlapping speech recognition coefficients, the problem of overlapping speech recognition error in multi-person conversations is solved, and efficient and accurate audio content extraction is achieved.

CN120356461AActive Publication Date: 2025-07-22ZHEJIANG HIPI NETWORK TECH CO LTD
View PDF 17 Cites 0 Cited by

Patent Information

Application Number
CN202510814178.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-07-22
Estimated Expiration
2045-06-18

AI Technical Summary

Technical Problem

The prior art is difficult to accurately distinguish and extract audio content of different speakers in overlapping voice signal segments in multi-person conversation scenarios, resulting in low speech recognition accuracy and low processing efficiency.

Method used

By extracting the target voice from multi-user conversation audio, dividing the voice signal segments and extracting speech feature parameters, clustering and analyzing the user feature similarity, setting overlapping speech recognition coefficients to separate overlapping audio content, and separating audio content using integrated learning strategies and multiple recognition branches.

Benefits of technology

It realizes the accurate extraction and distinction between audio content of each user in multi-user conversation, improves recognition accuracy and efficiency, and effectively solves the recognition error problem caused by overlapping speech.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356461A_ABST
    Figure CN120356461A_ABST
Patent Text Reader

Abstract

The invention relates to an audio content extraction method and system for audio feature recognition, and relates to the field of speech recognition. Target speech is extracted from multi-user dialogue audio, speech signal segments are divided, and speech feature parameters are extracted to construct vectors; recognizing overlapped voice signal segments, clustering user voice feature parameters to analyze user feature similarity, setting overlapped voice recognition coefficients according to the similarity to separate overlapped audio contents, and finally recognizing other user voice signal segments and integrating results. The technical effect of accurately extracting and distinguishing the audio content of each user in multiple users, especially dialogues, is achieved, and the problem of recognition errors caused by overlapped voices is effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of speech recognition, and particularly to an audio content extraction method and system for audio feature recognition. Background Art

[0002] In the field of audio processing and analysis, especially in multi-person conversation scenarios such as telephone call records, the accurate recognition and extraction of audio content have always been a challenging task. In a two-person conversation, due to the possible situation where both parties speak simultaneously, that is, the overlapping speech phenomenon, traditional audio content recognition methods often have difficulty accurately distinguishing and extracting the independent audio content of each speaker. In this case, the speech features of one speaker may be misclassified into the audio content of another speaker, resulting in errors in the recognition results and affecting subsequent applications such as speech analysis, transcription, or emotion recognition.

[0003] From a technical perspective, traditional audio processing techniques usually rely on simple audio energy threshold detection or basic spectrum analysis to distinguish speech from noise. However, when dealing with overlapping speech, these methods are inadequate. They cannot effectively distinguish the speech signals of different speakers within the same time period, especially when the speech energies are similar or the speech features are similar, and the error rate will increase significantly. To solve this problem, researchers have begun to explore more complex audio feature recognition and separation techniques, which aim to more accurately identify and separate the speech content of different speakers by deeply analyzing the speech feature parameters in the audio signal, such as pitch, volume, speech rate, and the time-frequency characteristics of speech. However, existing technologies still have problems such as low recognition accuracy and low processing efficiency when dealing with highly overlapping or speech-feature-similar two-person conversations. Summary of the Invention

[0004] Aiming at the technical problems in the prior art that it is difficult to accurately distinguish and extract the independent audio content of different speakers within the overlapping speech signal segment, resulting in low speech recognition accuracy and low processing efficiency, the present invention provides an audio content extraction method and system for audio feature recognition to solve this problem.

[0005] The technical solution of the present invention to solve the above technical problems is as follows: In a first aspect, the present invention provides an audio content extraction method for audio feature recognition, the method comprising: performing speech extraction within the target audio of multiple user conversations to obtain target speech; dividing the target speech into speech signal segments to obtain a speech signal segment sequence, and extracting the speech feature parameters of each speech signal segment to construct multiple speech feature parameter vectors; extracting the overlapping speech signal segments within the speech signal segment sequence, clustering the multiple speech feature parameter vectors for multiple users to obtain multiple user speech feature parameter sets, and screening to obtain multiple user speech signal segment sets, performing user feature similarity analysis based on the multiple user speech feature parameter sets to obtain user feature similarity; setting an overlapping speech recognition coefficient according to the user feature similarity, performing re-audio content separation recognition on the overlapping speech signal segments to obtain multiple overlapping audio contents, and performing audio content recognition on the other multiple user speech signal segments to obtain multiple user audio contents as the audio content extraction result.

[0006] In a second aspect, the present invention provides an audio content extraction system for audio feature recognition, the system comprising: a speech extraction module for performing speech extraction within the target audio of multiple user conversations to obtain target speech; a signal segmentation module for dividing the target speech into speech signal segments to obtain a speech signal segment sequence, and extracting the speech feature parameters of each speech signal segment to construct multiple speech feature parameter vectors; a user clustering module for extracting the overlapping speech signal segments within the speech signal segment sequence, clustering the multiple speech feature parameter vectors for multiple users to obtain multiple user speech feature parameter sets, and screening to obtain multiple user speech signal segment sets, performing user feature similarity analysis based on the multiple user speech feature parameter sets to obtain user feature similarity; a result extraction module for setting an overlapping speech recognition coefficient according to the user feature similarity, performing re-audio content separation recognition on the overlapping speech signal segments to obtain multiple overlapping audio contents, and performing audio content recognition on the other multiple user speech signal segments to obtain multiple user audio contents as the audio content extraction result.

[0007] The beneficial effects of the present invention are as follows: By extracting target speech in multi-user conversation audio, dividing speech signal segments and extracting speech feature parameters to construct vectors, further identifying overlapping speech signal segments and clustering and analyzing user feature similarity of user speech feature parameters, setting an overlapping speech recognition coefficient according to the similarity to separate overlapping audio contents, and finally identifying other user speech signal segments and integrating the results, the technical effect of accurately extracting and distinguishing the audio contents of each user in multi-users, especially in conversations, is achieved, and the problem of recognition errors caused by overlapping speech is effectively solved. Description of the Drawings

[0008] Figure 1Schematic flowchart of an audio content extraction method for audio feature recognition provided by the present invention.

[0009] Figure 2 Schematic structural diagram of an audio content extraction system for audio feature recognition provided by the present invention.

[0010] Explanation of reference numerals: voice extraction module 11, signal segmentation module 12, user clustering module 13, result extraction module 14. Detailed implementation manners

[0011] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts fall within the protection scope of the present invention.

[0012] In the description of the present invention, the terms "first" and "second" are only used for descriptive purposes, and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of the described features. In the description of the present invention, "a plurality" means two or more, unless otherwise specifically defined.

[0013] In the description of the present invention, the term "for example" is used to mean "serving as an example, illustration, or explanation". Any embodiment described as "for example" in the present invention is not necessarily construed as being more preferred or having more advantages than other embodiments. In order to enable any person skilled in the art to implement and use the present invention, the following description is given. In the following description, details are set forth for the purpose of explanation. It should be understood that those skilled in the art can recognize that the present invention can be implemented without using these specific details. In other instances, well-known structures and processes will not be described in detail to avoid unnecessary details from obscuring the description of the present invention. Therefore, the present invention is not intended to be limited to the embodiments shown, but is to be accorded the widest scope consistent with the principles and features disclosed herein.

[0014] Embodiment 1: As Figure 1 shown, the embodiment of the present invention provides an audio content extraction method for audio feature recognition, and the method includes: S10: Perform voice extraction in the target audio of multiple user conversations to obtain target voice.

[0015] Exemplarily, in the target audio processing flow of a multi-user conversation, the first step is to perform speech extraction, which aims to accurately locate and separate the speaking part, that is, the target speech, from the overall audio. This process is achieved by analyzing the energy changes of the audio signal on the time axis. Specifically, traverse each moment of the target audio, calculate the audio energy value at that moment, and compare it with the preset speech audio energy threshold. When the audio energy exceeds or equals the threshold, the time period is determined to be a voice activity, thereby extracting the part containing human speech; conversely, if the audio energy is lower than the threshold, it is determined to be a noise or silence segment and is excluded. For example, in an audio segment containing a conversation between two people, the system can effectively filter out background noise, breathing sounds, or short periods of silence through this mechanism, retaining only the voice content of the actual conversation between the two parties, and providing a pure data basis for subsequent analysis.

[0016] S20: Divide the target speech into speech signal segments to obtain a speech signal segment sequence, extract speech feature parameters of each speech signal segment, and construct a plurality of speech feature parameter vectors.

[0017] Optionally, after the target speech is extracted, the speech is further segmented to form an ordered sequence of speech signal segments, and the speech feature parameters in each signal segment are further extracted to construct multiple speech feature parameter vectors. Specifically, according to the time domain characteristics of the speech, such as energy change, zero crossing rate, etc., or combined with frequency domain analysis methods, continuous speech segments are automatically identified and segmented. Among them, each segment represents an independent speech signal segment, and these signal segments are arranged in time order to form a sequence of speech signal segments. Subsequently, for each speech signal segment, its inherent speech features are deeply analyzed, including but not limited to voiceprint features (such as d-vector, x-vector, which represent different levels of speech feature representation, d-vector usually focuses on the aggregation of short-term speech features, and x-vector captures richer speech context information through deep learning models), pitch, volume, and speech speed. By extracting these feature parameters, a multi-dimensional feature vector can be constructed for each speech signal segment, which comprehensively and concisely summarizes the speech characteristics of the signal segment. For example, in an audio clip containing a complex conversation, the above process can accurately segment the continuous voice stream into multiple independent signal segments, and construct a feature vector containing rich voice feature information for each signal segment, laying a data foundation for subsequent steps such as user clustering and feature similarity analysis.

[0018] S30: Extract the overlapping speech signal segments within the speech signal segment sequence, perform clustering on the multiple speech feature parameter vectors for multiple users to obtain multiple user speech feature parameter sets, and screen to obtain multiple user speech signal segment sets. According to the multiple user speech feature parameter sets, conduct user feature similarity analysis to obtain the user feature similarity.

[0019] Further, after obtaining the speech signal segment sequence and its corresponding speech feature parameter vectors, further perform the operation of extracting overlapping speech signal segments, and implement multi-user clustering analysis for the extracted multiple speech feature parameter vectors. Specifically, first identify the regions where overlapping speech may exist, i.e., overlapping speech signal segments, by detecting the abnormal superposition of audio energy in the speech signal segment sequence. Subsequently, use clustering algorithms (such as K-means, DBSCAN, etc.) to assign these vectors to different clusters according to the similarity metric (such as Euclidean distance, cosine similarity, etc.) between speech feature parameter vectors. Each cluster represents the speech feature set of a potential user, thus obtaining multiple user speech feature parameter sets. After that, further screen the speech signal segment sequence based on the parameter sets, merge the signal segments belonging to the same user, and form multiple user speech signal segment sets. On this basis, carry out user feature similarity analysis, and quantify and evaluate the similarity degree of speech features between users by calculating the similarity metrics (such as feature mean difference, covariance matrix similarity, etc.) between different user speech feature parameter sets. Generally speaking, the greater the user feature similarity, the closer their speech features are, and thus the greater the difficulty faced in the recognition process of the overlapping speech part; on the contrary, if the similarity is small, it means that the speech features between users are significantly different, and the recognition difficulty of overlapping speech is relatively low. For example, in an audio containing a two-person conversation with some overlapping speech, the system can accurately identify the overlapping region through the above process, effectively distinguish the speech signal segments of the two speakers based on the clustering result of the speech feature parameter vectors, and at the same time quantify and evaluate the similarity of their speech features, providing a key basis for the subsequent separation and recognition of overlapping speech content.

[0020] S40: Set the overlapping speech recognition coefficient according to the user feature similarity, perform separation and recognition of the overlapping audio content for the overlapping speech signal segments to obtain multiple overlapping audio contents, and perform audio content recognition on the other multiple user speech signal segments to obtain multiple user audio contents as the audio content extraction result.

[0021] Specifically, after completing the user feature similarity analysis, the overlapping speech recognition coefficient is flexibly set according to the obtained user feature similarity index. This coefficient aims to quantitatively reflect the impact of the proximity of different users' speech features on the difficulty of overlapping speech recognition. Specifically, when the user feature similarity is high, the system will correspondingly increase the overlapping speech recognition coefficient to cope with higher recognition challenges; conversely, if the similarity is low, the coefficient will be decreased to optimize the recognition efficiency and accuracy.

[0022] Subsequently, an ensemble learning strategy is adopted to integrate multiple overlapping audio content recognition branches based on different algorithms or models (these branches can be constructed based on technologies such as deep learning and pattern recognition and are specially trained to handle overlapping speech situations) to separately identify the previously marked overlapping speech signal segments. By synthesizing the recognition results of each branch, multiple overlapping audio contents belonging to different users in the overlapping part can be accurately extracted. At the same time, for the non-overlapping speech signal segments, conventional audio content recognition technologies are used to parse and identify the independent audio content of each user one by one. Finally, the recognition results of the overlapping part and the non-overlapping part are integrated to form a complete audio content extraction result, which not only includes the independent audio content of each user but also precisely distinguishes the speech contributions of different users in the overlapping area, providing high-quality data support for subsequent speech analysis, transcription, and other applications. For example, when processing a two-person conversation audio containing complex overlapping speech, the system can efficiently and accurately separate and identify the speech content of the two speakers in the overlapping area through the above process, while retaining the independent audio information of the non-overlapping part, thus outputting a detailed and accurate audio content extraction report.

[0023] In a preferred embodiment, in the target audio of multiple user conversations, speech extraction is performed to obtain the target speech, including: extracting the audio energy at each moment in the target audio of multiple user conversations; determining whether the audio energy is greater than or equal to the speech audio energy threshold. If so, it is speech; if not, it is noise or silence, and the target speech is obtained through judgment and screening.

[0024] Preferably, during the process of extracting speech from the target audio of multiple user conversations, the energy performance of the audio signal at each moment is first analyzed. Specifically, the complete timeline of the target audio is traversed, and the energy of the audio signal at each time point is calculated. This process aims to quantify the intensity or activity of the audio signal at that moment. Subsequently, the calculated audio energy value is compared with a pre-set speech audio energy threshold. The pre-set speech audio energy threshold is set based on the differences in energy characteristics between speech signals, noise, and silence, and is used to distinguish speech activity from non-speech activity. When the audio energy is greater than or equal to this threshold, the audio signal at that moment is determined to be speech activity, that is, it contains the sound of human speech; on the contrary, if the audio energy is lower than the threshold, it is classified as noise or silence, and these parts will be regarded as non-target content and excluded in subsequent processing. Through this series of judgment and screening operations, the part containing only speech activity, that is, the target speech, can be accurately extracted from the original target audio, laying a solid foundation for subsequent speech processing and analysis work. For example, in an audio containing multiple people's conversations, the system can effectively filter out background noise, environmental noise, and long silent periods through this process, and only retain the clearly distinguishable speech content, thus greatly improving the efficiency and accuracy of subsequent processing.

[0025] In a preferred embodiment, the target speech is segmented into speech signal segments to obtain a sequence of speech signal segments, and the speech feature parameters of each speech signal segment are extracted to construct a plurality of speech feature parameter vectors, including: segmenting the target speech into a plurality of speech signal segments, arranging them in chronological order to obtain a sequence of speech signal segments; extracting the speech feature parameters within each speech signal segment to obtain a plurality of speech feature parameter sets; constructing a plurality of speech feature parameter vectors according to the plurality of speech feature parameter sets.

[0026] Specifically, after the extraction of the target speech is completed, the operation of segmenting the speech signal segments is carried out. Specifically, based on the changing characteristics of the speech signal in the time domain, such as energy fluctuations, fundamental frequency fluctuations, etc., the target speech is automatically segmented to obtain a plurality of independent speech signal segments. These signal segments are arranged in the order of their appearance in the original speech to form an ordered sequence of speech signal segments. Subsequently, for each speech signal segment, the speech feature parameters inside it are extracted. These parameters include but are not limited to Mel Frequency Cepstral Coefficients (MFCCs), Linear Predictive Coding Coefficients (LPCCs), fundamental frequency (F0), and speech energy, etc., which can comprehensively describe the characteristics of the speech signal from different angles.

[0027] Through this step, a corresponding set of speech feature parameters is generated for each speech signal segment. Finally, using these sets of speech feature parameters, a multi-dimensional speech feature parameter vector is constructed for each speech signal segment. This vector takes each feature parameter as its component. Through the vectorization method, the system can process and analyze this speech feature information more efficiently. For example, when processing an audio containing a multi-person conversation, through the above process, the continuous speech stream can be accurately segmented into multiple independent signal segments, and a feature vector containing rich speech feature information can be constructed for each signal segment, providing strong data support for subsequent tasks such as user clustering and audio content recognition.

[0028] In a preferred embodiment, overlapping speech signal segments within the sequence of speech signal segments are extracted, clustering of the multiple speech feature parameter vectors for multiple users is performed to obtain multiple sets of user speech feature parameters, and multiple sets of user speech signal segments are obtained through screening, including: determining whether the audio energy within each speech signal segment in the sequence of speech signal segments is greater than or equal to an overlapping audio energy threshold. If so, it is an overlapping speech signal segment; if not, it is a normal speech signal segment. The overlapping speech signal segments and the sequence of normal speech signal segments are extracted; the number of users of the multiple users is obtained; according to the number of users, clustering of the multiple speech feature parameter vectors is performed to obtain multiple sets of user speech feature parameters. Among them, the distance between each original feature parameter vector and other speech feature parameter vectors is calculated, and the speech feature parameter vectors with the smallest distance are clustered into one category, and the clustering result of the number of users is obtained as multiple sets of user speech feature parameters; according to the multiple sets of user speech feature parameters, the sequence of normal speech signal segments is divided to obtain multiple sets of user speech signal segments for multiple users.

[0029] Exemplarily, during the process of extracting overlapping speech signal segments within the sequence of speech signal segments, each speech signal segment in the sequence of speech signal segments is judged one by one. By calculating its internal audio energy and comparing it with a preset overlapping audio energy threshold to determine whether this signal segment belongs to overlapping speech. If the audio energy is greater than or equal to the threshold, it is determined as an overlapping speech signal segment, which may contain the speech content of multiple users; otherwise, it is determined as a normal speech signal segment, which only contains the speech of a single user. Through this step, the sequence of speech signal segments can be accurately divided into two parts: overlapping speech signal segments and normal speech signal segments.

[0030] Subsequently, obtain the number of users involved in the current conversation scenario, which is crucial for subsequent clustering analysis. Immediately afterwards, perform clustering operations on the multiple extracted voice feature parameter vectors according to the number of users. During the clustering process, calculate the distances (such as Euclidean distance, cosine similarity, etc.) between each voice feature parameter vector and other vectors, and group the vectors with the smallest distance into one category, thereby obtaining a clustering result that matches the number of users, namely multiple user voice feature parameter sets.

[0031] Finally, further divide the segmented ordinary voice signal sequence based on the user voice feature parameter sets. Specifically, allocate each signal segment in the ordinary voice signal sequence to the corresponding user category according to the characteristics of each user voice feature parameter set, and finally form multiple user voice signal segment sets for multiple users. For example, in an audio containing a three-person conversation with partially overlapping voices, the system can accurately identify the overlapping voice signal segments through the above process, cluster the voice feature parameter vectors according to the number of users, and then reasonably allocate the ordinary voice signal segments to each user category, providing a clear data structure for subsequent audio content recognition and analysis.

[0032] In a preferred embodiment, according to the multiple user voice feature parameter sets, perform user feature similarity analysis to obtain user feature similarity, including: respectively calculate the means of the user voice feature parameters within the multiple user voice feature parameter sets to obtain multiple average user voice feature parameters, and construct multiple average user voice feature parameter vectors; according to the multiple average user voice feature parameter vectors, perform user feature deviation analysis calculation to obtain the user feature deviation degree; according to the user feature deviation degree, calculate to obtain the user feature similarity.

[0033] Furthermore, after obtaining multiple sets of user voice feature parameters, calculate the mean of the internal voice feature parameters for each set of user voice feature parameters respectively. The calculation of the mean aims to comprehensively reflect the overall level of each user's voice features. By calculating the mean, multiple average user voice feature parameters can be obtained, and these parameters constitute a simplified and representative description of each user's voice features. Subsequently, use the average user voice feature parameters to construct multiple average user voice feature parameter vectors. Each vector uniquely corresponds to a user and contains the average information of the user's voice features. Next, perform user feature deviation analysis and calculation based on these average user voice feature parameter vectors. Specifically, calculate the degree of difference between different user vectors, that is, the user feature deviation degree. This index is used to quantitatively evaluate the deviation degree between different users' voice features. The larger the deviation degree, the more significant the difference between different users' voice features; conversely, it indicates that the users' voice features are relatively close. Finally, further derive the user feature similarity based on the calculated user feature deviation degree. The similarity is inversely proportional to the deviation degree, that is, the smaller the deviation degree, the higher the similarity, indicating that the similarity between different users' voice features is stronger. For example, in the analysis of an audio containing a three-person conversation, calculate the means of the three sets of user voice feature parameters to construct three average user voice feature parameter vectors, and calculate the user feature deviation degree based on these vectors, and then obtain the user feature similarity. This result provides an important reference basis for subsequent overlapping speech recognition and audio content separation, and helps the system to more accurately identify and distinguish the speech content of different users.

[0034] In a preferred embodiment, set the overlapping speech recognition coefficient according to the user feature similarity, and perform re-audio content separation and recognition on the overlapping speech signal segment to obtain multiple overlapping audio contents, including: extracting data from historical audio content, collecting the maximum value of the user feature similarity within the historical time to obtain the maximum user feature similarity; calculating the ratio of the user feature similarity to the maximum user feature similarity as the overlapping speech recognition coefficient; calling a pre-trained overlapping audio content recognizer, and randomly selecting the overlapping audio content recognition branch with the proportion of the overlapping speech recognition coefficient in the overlapping audio content recognizer, where the overlapping audio content recognizer includes multiple overlapping audio content recognition branches; inputting multiple average user voice feature parameters and the overlapping speech signal segment into the selected overlapping audio content recognition branch to identify and obtain multiple branch overlapping audio content sets, where each branch overlapping audio content set includes the user overlapping audio contents of multiple users; in the multiple branch overlapping audio content sets, screen out the user overlapping audio contents of multiple users with the highest frequency of occurrence to obtain multiple overlapping audio contents.

[0035] Optionally, during the process of processing overlapping speech signal segments for re-audio content separation and recognition, data is extracted with reference to historical audio content, and the maximum value of the user feature similarity within the historical time is collected and determined. This value is defined as the maximum user feature similarity. This step aims to provide a benchmark for subsequent calculations to evaluate the relative position of the current user feature similarity in the historical data.

[0036] Subsequently, calculate the ratio of the current user feature similarity to the maximum user feature similarity. This ratio is used as the overlapping speech recognition coefficient to dynamically adjust the strategy or weight of overlapping speech recognition. For example, if the current user feature similarity is high, the ratio is close to 1, which may mean that a more refined recognition strategy is needed to distinguish the voices of different users; conversely, if the ratio is low, a relatively simple recognition method can be adopted.

[0037] Next, call a pre-trained overlapping audio content recognizer. This recognizer integrates multiple overlapping audio content recognition branches internally, and each branch is trained based on different algorithms or models to handle overlapping speech scenarios of different complexities. During the recognition process, randomly select the overlapping audio content recognition branches accounting for this coefficient according to the calculated overlapping speech recognition coefficient for subsequent processing. For example, select 4 out of 10 overlapping audio content recognition branches, which is 40%. Subsequently, input multiple average user speech feature parameters and the overlapping speech signal segments into the selected overlapping audio content recognition branches. These branches will analyze the input signals using their respective algorithms or models and identify the overlapping audio content belonging to different users, forming multiple branch overlapping audio content sets. Among them, each set contains the user overlapping audio content of multiple users recognized by the corresponding branch. Finally, conduct a comprehensive analysis of these branch overlapping audio content sets. By counting the occurrence frequency of each user's overlapping audio content in each set, screen out the user overlapping audio content of multiple users with the highest occurrence frequency as the final recognition result, that is, multiple overlapping audio contents. For example, in a complex two-person dialogue audio, the system can effectively adjust the recognition strategy according to the user feature similarity through the above process and utilize the collaborative effect of multiple recognition branches to accurately separate and identify the speech content of the two speakers in the overlapping area.

[0038] In a preferred embodiment, the steps of pre-training the overlapping audio content recognizer include: according to the historical data of multi-user audio content recognition, collecting a set of sample overlapping speech signal segments, multiple sets of sample user speech feature parameters, and collecting multiple sample overlapping audio contents extracted from each sample overlapping speech signal segment, and obtaining a set of multiple sample overlapping audio contents through annotation; constructing multiple overlapping audio content recognition branches; randomly partitioning the set of sample overlapping speech signal segments, multiple sets of sample user speech feature parameters, and the set of multiple sample overlapping audio contents to obtain multiple sets of overlapping audio content recognition training data, and training each of the multiple overlapping audio content recognition branches until convergence.

[0039] Specifically, when training the overlapping audio content recognizer, the collection of sample data is first carried out based on the historical data of multi-user audio content recognition. Specifically, a set of sample overlapping speech signal segments are selected from the historical data, and these signal segments contain the speech contents of multiple users and there is an overlapping phenomenon; at the same time, multiple sets of sample user speech feature parameters corresponding to each sample overlapping speech signal segment are collected, and these parameter sets comprehensively describe the feature information of the speech of each user in the signal segment. In addition, for each sample overlapping speech signal segment, multiple sample overlapping audio contents contained therein are further extracted and annotated, so as to form a set of multiple sample overlapping audio contents. These sets provide accurate label information for subsequent model training.

[0040] After the collection and annotation of the sample data are completed, multiple overlapping audio content recognition branches are further constructed. These branches are designed based on different algorithm architectures or model structures, aiming to capture the feature information in the overlapping speech signal from different angles or levels to improve the accuracy and robustness of recognition.

[0041] Subsequently, the set of collected sample overlapping speech signal segments, multiple sets of sample user speech feature parameters, and the set of multiple sample overlapping audio contents are randomly partitioned to generate multiple sets of overlapping audio content recognition training data. The purpose of this step is to increase the diversity and generalization ability of the training data and prevent the model from overfitting.

[0042] Finally, the training data is used to train each of the multiple overlapping audio content recognition branches until each branch reaches a converged state. During the training process, the parameters and structure of the model are adjusted according to the characteristics of the branch and the training data to optimize its recognition performance. For example, in a historical audio data containing a three-person conversation with partially overlapping speech, rich sample data can be collected and multiple effective overlapping audio content recognition branches can be constructed through the above process. After sufficient training, these branches can accurately identify the speech contents of different users in the overlapping area.

[0043] In a preferred embodiment, audio content recognition is performed on multiple other user voice signal segments to obtain multiple user audio contents as the audio content extraction result, including: performing audio content recognition on the normal voice signal segments other than the overlapping voice signal segments to obtain multiple user audio contents; and integrating the multiple overlapping audio contents and the multiple user audio contents as the audio content extraction result.

[0044] Specifically, after completing the audio content separation and recognition of the overlapping voice signal segments and obtaining multiple overlapping audio contents, next, audio content recognition will be performed on the normal voice signal segments other than the overlapping voice signal segments. This step aims to process the voice signal segments generated by a single user that are not covered by the overlapping voices. The system will use specific audio recognition algorithms or models to analyze these normal voice signal segments one by one, extract the voice features therein, and match them with the pre-constructed voice models or templates to identify the user audio content corresponding to each signal segment. These user audio contents represent the independent speaking parts of each user in the conversation.

[0045] Subsequently, the multiple obtained overlapping audio contents are integrated with the multiple currently recognized user audio contents. The integration process involves operations such as sorting, classifying, and removing redundant information of the audio content to ensure the integrity and accuracy of the final result. After integration, the obtained audio content extraction result will contain the complete speaking contents of all users, including both the voice contents of different users in the overlapping area (reflected by the overlapping audio contents) and the independent speeches of each user in the non-overlapping area (reflected by the user audio contents). For example, in an audio recording of a multi-person meeting, first, the different user speeches in the overlapping voice signal segments are identified and separated, then the remaining normal voice signal segments are independently recognized, and finally, these two parts of the content are integrated together to form a complete audio content extraction result of the meeting, providing comprehensive data support for subsequent applications such as speech transcription and content analysis.

[0046] The audio content extraction method for audio feature recognition provided by the embodiment of the present invention has at least the following technical effects: 1. By extracting the target voice, dividing the voice signal segments, clustering and analyzing the user voice feature parameters, and setting the overlapping voice recognition coefficient, the voice contents of each user can be efficiently separated and recognized from the audio of the multi-user conversation, significantly improving the accuracy and efficiency of audio content extraction.

[0047] 2. Dynamically adjusting the overlapping voice recognition coefficient using the user feature similarity and calling the overlapping audio content recognizer with multiple recognition branches for content separation can flexibly handle overlapping voice scenarios of different complexities, effectively improving the robustness and accuracy of overlapping voice recognition.

[0048] 3. By separately performing audio content recognition on the overlapping speech signal segments and ordinary speech signal segments and integrating the recognition results, a complete audio content extraction result can be provided, which includes the speech content of different users in the overlapping area and the independent speeches of each user in the non-overlapping area, providing comprehensive data support for subsequent speech processing and analysis.

[0049] Embodiment 2: As Figure 2 shown, based on the same inventive concept as the audio content extraction method for audio feature recognition provided in Embodiment 1, the present invention embodiment also provides an audio content extraction system for audio feature recognition. The system includes: A speech extraction module 11, configured to perform speech extraction in the target audio of multiple user conversations to obtain target speech.

[0050] A signal segmentation module 12, configured to divide the target speech into speech signal segments to obtain a sequence of speech signal segments, extract the speech feature parameters of each speech signal segment, and construct multiple speech feature parameter vectors.

[0051] A user clustering module 13, configured to extract the overlapping speech signal segments in the sequence of speech signal segments, perform clustering of multiple users on the multiple speech feature parameter vectors to obtain multiple sets of user speech feature parameters, sieve to obtain multiple sets of user speech signal segments, and perform user feature similarity analysis based on the multiple sets of user speech feature parameters to obtain user feature similarity.

[0052] A result extraction module 14, configured to set an overlapping speech recognition coefficient according to the user feature similarity, perform re-audio content separation recognition on the overlapping speech signal segments to obtain multiple overlapping audio contents, and perform audio content recognition on the other multiple user speech signal segments to obtain multiple user audio contents as the audio content extraction result.

[0053] Furthermore, the speech extraction module 11 is further configured to perform the following steps: In the target audio of multiple user conversations, extract the audio energy at each moment; determine whether the audio energy is greater than or equal to the speech audio energy threshold. If so, it is speech; if not, it is noise or silence, and perform screening to obtain the target speech.

[0054] Furthermore, the signal segmentation module 12 is further configured to perform the following steps: Perform speech signal segment division on the target speech to obtain multiple speech signal segments, arrange them in chronological order to obtain a speech signal segment sequence; extract speech feature parameters within each speech signal segment to obtain multiple speech feature parameter sets; construct multiple speech feature parameter vectors according to the multiple speech feature parameter sets.

[0055] Furthermore, the user clustering module 13 is further configured to perform the following steps: Judge whether the audio energy within each speech signal segment in the speech signal segment sequence is greater than or equal to the overlapping audio energy threshold. If so, it is an overlapping speech signal segment; if not, it is a normal speech signal segment. Extract and obtain the overlapping speech signal segments and the normal speech signal segment sequence; obtain the number of users of the multiple users; cluster the multiple speech feature parameter vectors according to the number of users to obtain multiple user speech feature parameter sets. Among them, calculate the distance between each cause feature parameter vector and other speech feature parameter vectors, and cluster the speech feature parameter vectors with the smallest distance into one category to obtain the clustering result of the number of users as multiple user speech feature parameter sets; divide the normal speech signal segment sequence according to the multiple user speech feature parameter sets to obtain multiple user speech signal segment sets of multiple users.

[0056] Furthermore, the user clustering module 13 is further configured to perform the following steps: Calculate the mean values of the user speech feature parameters within the multiple user speech feature parameter sets respectively to obtain multiple average user speech feature parameters, and construct multiple average user speech feature parameter vectors; perform user feature deviation analysis calculation according to the multiple average user speech feature parameter vectors to obtain the user feature deviation degree; calculate the user feature similarity according to the user feature deviation degree.

[0057] Furthermore, the result extraction module 14 is further configured to perform the following steps: Extract data according to historical audio content, collect the maximum value of the user feature similarity within the historical time, and obtain the maximum user feature similarity; calculate the ratio of the user feature similarity to the maximum user feature similarity as the overlapping speech recognition coefficient; call an overlapping audio content recognizer trained in advance, and randomly select an overlapping audio content recognition branch with the proportion of the overlapping speech recognition coefficient in the overlapping audio content recognizer, where the overlapping audio content recognizer includes multiple overlapping audio content recognition branches; input multiple average user speech feature parameters and overlapping speech signal segments into the selected overlapping audio content recognition branch to identify and obtain multiple branch overlapping audio content sets, where each branch overlapping audio content set includes the user overlapping audio content of multiple users; screen the user overlapping audio content of multiple users with the highest frequency of occurrence in the multiple branch overlapping audio content sets to obtain multiple overlapping audio contents.

[0058] Furthermore, the result extraction module 14 is further configured to perform the following steps: According to the historical data of multi-user audio content recognition, collect a set of sample overlapping speech signal segments, multiple sets of sample user speech feature parameters, and collect multiple sample overlapping audio contents extracted from each sample overlapping speech signal segment, and label to obtain multiple sample overlapping audio content sets; construct multiple overlapping audio content recognition branches; randomly divide the set of sample overlapping speech signal segments, multiple sets of sample user speech feature parameters, and multiple sets of sample overlapping audio contents to obtain multiple pieces of overlapping audio content recognition training data, and train each of the multiple overlapping audio content recognition branches until convergence.

[0059] Furthermore, the result extraction module 14 is further configured to perform the following steps: Perform audio content recognition on the ordinary speech signal segments other than the overlapping speech signal segments to obtain multiple user audio contents; integrate the multiple overlapping audio contents and multiple user audio contents as the audio content extraction result.

[0060] Through the foregoing detailed description of an audio content extraction method for audio feature recognition in this specification, those skilled in the art can clearly know an audio content extraction system for audio feature recognition in this embodiment. For the system disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the description in the method part.

[0061] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. An audio content extraction method for audio feature recognition, characterized in that The method includes: Performing speech extraction within the target audio of multiple user conversations to obtain target speech; Dividing the target speech into speech signal segments to obtain a sequence of speech signal segments, extracting the speech feature parameters of each speech signal segment, and constructing multiple speech feature parameter vectors; Extracting the overlapping speech signal segments within the sequence of speech signal segments, clustering the multiple speech feature parameter vectors for multiple users to obtain multiple sets of user speech feature parameters, screening to obtain multiple sets of user speech signal segments, and performing user feature similarity analysis based on the multiple sets of user speech feature parameters to obtain user feature similarity; Setting an overlapping speech recognition coefficient according to the user feature similarity, performing separation and recognition of the overlapping audio content for the overlapping speech signal segments to obtain multiple overlapping audio contents, and performing audio content recognition on the other multiple user speech signal segments to obtain multiple user audio contents as the audio content extraction result.

2. The method for extracting audio content for audio feature recognition according to claim 1, wherein Performing speech extraction within the target audio of multiple user conversations to obtain target speech, including: Extracting the audio energy at each moment within the target audio of multiple user conversations; Judging whether the audio energy is greater than or equal to the speech audio energy threshold. If so, it is speech; if not, it is noise or silence, and screening to obtain the target speech.

3. The audio content extraction method for audio feature recognition according to claim 1, wherein Dividing the target speech into speech signal segments to obtain a sequence of speech signal segments, and extracting the speech feature parameters of each speech signal segment to construct multiple speech feature parameter vectors, including: Dividing the target speech into multiple speech signal segments, arranging them in chronological order to obtain a sequence of speech signal segments; Extracting the speech feature parameters within each speech signal segment to obtain multiple sets of speech feature parameters; Constructing multiple speech feature parameter vectors according to the multiple sets of speech feature parameters.

4. The method for extracting audio content for audio feature recognition according to claim 1, wherein Extracting the overlapping speech signal segments within the sequence of speech signal segments, clustering the multiple speech feature parameter vectors for multiple users to obtain multiple sets of user speech feature parameters, and screening to obtain multiple sets of user speech signal segments, including: Judging whether the audio energy within each speech signal segment in the sequence of speech signal segments is greater than or equal to the overlapping audio energy threshold. If so, it is an overlapping speech signal segment; if not, it is an ordinary speech signal segment, and extracting to obtain the overlapping speech signal segments and the sequence of ordinary speech signal segments; Obtaining the number of users of the multiple users; Clustering the multiple speech feature parameter vectors according to the number of users to obtain multiple sets of user speech feature parameters. Among them, calculating the distance between each cause feature parameter vector and other speech feature parameter vectors, clustering the speech feature parameter vectors with the smallest distance into one category, and obtaining the clustering result of the number of users as multiple sets of user speech feature parameters; Dividing the sequence of ordinary speech signal segments according to the multiple sets of user speech feature parameters to obtain multiple sets of user speech signal segments for multiple users.

5. The audio content extraction method for audio feature recognition according to claim 1, wherein Performing user feature similarity analysis based on the multiple sets of user speech feature parameters to obtain user feature similarity, including: Calculate the mean values of the user voice feature parameters within the multiple sets of user voice feature parameters respectively to obtain multiple average user voice feature parameters, and construct multiple average user voice feature parameter vectors; Perform user feature deviation analysis calculation based on the multiple average user voice feature parameter vectors to obtain the user feature deviation degree; Calculate the user feature similarity based on the user feature deviation degree.

6. The method for extracting audio content by audio feature recognition according to claim 1, characterized in that, Set the overlapping speech recognition coefficient according to the user feature similarity, and perform re-audio content separation and recognition on the overlapping speech signal segments to obtain multiple overlapping audio contents, including: Extract data from the historical audio content, collect the maximum value of the user feature similarity within the historical time to obtain the maximum user feature similarity; Calculate the ratio of the user feature similarity to the maximum user feature similarity as the overlapping speech recognition coefficient; Call the pre-trained overlapping audio content recognizer, and randomly select the overlapping audio content recognition branches accounting for the overlapping speech recognition coefficient within the overlapping audio content recognizer, where the overlapping audio content recognizer includes multiple overlapping audio content recognition branches; Input the multiple average user voice feature parameters and the overlapping speech signal segments into the selected overlapping audio content recognition branches to recognize and obtain multiple branch overlapping audio content sets, where each branch overlapping audio content set includes the user overlapping audio contents of multiple users; In the multiple branch overlapping audio content sets, screen the user overlapping audio contents of multiple users with the highest occurrence frequency to obtain multiple overlapping audio contents.

7. The method for extracting audio content for audio feature recognition according to claim 6, wherein The steps for pre-training the overlapping audio content recognizer include: According to the historical data of multi-user audio content recognition, collect the sample overlapping speech signal segment set, multiple sample user voice feature parameter sets, and collect multiple sample overlapping audio contents extracted from each sample overlapping speech signal segment, and label to obtain multiple sample overlapping audio content sets; Construct multiple overlapping audio content recognition branches; Randomly divide the sample overlapping speech signal segment set, multiple sample user voice feature parameter sets, and multiple sample overlapping audio content sets to obtain multiple sets of overlapping audio content recognition training data, and train the multiple overlapping audio content recognition branches respectively until convergence.

8. The method for extracting audio content for audio feature recognition according to claim 1, characterized in that Perform audio content recognition on other multiple user speech signal segments to obtain multiple user audio contents as the audio content extraction result, including: Perform audio content recognition on the ordinary speech signal segments other than the overlapping speech signal segments to obtain multiple user audio contents; Integrate the multiple overlapping audio contents and multiple user audio contents as the audio content extraction result.

9. An audio content extraction system for audio feature recognition, characterized in that, An audio content extraction method for implementing the audio feature recognition according to any one of claims 1-8, the system includes: A voice extraction module, configured to perform voice extraction in the target audio of multiple user conversations to obtain the target voice; A signal segmentation module, configured to perform speech signal segment division on the target voice to obtain a speech signal segment sequence, and extract the speech feature parameters of each speech signal segment to construct multiple speech feature parameter vectors; The user clustering module is used to extract overlapping speech signal segments within the sequence of speech signal segments, cluster the multiple speech feature parameter vectors for multiple users to obtain multiple sets of user speech feature parameters, and screen to obtain multiple sets of user speech signal segments. According to the multiple sets of user speech feature parameters, user feature similarity analysis is performed to obtain user feature similarity; The result extraction module is used to set an overlapping speech recognition coefficient according to the user feature similarity, perform overlapping audio content separation and recognition on the overlapping speech signal segments to obtain multiple overlapping audio contents, and perform audio content recognition on other multiple user speech signal segments to obtain multiple user audio contents as the audio content extraction result.

Citation Information

Patent Citations

  • Audio file retrieval method and device

    CN107918663A

  • Overlapped speech recognition method, device, computer equipment and storage medium

    CN111145782A

  • End-to-end multi-speaker overlapping speech recognition

    CN115485768A

  • Speech recognition method and device, processor, memory and electronic equipment

    CN118072734A

  • Target speaker speech recognition method based on lightweight prompt fine tuning

    CN118136017A