Audio content extraction method and system for audio feature recognition
By extracting target speech, dividing signal segments from multi-user conversation audio, and performing user clustering analysis, setting overlap speech recognition coefficients, the problem of low overlap speech recognition accuracy in the prior art is solved, and efficient audio content extraction is achieved.
Patent Information
- Application Number
- CN202510814178.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2045-06-18
AI Technical Summary
The existing audio processing technology is difficult to accurately distinguish and extract independent audio content from overlapping voice signal segments in multi-person conversation scenarios, resulting in low speech recognition accuracy and low processing efficiency.
By extracting target speech from multi-user conversation audio, dividing speech signal segments and extracting speech feature parameters, performing user clustering analysis and similarity evaluation, setting overlapping speech recognition coefficients, and separating overlapping audio content using an integrated learning strategy.
It realizes the accurate extraction and distinction between audio content of each user in multi-user conversation, improves recognition accuracy and efficiency, and effectively solves the recognition error problem caused by overlapping speech.
Smart Images

Figure CN120356461B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of speech recognition, and in particular to an audio content extraction method and system for audio feature recognition. Background Art
[0002] In the field of audio processing and analysis, especially in multi-person conversation scenarios such as phone call recordings, accurately identifying and extracting audio content has always been a challenging task. In two-person conversations, due to the possibility of both parties speaking simultaneously, a phenomenon known as overlapping speech, traditional audio content recognition methods often struggle to accurately distinguish and extract the independent audio content of each speaker. In such cases, the speech features of one speaker may be misclassified as part of the audio content of another speaker, resulting in errors in the recognition results and affecting subsequent applications such as speech analysis, transcription, or emotion recognition.
[0003] From a technical perspective, traditional audio processing techniques typically rely on simple audio energy threshold detection or basic spectrum analysis to distinguish speech from noise. However, these methods struggle when dealing with overlapping speech. They cannot effectively distinguish speech signals from different speakers within the same time period, especially when the speech energy or speech features are similar, resulting in a significant increase in error rates. To address this issue, researchers have begun exploring more complex audio feature recognition and separation technologies. These technologies aim to more accurately identify and separate the speech content of different speakers by deeply analyzing speech feature parameters in audio signals, such as pitch, volume, speaking rate, and the time-frequency characteristics of speech. However, existing technologies still suffer from low recognition accuracy and processing efficiency when dealing with two-person conversations with high overlap or similar speech features. Summary of the Invention
[0004] The present invention aims to solve the technical problem in the prior art that it is difficult to accurately distinguish and extract the independent audio content of different speakers in overlapping speech signal segments, resulting in low speech recognition accuracy and low processing efficiency. It provides an audio content extraction method and system for audio feature recognition to solve the problem.
[0005] The technical solution of the present invention to solve the above technical problems is as follows:
[0006] In a first aspect, the present invention provides an audio content extraction method for audio feature recognition, the method comprising: performing speech extraction in target audio of a conversation between multiple users to obtain a target speech; dividing the target speech into speech signal segments to obtain a speech signal segment sequence, and extracting speech feature parameters of each speech signal segment to construct multiple speech feature parameter vectors; extracting overlapping speech signal segments in the speech signal segment sequence, clustering the multiple speech feature parameter vectors into multiple users to obtain multiple user speech feature parameter sets, and screening to obtain multiple user speech signal segment sets, performing user feature similarity analysis based on the multiple user speech feature parameter sets to obtain user feature similarity; setting overlapping speech recognition coefficients based on the user feature similarity, performing heavy audio content separation and recognition on the overlapping speech signal segments to obtain multiple overlapping audio contents, and performing audio content recognition on other multiple user speech signal segments to obtain multiple user audio contents as audio content extraction results.
[0007] In a second aspect, the present invention provides an audio content extraction system for audio feature recognition, the system comprising: a speech extraction module for performing speech extraction in a target audio of a conversation between multiple users to obtain a target speech; a signal segmentation module for dividing the target speech into speech signal segments to obtain a speech signal segment sequence, and extracting speech feature parameters of each speech signal segment to construct multiple speech feature parameter vectors; a user clustering module for extracting overlapping speech signal segments in the speech signal segment sequence, clustering the multiple speech feature parameter vectors into multiple users to obtain multiple user speech feature parameter sets, and screening to obtain multiple user speech signal segment sets, performing user feature similarity analysis based on the multiple user speech feature parameter sets to obtain user feature similarity; a result extraction module for setting overlapping speech recognition coefficients based on the user feature similarity, performing heavy audio content separation and recognition on the overlapping speech signal segments to obtain multiple overlapping audio contents, and performing audio content recognition on other multiple user speech signal segments to obtain multiple user audio contents as audio content extraction results.
[0008] The beneficial effects of the present invention are: by extracting the target voice from the multi-user conversation audio, dividing the voice signal segments and extracting the voice feature parameters to construct a vector, and then identifying the overlapping voice signal segments and clustering the user voice feature parameters to analyze the user feature similarity, setting the overlapping voice recognition coefficient according to the similarity to separate the overlapping audio content, and finally identifying the other user voice signal segments and integrating the results, the technical effect of accurately extracting and distinguishing the audio content of each user, especially in conversations, is achieved, and the recognition error problem caused by overlapping voice is effectively solved. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] Figure 1A flowchart of an audio content extraction method for audio feature recognition provided by the present invention.
[0010] Figure 2 A structural diagram of an audio content extraction system for audio feature recognition provided by the present invention.
[0011] Description of reference numerals: speech extraction module 11 , signal segmentation module 12 , user clustering module 13 , result extraction module 14 . DETAILED DESCRIPTION
[0012] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making any creative efforts shall fall within the scope of protection of the present invention.
[0013] In the description of the present invention, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of the technical features indicated. Therefore, a feature specified as "first" or "second" may explicitly or implicitly include one or more of the specified features. In the description of the present invention, "plurality" means two or more, unless otherwise specifically defined.
[0014] In the description of the present invention, the term "for example" is used to mean "used as an example, illustration or illustration". Any embodiment of the present invention described as "for example" is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is given to enable any person skilled in the art to implement and use the present invention. In the following description, details are listed for the purpose of explanation. It should be understood that a person of ordinary skill in the art can recognize that the present invention can be implemented without using these specific details. In other examples, well-known structures and processes are not elaborated in detail to avoid obscuring the description of the present invention with unnecessary details. Therefore, the present invention is not intended to be limited to the embodiments shown, but is consistent with the widest scope consistent with the principles and features disclosed herein.
[0015] Example 1:
[0016] like Figure 1 As shown, an embodiment of the present invention provides an audio content extraction method for audio feature recognition, the method comprising:
[0017] S10: Perform voice extraction in the target audio of the conversation between multiple users to obtain the target voice.
[0018] For example, in the target audio processing flow of a multi-user conversation, the first step is voice extraction, which aims to accurately locate and separate the speaking part, that is, the target voice, from the overall audio. This process is achieved by analyzing the energy changes of the audio signal on the time axis. Specifically, each moment of the target audio is traversed, the audio energy value at that moment is calculated, and it is compared with the preset voice audio energy threshold. When the audio energy exceeds or equals the threshold, the time period is determined to be voice activity, and the part containing human voice is extracted; conversely, if the audio energy is lower than the threshold, it is determined to be a noise or silence segment and is excluded. For example, in an audio segment containing a conversation between two people, the system can effectively filter out background noise, breathing sounds or short periods of silence through this mechanism, retaining only the voice content of the actual conversation between the two parties, providing a pure data basis for subsequent analysis.
[0019] S20: Divide the target speech into speech signal segments to obtain a speech signal segment sequence, extract speech feature parameters of each speech signal segment, and construct multiple speech feature parameter vectors.
[0020] Optionally, after extracting the target speech, the speech is segmented to form an ordered sequence of speech signal segments. Speech feature parameters within each segment are then extracted to construct multiple speech feature parameter vectors. Specifically, continuous speech segments are automatically identified and segmented based on time-domain characteristics of speech, such as energy variation and zero-crossing rate, or by combining frequency-domain analysis. Each segment represents an independent speech signal segment, and these segments are arranged in chronological order to form a sequence of speech signal segments. Subsequently, each speech signal segment is analyzed for its inherent speech features, including but not limited to voiceprint features (such as d-vectors and x-vectors, which represent different levels of speech feature representation. d-vectors typically focus on the aggregation of short-term speech features, while x-vectors capture richer speech context through deep learning models), pitch, volume, and speech rate. By extracting these feature parameters, a multi-dimensional feature vector is constructed for each speech signal segment, which comprehensively and concisely summarizes the speech characteristics of the signal segment. For example, in an audio segment containing a complex conversation, the above process can accurately segment the continuous voice stream into multiple independent signal segments, and construct a feature vector containing rich voice feature information for each signal segment, laying the data foundation for subsequent steps such as user clustering and feature similarity analysis.
[0021] S30: extract overlapping speech signal segments within the speech signal segment sequence, cluster the multiple speech feature parameter vectors into multiple users to obtain multiple user speech feature parameter sets, and screen to obtain multiple user speech signal segment sets, perform user feature similarity analysis based on the multiple user speech feature parameter sets to obtain user feature similarity.
[0022] After obtaining the speech signal segment sequence and its corresponding speech feature parameter vectors, the system extracts overlapping speech signal segments and applies multi-user clustering analysis to the extracted multiple speech feature parameter vectors. Specifically, the system first detects abnormal overlap of audio energy within the speech signal segment sequence to identify regions of possible overlapping speech, i.e., overlapping speech signal segments. Subsequently, a clustering algorithm (such as K-means or DBSCAN) is used to assign these vectors to clusters based on similarity metrics (e.g., Euclidean distance or cosine similarity) between the speech feature parameter vectors. Each cluster represents a potential user's speech feature set, thereby obtaining multiple user speech feature parameter sets. The speech signal segment sequence is then further screened based on the parameter sets, and signal segments belonging to the same user are merged to form multiple user speech signal segment sets. Based on this, user feature similarity analysis is performed to quantitatively assess the degree of similarity between user speech features by calculating similarity metrics (e.g., feature mean difference, covariance matrix similarity, etc.) between different user speech feature parameter sets. Generally speaking, the greater the similarity between user features, the closer their speech features are, and the greater the difficulty in identifying overlapping speech. Conversely, a lower similarity indicates significant differences in speech features between users, making it relatively easier to identify overlapping speech. For example, in an audio segment containing a two-person conversation with some overlapping speech, the system can accurately identify the overlapping region through the above process and, based on the clustering results of the speech feature parameter vectors, effectively distinguish the speech signal segments of the two speakers. It also quantitatively evaluates the similarity of their speech features, providing a key basis for the subsequent separation and recognition of overlapping speech content.
[0023] S40: setting overlapping speech recognition coefficients according to the user feature similarity, performing audio content separation and recognition on the overlapping speech signal segments to obtain multiple overlapping audio contents, and performing audio content recognition on multiple other user speech signal segments to obtain multiple user audio contents as audio content extraction results.
[0024] Specifically, after completing the user feature similarity analysis, the overlapping speech recognition coefficient is flexibly set based on the resulting user feature similarity index. This coefficient is designed to quantify the impact of the similarity between different users' voice features on the difficulty of overlapping speech recognition. Specifically, when user feature similarity is high, the system will increase the overlapping speech recognition coefficient accordingly to meet the greater recognition challenge; conversely, when the similarity is low, the coefficient will be reduced to optimize recognition efficiency and accuracy.
[0025] Subsequently, an ensemble learning strategy is employed to integrate multiple overlapping audio content recognition branches based on different algorithms or models (these branches can be built using techniques such as deep learning and pattern recognition and specifically trained to handle overlapping speech). The previously marked overlapping speech signal segments are then separated and recognized. By combining the recognition results from each branch, the system accurately extracts the overlapping audio content belonging to different users within the overlapping portion. Meanwhile, conventional audio content recognition techniques are used to analyze and identify the independent audio content of each user within the non-overlapping speech signal segments. Finally, the recognition results from the overlapping and non-overlapping portions are combined to form a complete audio content extraction result. This result not only contains the independent audio content of each user but also accurately distinguishes the speech contributions of different users in the overlapping region, providing high-quality data support for subsequent applications such as speech analysis and transcription. For example, when processing an audio segment of a two-person conversation with complex overlapping speech, the system's aforementioned process can efficiently and accurately separate and identify the speech content of the two speakers in the overlapping region while preserving the independent audio information in the non-overlapping portion, thereby outputting a detailed and accurate audio content extraction report.
[0026] In a preferred embodiment, speech extraction is performed in the target audio of a conversation between multiple users to obtain the target speech, including: extracting the audio energy at each moment in the target audio of a conversation between multiple users; judging whether the audio energy is greater than or equal to a speech audio energy threshold, if so, it is speech, if not, it is noise or silence, and the target speech is obtained by judgment and screening.
[0027] Preferably, in the process of voice extraction for the target audio of multiple user conversations, the energy performance of the audio signal at each moment is first analyzed. Specifically, the complete timeline of the target audio is traversed, and the energy of the audio signal at each time point is calculated. This process aims to quantify the intensity or activity of the audio signal at that moment. Subsequently, the calculated audio energy value is compared with a pre-set voice audio energy threshold. The pre-set voice audio energy threshold is set based on the difference in energy characteristics between the voice signal and noise and silence, and is used to distinguish between voice activity and non-voice activity. When the audio energy is greater than or equal to this threshold, the audio signal at that moment is determined to be voice activity, that is, it contains the sound of human speech; on the contrary, if the audio energy is lower than the threshold, it is classified as noise or silence, and these parts will be regarded as non-target content and excluded in subsequent processing. Through this series of judgment and screening operations, the part containing only voice activity, that is, the target voice, can be accurately extracted from the original target audio, laying a solid foundation for subsequent voice processing and analysis. For example, in an audio clip containing multiple people talking, the system can effectively filter out background noise, environmental noise, and long periods of silence through this process, retaining only clearly discernible voice content, thereby greatly improving the efficiency and accuracy of subsequent processing.
[0028] In a preferred embodiment, the target speech is divided into speech signal segments to obtain a sequence of speech signal segments, and speech feature parameters of each speech signal segment are extracted to construct a plurality of speech feature parameter vectors, including: dividing the target speech into speech signal segments to obtain a plurality of speech signal segments, arranging them in chronological order to obtain a sequence of speech signal segments; extracting the speech feature parameters in each speech signal segment to obtain a plurality of speech feature parameter sets; and constructing a plurality of speech feature parameter vectors based on the plurality of speech feature parameter sets.
[0029] Specifically, after extracting the target speech, the speech signal is segmented. Specifically, the target speech is automatically segmented based on its temporal characteristics, such as energy fluctuations and fundamental frequency fluctuations, to produce multiple independent speech signal segments. These segments are arranged in the order in which they appear in the original speech, forming an ordered sequence of speech signal segments. Subsequently, for each speech signal segment, the internal speech feature parameters are extracted. These parameters include, but are not limited to, Mel-Frequency Cepstral Coefficients (MFCCs), Linear Predictive Coding Coefficients (LPCCs), fundamental frequency (F0), and speech energy, which can comprehensively describe the characteristics of the speech signal from different perspectives.
[0030] Through this step, a corresponding speech feature parameter set is generated for each speech signal segment. Finally, these speech feature parameter sets are used to construct a multi-dimensional speech feature parameter vector for each speech signal segment. This vector uses each feature parameter as its component. Through vectorization, the system can more efficiently process and analyze these speech feature information. For example, when processing an audio segment containing a conversation between multiple people, the above process can accurately divide the continuous speech stream into multiple independent signal segments, and construct a feature vector containing rich speech feature information for each signal segment, providing strong data support for subsequent tasks such as user clustering and audio content recognition.
[0031] In a preferred embodiment, overlapping speech signal segments in the speech signal segment sequence are extracted, the multiple speech feature parameter vectors are clustered into multiple users to obtain multiple user speech feature parameter sets, and multiple user speech signal segment sets are screened, including: judging whether the audio energy in each speech signal segment in the speech signal segment sequence is greater than or equal to the overlapping audio energy threshold, if so, it is an overlapping speech signal segment, if not, it is an ordinary speech signal segment, extracting and obtaining overlapping speech signal segments and ordinary speech signal segment sequences; obtaining the number of users of the multiple users; clustering the multiple speech feature parameter vectors according to the number of users to obtain multiple user speech feature parameter sets, wherein the distance between each cause feature parameter vector and other speech feature parameter vectors is calculated, and the speech feature parameter vectors with the smallest distance are clustered into one class to obtain the clustering result of the number of users as multiple user speech feature parameter sets; dividing the ordinary speech signal segment sequence according to the multiple user speech feature parameter sets to obtain multiple user speech signal segment sets for multiple users.
[0032] Exemplarily, during the extraction of overlapping speech signal segments within a speech signal segment sequence, each segment is individually evaluated by calculating its internal audio energy and comparing it with a preset overlapping audio energy threshold to determine whether it represents overlapping speech. If the audio energy is greater than or equal to the threshold, the segment is considered an overlapping speech signal segment, potentially containing speech from multiple users. Otherwise, the segment is considered a normal speech signal segment, containing only the speech of a single user. This step accurately divides the speech signal segment sequence into overlapping speech signal segments and normal speech signal segments.
[0033] Next, the number of users involved in the current conversation scenario is obtained. This information is crucial for subsequent clustering analysis. Next, the extracted speech feature parameter vectors are clustered according to the number of users. During the clustering process, the distance (e.g., Euclidean distance, cosine similarity, etc.) between each speech feature parameter vector and other vectors is calculated. Vectors with the smallest distance are grouped together, resulting in clustering results that match the number of users, namely, multiple sets of user speech feature parameters.
[0034] Finally, the divided ordinary voice signal segment sequence is further divided according to the user voice feature parameter set. Specifically, each signal segment in the ordinary voice signal segment sequence is assigned to the corresponding user category according to the characteristics of each user voice feature parameter set, and finally a plurality of user voice signal segment sets of multiple users are formed. For example, in an audio segment containing a conversation between three people and with some overlapping voices, the system can accurately identify the overlapping voice signal segments through the above process, and cluster the voice feature parameter vectors according to the number of users, and then reasonably assign the ordinary voice signal segments to each user category, providing a clear data structure for subsequent audio content recognition and analysis.
[0035] In a preferred embodiment, user feature similarity analysis is performed based on the multiple user voice feature parameter sets to obtain user feature similarity, including: respectively calculating the mean of the user voice feature parameters in the multiple user voice feature parameter sets to obtain multiple average user voice feature parameters, and constructing multiple average user voice feature parameter vectors; performing user feature deviation analysis and calculation based on the multiple average user voice feature parameter vectors to obtain user feature deviation degree; and calculating user feature similarity based on the user feature deviation degree.
[0036] Furthermore, after obtaining multiple user voice feature parameter sets, the mean of the voice feature parameters within each set is calculated. The mean calculation aims to comprehensively reflect the overall level of each user's voice features. By taking the mean, multiple average user voice feature parameters are obtained, which form a simplified and representative description of each user's voice features. Subsequently, multiple average user voice feature parameter vectors are constructed using the average user voice feature parameters. Each vector uniquely corresponds to a user and contains average information about that user's voice features. Next, user feature deviation analysis is performed based on these average user voice feature parameter vectors. Specifically, the degree of difference between different user vectors, known as the user feature deviation, is calculated. This metric quantifies the degree of deviation between different user voice features. A greater deviation indicates more significant differences between different user voice features; conversely, a lower deviation indicates closer user voice features. Finally, the user feature similarity is derived based on the calculated user feature deviation. Similarity and deviation are inversely proportional: smaller deviations correspond to higher similarity, indicating greater similarity between different user voice features. For example, in an audio analysis of a three-person conversation, the system calculates the mean of the three user voice feature parameter sets to construct three average user voice feature parameter vectors. Based on these vectors, the user feature deviation is calculated, and the user feature similarity is derived. This result provides an important reference for subsequent overlapping speech recognition and audio content separation, helping the system to more accurately identify and distinguish the speech content of different users.
[0037] In a preferred embodiment, an overlapping speech recognition coefficient is set according to the user feature similarity, and the overlapping speech signal segments are subjected to overlapping audio content separation and recognition to obtain multiple overlapping audio contents, including: extracting data based on historical audio content, collecting the maximum value of user feature similarity within a historical time period, and obtaining a maximum user feature similarity; calculating the ratio of the user feature similarity to the maximum user feature similarity as the overlapping speech recognition coefficient; calling a pre-trained overlapping audio content identifier, and randomly selecting an overlapping audio content recognition branch with a proportion of the overlapping speech recognition coefficient within the overlapping audio content identifier, wherein the overlapping audio content identifier includes multiple overlapping audio content recognition branches; inputting multiple average user speech feature parameters and overlapping speech signal segments into the selected overlapping audio content recognition branch to identify and obtain multiple branch overlapping audio content sets, wherein each branch overlapping audio content set includes user overlapping audio content of multiple users; and screening the user overlapping audio content of multiple users with the highest occurrence frequency in the multiple branch overlapping audio content sets to obtain multiple overlapping audio contents.
[0038] Optionally, when processing overlapping speech signal segments for heavy audio content separation and recognition, historical audio content extraction data is referenced to collect and determine the maximum value of user feature similarity over the historical period. This value is defined as the maximum user feature similarity. This step provides a benchmark for subsequent calculations to assess the relative position of the current user feature similarity within the historical data.
[0039] The system then calculates the ratio of the current user's feature similarity to the maximum user's feature similarity. This ratio is used as the overlapping speech recognition coefficient to dynamically adjust the overlapping speech recognition strategy or weight. For example, if the current user's feature similarity is high, a ratio close to 1 may indicate that a more sophisticated recognition strategy is needed to distinguish between different users' voices. Conversely, if the ratio is low, a relatively simple recognition method can be used.
[0040] Next, a pre-trained overlapping audio content identifier is called. This identifier integrates multiple overlapping audio content recognition branches, each trained based on a different algorithm or model to handle overlapping speech scenarios of varying complexity. During the recognition process, the overlapping audio content recognition branches with the calculated overlapping speech recognition coefficient are randomly selected for subsequent processing. For example, four overlapping audio content recognition branches representing 40% of the 10 overlapping audio content recognition branches are selected. Subsequently, multiple average user speech feature parameters and overlapping speech signal segments are input into the selected overlapping audio content recognition branches. These branches use their respective algorithms or models to analyze the input signals and identify overlapping audio content belonging to different users, forming multiple branch overlapping audio content sets. Each set contains the overlapping audio content of multiple users identified by the corresponding branch. Finally, these branch overlapping audio content sets are comprehensively analyzed. By counting the frequency of occurrence of each user's overlapping audio content in each set, the overlapping audio content of multiple users with the highest frequency of occurrence is selected as the final recognition result, namely, the multiple overlapping audio content. For example, in a complex audio recording of a two-person conversation, the system can effectively adjust the recognition strategy based on the similarity of user features through the above process, and utilize the synergy of multiple recognition branches to accurately separate and identify the speech content of the two speakers in the overlapping area.
[0041] In a preferred embodiment, the step of pre-training the overlapping audio content identifier includes: collecting a set of sample overlapping voice signal segments and a plurality of sample user voice feature parameter sets based on historical data of multi-user audio content recognition, and collecting a plurality of sample overlapping audio contents extracted from each sample overlapping voice signal segment, and labeling to obtain a plurality of sample overlapping audio content sets; constructing a plurality of overlapping audio content recognition branches; randomly dividing the set of sample overlapping voice signal segments, the plurality of sample user voice feature parameter sets, and the plurality of sample overlapping audio content sets to obtain a plurality of overlapping audio content recognition training data, and training the plurality of overlapping audio content recognition branches respectively until convergence.
[0042] In detail, when training the overlapping audio content identifier, sample data is first collected based on the historical data of multi-user audio content recognition. Specifically, a set of sample overlapping voice signal segments is screened out from the historical data. These signal segments contain the voice content of multiple users and overlap; at the same time, a set of multiple sample user voice feature parameters corresponding to each sample overlapping voice signal segment is collected. These parameter sets comprehensively describe the feature information of each user's voice in the signal segment. In addition, for each sample overlapping voice signal segment, the multiple sample overlapping audio contents contained therein are further extracted and labeled to form multiple sample overlapping audio content sets. These sets provide accurate label information for subsequent model training.
[0043] After completing the collection and labeling of sample data, multiple overlapping audio content recognition branches are constructed. These branches are designed based on different algorithm architectures or model structures, aiming to capture the characteristic information in overlapping speech signals from different angles or levels to improve the accuracy and robustness of recognition.
[0044] Subsequently, the collected sets of overlapping speech signal segments, multiple sets of user speech feature parameters, and multiple sets of overlapping audio content are randomly divided to generate multiple sets of overlapping audio content recognition training data. This step aims to increase the diversity and generalization ability of the training data and prevent model overfitting.
[0045] Finally, the training data is used to train the multiple overlapping audio content recognition branches separately until each branch reaches convergence. During the training process, the model's parameters and structure are adjusted based on the characteristics of the branches and the training data to optimize its recognition performance. For example, in a historical audio data segment containing a three-person conversation with some overlapping speech, the above process can collect rich sample data and construct multiple effective overlapping audio content recognition branches. After sufficient training, these branches can accurately identify the speech content of different users in the overlapping area.
[0046] In a preferred embodiment, audio content recognition is performed on multiple other user voice signal segments to obtain multiple user audio contents as audio content extraction results, including: performing audio content recognition on ordinary voice signal segments other than the overlapping voice signal segments to obtain multiple user audio contents; and integrating the multiple overlapping audio contents and the multiple user audio contents as the audio content extraction result.
[0047] Specifically, after completing the audio content separation and recognition of overlapping speech signal segments and obtaining multiple overlapping audio contents, the next step is to perform audio content recognition on ordinary speech signal segments other than the overlapping speech signal segments. This step is intended to process speech signal segments generated by a single user that are not covered by overlapping speech. The system will use a specific audio recognition algorithm or model to analyze these ordinary speech signal segments one by one, extract the speech features therein, and match them with pre-built speech models or templates to identify the user audio content corresponding to each signal segment. These user audio contents represent the independent speech parts of each user in the conversation.
[0048] Subsequently, the multiple overlapping audio contents obtained are integrated with the multiple user audio contents currently identified. The integration process involves operations such as sorting, categorizing, and removing redundant information of the audio content to ensure the completeness and accuracy of the final result. After integration, the audio content extraction results obtained will contain the complete speech content of all users, including both the voice content of different users in the overlapping area (reflected by overlapping audio content) and the independent speech of each user in the non-overlapping area (reflected by user audio content). For example, in an audio recording of a multi-person meeting, the speeches of different users in the overlapping voice signal segments are first identified and separated, and then the remaining ordinary voice signal segments are independently identified. Finally, the two parts of content are integrated together to form a complete conference audio content extraction result, which provides comprehensive data support for subsequent applications such as voice transcription and content analysis.
[0049] The embodiment of the present invention provides an audio content extraction method for audio feature recognition, which has at least the following technical effects:
[0050] 1. By extracting the target speech, dividing the speech signal into segments, clustering and analyzing the user speech feature parameters, and setting the overlapping speech recognition coefficient, it is possible to efficiently separate and identify the speech content of each user from the audio of multi-user conversations, significantly improving the accuracy and efficiency of audio content extraction.
[0051] 2. The overlapping speech recognition coefficient is dynamically adjusted using user feature similarity, and an overlapping audio content recognizer containing multiple recognition branches is called to perform content separation. This can flexibly cope with overlapping speech scenarios of varying complexity, effectively improving the robustness and accuracy of overlapping speech recognition.
[0052] 3. By performing audio content recognition on overlapping speech signal segments and ordinary speech signal segments separately and integrating the recognition results, it can provide complete audio content extraction results, which include both the speech content of different users in the overlapping area and the independent speeches of each user in the non-overlapping area, providing comprehensive data support for subsequent speech processing and analysis.
[0053] Example 2:
[0054] like Figure 2 As shown, based on the same inventive concept as the audio content extraction method for audio feature recognition provided in Example 1, an embodiment of the present invention further provides an audio content extraction system for audio feature recognition, the system comprising:
[0055] The speech extraction module 11 is used to extract speech from the target audio of the conversation between multiple users to obtain the target speech.
[0056] The signal segmentation module 12 is used to divide the target speech into speech signal segments, obtain a speech signal segment sequence, extract speech feature parameters of each speech signal segment, and construct multiple speech feature parameter vectors.
[0057] The user clustering module 13 is used to extract overlapping speech signal segments within the speech signal segment sequence, cluster the multiple speech feature parameter vectors into multiple users to obtain multiple user speech feature parameter sets, and screen to obtain multiple user speech signal segment sets. Based on the multiple user speech feature parameter sets, user feature similarity analysis is performed to obtain user feature similarity.
[0058] The result extraction module 14 is used to set the overlapping speech recognition coefficient according to the user feature similarity, perform audio content separation and recognition on the overlapping speech signal segments to obtain multiple overlapping audio contents, and perform audio content recognition on multiple other user speech signal segments to obtain multiple user audio contents as audio content extraction results.
[0059] Furthermore, the speech extraction module 11 is further configured to perform the following steps:
[0060] Extract the audio energy at each moment in the target audio of multiple user conversations; determine whether the audio energy is greater than or equal to the voice audio energy threshold. If so, it is voice; if not, it is noise or silence. The target voice is then filtered out.
[0061] Furthermore, the signal segmentation module 12 is further configured to perform the following steps:
[0062] The target speech is divided into speech signal segments to obtain multiple speech signal segments, which are arranged in chronological order to obtain a speech signal segment sequence; speech feature parameters in each speech signal segment are extracted to obtain multiple speech feature parameter sets; and multiple speech feature parameter vectors are constructed based on the multiple speech feature parameter sets.
[0063] Furthermore, the user clustering module 13 is further configured to perform the following steps:
[0064] Determine whether the audio energy in each speech signal segment in the speech signal segment sequence is greater than or equal to the overlapping audio energy threshold, if so, it is an overlapping speech signal segment, if not, it is an ordinary speech signal segment, and extract and obtain overlapping speech signal segments and ordinary speech signal segment sequences; obtain the number of users of the multiple users; cluster the multiple speech feature parameter vectors according to the number of users to obtain multiple user speech feature parameter sets, wherein the distance between each cause feature parameter vector and other speech feature parameter vectors is calculated, and the speech feature parameter vectors with the smallest distance are clustered into one class to obtain the clustering result of the number of users as multiple user speech feature parameter sets; divide the ordinary speech signal segment sequence according to the multiple user speech feature parameter sets to obtain multiple user speech signal segment sets for multiple users.
[0065] Furthermore, the user clustering module 13 is further configured to perform the following steps:
[0066] Calculate the mean of the user voice feature parameters in the multiple user voice feature parameter sets respectively to obtain multiple average user voice feature parameters, and construct multiple average user voice feature parameter vectors; perform user feature deviation analysis and calculation based on the multiple average user voice feature parameter vectors to obtain user feature deviation degree; and calculate user feature similarity based on the user feature deviation degree.
[0067] Furthermore, the result extraction module 14 is further configured to perform the following steps:
[0068] Extract data based on historical audio content, collect the maximum value of user feature similarity within the historical time to obtain the maximum user feature similarity; calculate the ratio of the user feature similarity to the maximum user feature similarity as the overlapping speech recognition coefficient; call a pre-trained overlapping audio content identifier, and randomly select an overlapping audio content recognition branch with a proportion of the overlapping speech recognition coefficient within the overlapping audio content identifier, wherein the overlapping audio content identifier includes multiple overlapping audio content recognition branches; input multiple average user voice feature parameters and overlapping voice signal segments into the selected overlapping audio content recognition branch to identify and obtain multiple branch overlapping audio content sets, wherein each branch overlapping audio content set includes user overlapping audio content of multiple users; and screen the user overlapping audio content of multiple users with the highest occurrence frequency from the multiple branch overlapping audio content sets to obtain multiple overlapping audio contents.
[0069] Furthermore, the result extraction module 14 is further configured to perform the following steps:
[0070] Based on historical data of multi-user audio content recognition, a set of sample overlapping voice signal segments and a plurality of sample user voice feature parameter sets are collected, and a plurality of sample overlapping audio contents extracted from each sample overlapping voice signal segment are collected and annotated to obtain a plurality of sample overlapping audio content sets; a plurality of overlapping audio content recognition branches are constructed; the set of sample overlapping voice signal segments, the plurality of sample user voice feature parameter sets, and the plurality of sample overlapping audio content sets are randomly divided to obtain a plurality of overlapping audio content recognition training data, and the plurality of overlapping audio content recognition branches are trained respectively until convergence.
[0071] Furthermore, the result extraction module 14 is further configured to perform the following steps:
[0072] Audio content recognition is performed on the common speech signal segments other than the overlapping speech signal segments to obtain a plurality of user audio contents; and the plurality of overlapping audio contents and the plurality of user audio contents are integrated as an audio content extraction result.
[0073] Through the detailed description of the audio content extraction method for audio feature recognition in the foregoing specification, those skilled in the art can clearly understand the audio content extraction system for audio feature recognition in this embodiment. For the system disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant details can be referred to the method description.
[0074] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present application. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application is not limited to the embodiments shown herein, but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. An audio content extraction method for audio feature recognition, characterized in that: The method comprises: Perform speech extraction in target audio of multiple user conversations to obtain target speech; Dividing the target speech into speech signal segments to obtain a speech signal segment sequence, and extracting speech feature parameters of each speech signal segment to construct multiple speech feature parameter vectors; Extracting overlapping speech signal segments within the speech signal segment sequence, clustering the multiple speech feature parameter vectors into multiple users to obtain multiple user speech feature parameter sets, screening to obtain multiple user speech signal segment sets, and performing user feature similarity analysis based on the multiple user speech feature parameter sets to obtain user feature similarities; The method further comprises: setting an overlapping speech recognition coefficient according to the user feature similarity, performing audio content separation and recognition on the overlapping speech signal segments to obtain a plurality of overlapping audio contents, and performing audio content recognition on the other plurality of user speech signal segments to obtain a plurality of user audio contents as audio content extraction results, including: Extract data based on historical audio content, collect the maximum value of user feature similarity within the historical time, and obtain the maximum user feature similarity; Calculating a ratio of the user feature similarity to the maximum user feature similarity as an overlapping speech recognition coefficient; Calling a pre-trained overlapping audio content identifier, and randomly selecting an overlapping audio content recognition branch with a proportion of the overlapping speech recognition coefficient in the overlapping audio content identifier, wherein the overlapping audio content identifier includes a plurality of overlapping audio content recognition branches; Inputting a plurality of average user voice feature parameters and overlapping voice signal segments into the selected overlapping audio content recognition branch to identify and obtain a plurality of branch overlapping audio content sets, wherein each branch overlapping audio content set includes user overlapping audio content of a plurality of users; In the multiple branch overlapping audio content sets, filtering the user overlapping audio contents of multiple users with the highest appearance frequency to obtain multiple overlapping audio contents; The other multiple user voice signal segments are common voice signal segments other than the overlapping voice signal segments.
2. The audio content extraction method for audio feature recognition according to claim 1, characterized in that: Perform speech extraction on the target audio of multiple user conversations to obtain the target speech, including: Extract the audio energy at each moment in the target audio of multiple user conversations; Determine whether the audio energy is greater than or equal to the speech audio energy threshold. If so, it is speech; if not, it is noise or silence. The target speech is obtained by screening.
3. The audio content extraction method for audio feature recognition according to claim 1, characterized in that: Dividing the target speech into speech signal segments to obtain a speech signal segment sequence, extracting speech feature parameters of each speech signal segment, and constructing multiple speech feature parameter vectors, including: Dividing the target speech into speech signal segments to obtain a plurality of speech signal segments, and arranging the segments in chronological order to obtain a speech signal segment sequence; Extracting speech feature parameters in each speech signal segment to obtain multiple speech feature parameter sets; A plurality of speech feature parameter vectors are constructed according to the plurality of speech feature parameter sets.
4. The audio content extraction method for audio feature recognition according to claim 1, characterized in that: Extracting overlapping speech signal segments within the speech signal segment sequence, clustering the multiple speech feature parameter vectors into multiple users to obtain multiple user speech feature parameter sets, and screening to obtain multiple user speech signal segment sets, including: Determine whether the audio energy in each speech signal segment in the speech signal segment sequence is greater than or equal to an overlapping audio energy threshold; if so, it is an overlapping speech signal segment; if not, it is a normal speech signal segment; extract and obtain a sequence of overlapping speech signal segments and normal speech signal segments; Obtain the number of users of the plurality of users; Clustering the multiple voice feature parameter vectors according to the number of users to obtain multiple user voice feature parameter sets, wherein the distance between each cause feature parameter vector and other voice feature parameter vectors is calculated, and the voice feature parameter vectors with the smallest distance are clustered into one class, thereby obtaining clustering results based on the number of users as the multiple user voice feature parameter sets; The ordinary voice signal segment sequence is divided according to the multiple user voice feature parameter sets to obtain multiple user voice signal segment sets for multiple users.
5. The audio content extraction method for audio feature recognition according to claim 1, characterized in that: Performing user feature similarity analysis based on the multiple user voice feature parameter sets to obtain user feature similarity includes: Calculating the mean values of the user voice feature parameters in the multiple user voice feature parameter sets respectively to obtain multiple average user voice feature parameters, and constructing multiple average user voice feature parameter vectors; Performing user feature deviation analysis and calculation based on the multiple average user voice feature parameter vectors to obtain a user feature deviation degree; The user feature similarity is calculated based on the user feature deviation.
6. The audio content extraction method for audio feature recognition according to claim 1, characterized in that: The step of pre-training the overlapping audio content identifier comprises: Based on historical data of multi-user audio content recognition, a set of sample overlapping voice signal segments and a set of multiple sample user voice feature parameters are collected, and multiple sample overlapping audio contents extracted from each sample overlapping voice signal segment are collected and annotated to obtain a set of multiple sample overlapping audio contents; Construct multiple overlapping audio content recognition branches; The sample overlapping speech signal segment set, the multiple sample user speech feature parameter sets and the multiple sample overlapping audio content sets are randomly divided to obtain multiple overlapping audio content recognition training data, and the multiple overlapping audio content recognition branches are trained respectively until convergence.
7. The audio content extraction method for audio feature recognition according to claim 1, characterized in that: Performing audio content recognition on multiple other user voice signal segments to obtain multiple user audio contents as audio content extraction results includes: Performing audio content recognition on the common voice signal segments other than the overlapping voice signal segments to obtain multiple user audio contents; The multiple overlapping audio contents and the multiple user audio contents are integrated as an audio content extraction result.
8. An audio content extraction system for audio feature recognition, characterized in that: An audio content extraction method for implementing an audio feature recognition method according to any one of claims 1 to 7, the system comprising: A speech extraction module is used to extract speech from target audio of multiple user conversations to obtain target speech; A signal segmentation module is used to divide the target speech into speech signal segments, obtain a speech signal segment sequence, extract speech feature parameters of each speech signal segment, and construct multiple speech feature parameter vectors; a user clustering module configured to extract overlapping speech signal segments within the speech signal segment sequence, cluster the multiple speech feature parameter vectors into multiple users to obtain multiple user speech feature parameter sets, screen the multiple user speech signal segment sets to obtain multiple user speech feature parameter sets, and perform user feature similarity analysis based on the multiple user speech feature parameter sets to obtain user feature similarities; A result extraction module is configured to set overlapping speech recognition coefficients based on the user feature similarity, perform audio content separation and recognition on the overlapping speech signal segments to obtain multiple overlapping audio contents, and perform audio content recognition on multiple other user speech signal segments to obtain multiple user audio contents as audio content extraction results, including: Extract data based on historical audio content, collect the maximum value of user feature similarity within the historical time, and obtain the maximum user feature similarity; Calculating a ratio of the user feature similarity to the maximum user feature similarity as an overlapping speech recognition coefficient; Calling a pre-trained overlapping audio content identifier, and randomly selecting an overlapping audio content recognition branch with a proportion of the overlapping speech recognition coefficient in the overlapping audio content identifier, wherein the overlapping audio content identifier includes a plurality of overlapping audio content recognition branches; Inputting a plurality of average user voice feature parameters and overlapping voice signal segments into the selected overlapping audio content recognition branch to identify and obtain a plurality of branch overlapping audio content sets, wherein each branch overlapping audio content set includes user overlapping audio content of a plurality of users; In the multiple branch overlapping audio content sets, filtering the user overlapping audio contents of multiple users with the highest appearance frequency to obtain multiple overlapping audio contents; The other multiple user voice signal segments are common voice signal segments other than the overlapping voice signal segments.
Citation Information
Patent Citations
Speech recognition method and electronic equipment
CN120319249A