Voiceprint processing method, and voiceprint processing method and device for live audio

By pairing the to-processed audio segmentation, clustering and similarity, and combining speech speed analysis to determine whether the audio set is singing audio, the problem of low singing voiceprint extraction efficiency in the live broadcast field is solved, and efficient and accurate voiceprint extraction and content management are achieved.

CN120199274APending Publication Date: 2025-06-24GUANGZHOU HUYA TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510166816.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently extract singing voiceprints in the field of live broadcast, resulting in low content management efficiency and difficulty in detecting violations.

Method used

By dividing the to-processed audio into audio segments, clustering to form an audio set, and determining whether the audio set is singing audio through similarity pairing and speech speed analysis, vocalprint processing is performed to extract singing voiceprints.

Benefits of technology

It improves the efficiency and accuracy of singing voiceprint extraction, and enhances the ability to manage live content and detect violations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120199274A_ABST
    Figure CN120199274A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of voice processing, and relates to a voiceprint processing method, and a voiceprint processing method and device for live audio, and the voiceprint processing method comprises the steps: segmenting a to-be-processed audio into a plurality of audio segments; clustering the plurality of audio clips to obtain a plurality of audio sets; performing similarity pairing on all the audio sets to obtain a first audio set and a second audio set which are paired; performing speech speed analysis on the audio clips of the first audio set and the second audio set to obtain a corresponding first speech speed analysis result and a corresponding second speech speed analysis result; and judging whether the first audio set is a singing audio set or the second audio set is a singing audio set according to the first speech speed analysis result and the second speech speed analysis result, and performing voiceprint processing on all segments in the singing audio set to obtain a singing voiceprint result of the voice corresponding to the singing audio set. According to the method and the device, the singing audio is accurately judged, and the singing voiceprint is automatically extracted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of speech processing, and more particularly, to a voiceprint processing method, a voiceprint processing method and apparatus for live audio. Background Art

[0002] In the era when Internet products are prevalent, voiceprint recognition technology has developed. Especially in the live broadcast field, voiceprint recognition technology plays a key role. As a real-time interactive entertainment activity based on the Internet, live broadcast requires a series of supervision measures such as live broadcast monitoring. Based on the singing voiceprint of the anchor, the system can monitor the live content in real time, judge whether the anchor is singing, which helps the platform automatically identify and mark the music performance part in the live broadcast, improve the efficiency of content management, and can also be used for content review to ensure that the live content complies with the platform regulations and prevent illegal acts such as unauthorized music performances. Voiceprint recognition technology is inseparable from the extraction of voiceprints. As an important identification mark, voiceprints have a certain degree of non-repeatability, and the voiceprints generated by different vocal activities are different. When extracting the singing voiceprint during the live broadcast, it also involves judging whether it belongs to the singing voiceprint. Therefore, in order to ensure the efficiency of extracting the singing voiceprint, a complete technical solution is needed to achieve the extraction of the singing voiceprint. Summary of the Invention

[0003] The present invention aims to overcome at least one defect (shortcoming) of the above-mentioned prior art, and provides a voiceprint processing method, a voiceprint processing method and apparatus for live audio, so as to achieve the effect of extracting the singing voiceprint.

[0004] According to the first aspect of the present application, there is provided a voiceprint processing method, including:

[0005] Segmenting the audio to be processed into a plurality of audio segments;

[0006] Clustering the plurality of audio segments to obtain a plurality of audio sets;

[0007] Performing similarity pairing on all the audio sets to obtain a paired first audio set and a second audio set;

[0008] Performing speech rate analysis on the audio segments of the first audio set and the second audio set respectively to obtain corresponding first speech rate analysis results and second speech rate analysis results;

[0009] Judging whether the first audio set is a singing audio set or the second audio set is a singing audio set according to the first speech rate analysis result and the second speech rate analysis result, and performing voiceprint processing on all the audio segments in the singing audio set to obtain a singing voiceprint result corresponding to the voice of the singing audio set.

[0010] The voiceprint processing method first obtains the audio to be processed, splits the audio to be processed into smaller segments of audio, and at the same time removes the silent or noisy segments to obtain audio segments with human voices. Then, the audio segments are clustered to obtain multiple audio sets. The similarity pairing of the multiple audio sets can obtain similar audio sets, which is the basis for judging singing segments. Then, through speech rate analysis, it is further determined whether the music set is a singing segment, so as to extract the singing voiceprint of the audio set judged as singing audio. The efficiency and accuracy of voiceprint extraction are high.

[0011] Optionally, the clustering of the several audio segments to obtain multiple audio sets includes:

[0012] Converting the several audio segments into corresponding several audio vectors;

[0013] Clustering the several audio vectors to obtain multiple audio sets.

[0014] Converting the audio segments into audio vectors for clustering. Audio vectors are conducive to analysis and processing, can reduce the difficulty of data processing, and at the same time vector data can well retain and store the data information contained in the audio segments themselves. Vector data is more easily recognized by clustering algorithms, improving the accuracy of clustering.

[0015] Optionally, the similarity pairing of all audio sets to obtain the paired first audio set and second audio set includes:

[0016] Randomly extracting the same number of audio segments from each of the audio sets;

[0017] Performing voiceprint extraction on the audio segments extracted from each of the audio sets to obtain a set of audio segment voiceprints corresponding to each of the audio sets;

[0018] Calculating the similarity value pairwise between each audio segment voiceprint in a set of audio segment voiceprints and each audio segment voiceprint in another set of audio segment voiceprints;

[0019] Obtaining the maximum value among all the similarity calculation values, and the two audio sets corresponding to the maximum value are used as the paired first audio set and second audio set.

[0020] The similarity pairing pairs the audio segments in different audio sets pairwise, randomly selects the same number of audio segments from the several audio sets, and then performs similarity pairing on the selected audio segments. According to the similarity pairing result, a set of similar audio sets is screened out, improving the accuracy of judging singing segments. When performing similarity pairing, voiceprint or other features can be paired.

[0021] Optionally, the method further includes:

[0022] Concatenating the audio segments in the audio set to obtain a first concatenated audio;

[0023] Performing duration statistics on the first concatenated audio to obtain a duration statistics result;

[0024] Selecting several candidate audio sets from all the audio sets according to the duration statistics result;

[0025] The step of performing similarity pairing on all the audio sets to obtain the paired first audio set and second audio set specifically includes: performing similarity pairing on all the candidate audio sets to obtain the paired first audio set and second audio set.

[0026] Concatenating the audio segments in the audio set can obtain audio segments with relatively continuous duration and relatively complete audio information. Performing duration statistics on the concatenated audio to obtain a duration statistics result, selecting the concatenated audio with appropriate duration according to the duration statistics result, eliminating the concatenated audio with too short duration, and using the audio set corresponding to the appropriate duration statistics result as the candidate audio set. The candidate audio set will continue to be processed for voiceprint. Concatenating the audio is also a means to further process the audio set. Facing relatively complex audio data, the role of concatenating the audio can screen the audio set to reduce data processing and improve the efficiency of voiceprint processing.

[0027] Optionally, the step of respectively performing speech rate analysis on the audio segments of the first audio set and the second audio set to obtain corresponding first and second speech rate analysis results includes:

[0028] Respectively performing speech recognition on the audio segments of the first audio set and the second audio set to obtain a text recognition result and corresponding text timestamp information;

[0029] Performing speech rate analysis according to the text recognition result and the corresponding text timestamp information to obtain corresponding first and second speech rate analysis results.

[0030] Based on the text recognition result and the corresponding text timestamp information, the corresponding sound speed can be obtained. The speech rate is an intuitive audio feature that can be quickly extracted through simple signal processing methods. It can quickly determine whether an audio is a singing segment in a real-time scenario. Compared with complex voiceprint recognition or deep learning models, the judgment based on the speech rate usually only requires simple feature extraction and threshold comparison, with low computational cost and is suitable for running on resource-constrained devices. At the same time, it is not sensitive to background noise. The judgment of the speech rate mainly depends on the rhythm and speed of the audio, rather than the specific timbre or content. The speech rate can be combined with other audio features (such as pitch, timbre, energy) to provide a more comprehensive analysis.

[0031] Optionally, the judging whether the first audio set is a singing audio set or the second audio set is a singing audio set according to the first speech rate analysis result and the second speech rate analysis result includes:

[0032] Preset a speech rate analysis result threshold;

[0033] Judge whether the corresponding audio segment is a singing segment according to the comparison between the speech rate analysis result and the speech rate analysis result threshold;

[0034] Respectively count the number of audio segments judged as singing segments in the first audio set and the second audio set to obtain the first audio segment number and the second audio segment number;

[0035] Obtain the maximum value of the first audio segment number and the second audio segment number, and the audio set corresponding to the maximum value is the singing audio set.

[0036] Using a preset speech rate threshold can quickly determine whether an audio segment is singing content without complex feature extraction or model training, with low computational cost, suitable for real-time scenarios, and able to respond quickly. The speech rate threshold method has strong robustness to background noise and accompaniment music; even in a complex audio environment, by analyzing the overall rhythm and speed of the audio segment, it can still effectively judge whether someone is singing. At the same time, the preset speech rate threshold is easy to implement and deploy, and can complete the judgment through simple signal processing and threshold comparison, without complex deep learning models or a large amount of labeled data, and does not depend on a specific language or music style, and is widely used in different types of music and singing scenarios. In practical applications, the speech rate threshold can be dynamically adjusted according to the needs of different scenarios.

[0037] Optionally, the performing voiceprint processing on all audio segments in the singing audio set to obtain a singing voiceprint result corresponding to the human voice to which the singing audio set belongs includes:

[0038] Stitch all audio segments in the singing audio set to obtain a second stitched audio;

[0039] Extract the voiceprint of the second spliced audio to obtain the singing voiceprint result corresponding to the voices of the singing audio set.

[0040] Performing voiceprint extraction based on the judgment of the audio can ensure that the obtained voiceprint only contains the singing voiceprint. The extraction of the singing voiceprint can provide users with personalized customization services, sound effect optimization and other services in a variety of application scenarios; there are many means to extract the singing voiceprint. Using deep learning technology makes the voiceprint extraction highly robust to environmental noise, device differences and language changes. Even in complex background music or noisy environments, it can accurately extract voiceprint features. Modern voiceprint extraction technology can achieve extremely high recognition accuracy. Voiceprint extraction is a non-contact biometric technology. Users do not need to directly contact the device and can complete the recognition only through the voice. The singing voiceprint extraction process is natural and smooth. Users can complete the collection and recognition of the voiceprint, music creation and analysis without additional operations when singing.

[0041] According to the second aspect of the present application, there is provided a method for processing the voiceprint of live audio, including:

[0042] Obtain multiple live audios of the same anchor in different sessions;

[0043] Use the voiceprint processing method of the first aspect to process each of the multiple live audios respectively to obtain the singing voiceprint result of each live audio;

[0044] Cluster the singing voiceprint results of all the live audios to obtain multiple voiceprint result sets corresponding to the same anchor; select the voiceprint result set with the largest number of clusters from the multiple voiceprint result sets as the singing voiceprint of the anchor.

[0045] The voiceprint processing method of the live audio aims at the live audio of the anchor, processes the live audio stream in real time, quickly extracts the voiceprint features and performs matching, and real-time identifies the identity of the anchor based on the uniqueness of the anchor's voice. The premise for identification is to have a reference voiceprint sample library. The voiceprint processing method of the live audio processes the live audio of the anchor to obtain the voiceprint result and thus automatically establishes a voiceprint sample library.

[0046] According to the third aspect of the present application, there is provided a voiceprint processing device, including:

[0047] A segmentation module for segmenting the audio to be processed into several audio segments;

[0048] A clustering module for clustering the several audio segments to obtain multiple audio sets;

[0049] A pairing module for pairing all audio sets by similarity to obtain a paired first audio set and a second audio set;

[0050] A speech rate analysis module for separately analyzing the speech rates of the audio segments of the first audio set and the second audio set to obtain corresponding first and second speech rate analysis results;

[0051] A judgment processing module for judging whether the first audio set is a singing audio set or the second audio set is a singing audio set according to the first and second speech rate analysis results, and performing voiceprint processing on all segments in the singing audio set to obtain a singing voiceprint result corresponding to the voice of the singing audio set.

[0052] For the voiceprint processing device, each module is an essential part of the whole process, and the modules cooperate with each other. The audio segments to be processed are first segmented into several audio segments by a segmentation model, and then the several audio segments are clustered by a clustering module to obtain multiple audio sets. The pairing module is used to pair the multiple audio sets by similarity to obtain the paired audio sets, and then the speech rate analysis module is used to analyze the speech rate results of the paired audio sets. Based on the speech rate analysis results, the judgment processing module judges whether it is a singing segment, and finally extracts the voiceprint result of the audio set judged as a singing segment to obtain a singing voiceprint result.

[0053] According to the fourth aspect of the present application, there is provided a voiceprint processing device for live audio, including:

[0054] A live audio acquisition module for acquiring multiple live audios of the same anchor in different sessions;

[0055] A live singing voiceprint acquisition module for respectively processing the multiple live audios by the method of the first aspect to obtain the singing voiceprint results of each live audio;

[0056] An anchor voiceprint acquisition module for clustering the singing voiceprint results of all the live audios to obtain multiple voiceprint result sets corresponding to the same anchor; and selecting the voiceprint result set with the largest number of clusters from the multiple voiceprint result sets as the singing voiceprint of the anchor.

[0057] According to the fifth aspect of the present application, there is provided an electronic device, including:

[0058] A memory for storing one or more computer programs;

[0059] A processor, when the one or more computer programs are executed by the processor, implementing the voiceprint processing method described in the first aspect or the voiceprint processing method for live audio described in the second aspect.

[0060] The electronic device provides a complete set of equipment, which facilitates audio processing and extracting singing voiceprint results. The device has a fast processing speed for the audio stream, is accurate and effective, and has certain advantages.

[0061] According to the sixth aspect of the present application, there is provided a computer-readable storage medium storing computer instructions for causing a processor to implement the voiceprint processing method described in the first aspect or the voiceprint processing method for live audio in the second aspect when executed.

[0062] Based on any of the above aspects, a voiceprint processing method, a voiceprint processing method for live audio, a device, an electronic device, and a computer storage medium provided by the embodiments of the present application can process audio voiceprints to obtain singing voiceprint results, automatically establish a voiceprint library, save time and effort, and can extract accurate singing voiceprint results, laying a foundation for voiceprint recognition and other voiceprint applications. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0064] Figure 1 It is a schematic application scenario diagram of a voiceprint processing method or a voiceprint processing method for live audio provided in this embodiment.

[0065] Figure 2 It is a flowchart of a voiceprint processing method provided in this embodiment.

[0066] Figure 3 It is a flowchart of a voiceprint processing method for live audio provided in this embodiment.

[0067] Figure 4 It is a schematic diagram of the functional modules of a voiceprint processing device provided in this embodiment.

[0068] Figure 5 It is a schematic diagram of the functional modules of a voiceprint processing device for live audio provided in this embodiment.

[0069] Figure 6 It is a schematic diagram of the structure of the electronic device provided in this embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0070] The accompanying drawings of this application are only for illustrative purposes and should not be construed as limiting the application. To better illustrate the following embodiments, some components in the drawings may be omitted, enlarged, or reduced, which do not represent the dimensions of the actual product; for those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.

[0071] To enable those skilled in the art to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.

[0072] It should be noted that the terms "first", "second", etc. in the specification, claims, and above-mentioned accompanying drawings of this application are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0073] Currently, when detecting in real time whether a host is singing based on the voiceprint solution, the key lies in the need to pre-store the singing voiceprints of the host in the database in advance. However, the traditional method of voiceprint storage in the database relies on manual search and manual annotation. This method not only consumes time and effort but also is prone to errors, resulting in low storage efficiency. To solve this problem, it is necessary to improve the efficiency of voiceprint collection and storage while reducing labor costs.

[0074] This embodiment provides a technical solution that can solve the above problems. The following will describe the specific implementation manners of this application in detail in conjunction with the accompanying drawings.

[0075] Exemplarily, it is a schematic diagram of the application scenario of a voiceprint processing method provided for the embodiments of this application and the voiceprint processing method of live audio. As Figure 1As shown, the application scenario at least includes a server 100 and a terminal 200 that can communicate with the server 100. The server 100 has the functions of transmitting and receiving audio and video, and also has the functions of analyzing and processing the received audio and video; the terminal device 200 has the functions of transmitting and receiving audio and video, also has the functions of analyzing and processing audio and video, and may also have the function of playing audio and video.

[0076] It is understandable that the server 100 can be an independent electronic device or a cluster composed of multiple electronic devices; the terminal 200 can be a smart phone terminal, a personal computer, a tablet computer, a vehicle-mounted terminal, etc., but is not limited thereto.

[0077] In an implementable manner, the server 100 and the terminal 200 can respectively execute the voiceprint processing method provided by the embodiments of the present application or the voiceprint processing method of live audio. Alternatively, optionally, part of the voiceprint processing method provided by the embodiments of the present application or the voiceprint processing method of live audio is executed in the server 100, and part is executed in the terminal 200.

[0078] As Figure 2 shown, the present embodiment provides a voiceprint processing method, which may include the following steps:

[0079] S110: Split the audio to be processed into several audio segments;

[0080] In this embodiment, the method of splitting the audio to be processed into several audio segments includes audio editing software, using an online audio cutting tool, using a programming language for automated cutting, etc. It can also be splitting based on an energy threshold, splitting based on frequency domain features, splitting based on a statistical model, splitting based on deep learning, etc. In the splitting based on deep learning, it also includes pyannote VAD (Voice Activity Detection): an end-to-end deep learning model based on CNN and RNN, which can output the voice activity probability at each moment; FSMN-Monophone VAD (Voice Activity Detection): based on the FSMN network structure, which can effectively capture long-distance dependencies in the time series and has a lower computational complexity, etc.

[0081] In an alternative implementation, the audio to be processed is obtained by collecting the audio in the video, and is segmented using a VAD tool (Voice Activity Detection) to obtain multiple audio segments with human voices. Then, an overlap intersection is performed on each audio segment with human voices. The overlap intersection is implemented according to the principle of 1.5 seconds per segment and 0.75 seconds between segments, and the audible audio segments are segmented into multiple 1.5-second seg (segment) small segments by using the overlap intersection. The segmentation method includes but is not limited to this.

[0082] S120. Cluster the several audio segments to obtain multiple audio sets;

[0083] In this embodiment, when clustering the several audio segments to obtain multiple audio sets, the clustering method can be partitioning-based clustering, hierarchical clustering, density-based clustering, etc., or model-based clustering, etc. The clustering condition can be the characteristic attributes of the sound or audio features suitable for clustering, such as MelSpectrum, MFCC (Mel-Frequency Cepstral Coefficients), or voiceprint features, etc. A suitable clustering algorithm is selected according to the characteristics of the audio data, such as K-means, hierarchical clustering, spectral clustering, or a deep learning-based clustering method.

[0084] In an alternative implementation, a voiceprint extraction model is used to convert the seg segment audio into vectors, and then these vectors are clustered. A clustering algorithm is used to cluster the converted vectors. The clustering is based on the characteristic attributes of the sound as the clustering condition, so as to cluster the audio segments belonging to the same person or similar ones into an audio set, and each audio segment is marked. Each audio segment has a speaker-id, representing different speakers. All the seg audio segments are clustered to obtain multiple audio sets. The clustering implementation includes but is not limited to this.

[0085] Exemplarily, a certain live broadcast of a certain anchor is obtained. In the audio of the live broadcast, the audio of the anchor speaking is from 0 to 20 minutes, the audio of the anchor singing is from 20 to 30 minutes, and the audio of the connected anchor speaking is from 30 to 40 minutes. The clustering algorithm can divide the live broadcast audio into speaker-id 1 from 0 to 20, speaker-id 2 from 20 to 30, and speaker-id 3 from 30 to 40 in sequence. Then, by methods such as pairing and speech rate determination, it is determined that speaker-id 2 is the voiceprint of the anchor singing.

[0086] S130. Perform similarity pairing on all audio sets to obtain the paired first audio set and second audio set;

[0087] In this embodiment, the similarity pairing can be cosine similarity, or it can be through MFCC (Mel-Frequency Cepstral Coefficients) to describe the timbre characteristics of the audio; Mel Spectrum: effectively describes the spectral characteristics of the audio; RMS energy: represents the average power of the audio signal and is used to distinguish the volume size; Zero-Crossing Rate: reflects the frequency characteristics of the audio signal; Spectral Centroid: describes the frequency distribution of the audio signal, etc.

[0088] In an alternative implementation, splice adjacent seg segments or those belonging to the same speaker-id, that is, they belong to one audio set. Splice the seg segments in the audio set to obtain multiple segments of audio belonging to different speaker-ids. Perform duration statistics on the spliced audio, and select several candidate audio sets according to the statistical duration results. The candidate audio sets are used for similarity pairing, and the spliced audio is not specifically limited.

[0089] In a preferred implementation, use the candidate audio sets for similarity pairing. Randomly select 6 segments from each candidate audio set to extract voiceprints. If there are three candidate audio sets, then there are 18 segment voiceprints. Calculate the cosine similarity between the 18 segment voiceprints pairwise, find the two segment voiceprints with the highest similarity, and respectively consider them as the voiceprints during singing and speaking. Determine the corresponding audio sets according to the segment voiceprints with the highest similarity, and then the audio sets for singing and speaking can be determined. Since it is impossible to distinguish singing audio and speaking audio in this case, it is also necessary to judge the singing audio according to the similarity pairing result. The similarity pairing methods include but are not limited to this.

[0090] S140. Perform speech rate analysis on the audio segments of the first audio set and the second audio set respectively to obtain the corresponding first speech rate analysis result and second speech rate analysis result;

[0091] In this embodiment, the method for performing speech rate analysis can be based on audio analysis tools, Python, or audio feature analysis, or it can also be deep learning, etc.

[0092] In an alternative implementation, for the set of speech audio and the set of singing audio determined according to the similarity matching results, perform Automatic Speech Recognition (ASR) on all audio segments in the set of speech audio and the set of singing audio. Obtain the text and word timestamps of the audio segments through automatic speech recognition, calculate the speech rate of each audio segment based on these two pieces of information, set a preset speech rate threshold, and determine the set of singing audio in the two audio sets obtained after similarity matching according to the speech rate threshold. The judgment method includes but is not limited to...

[0093] S150. According to the first speech rate analysis result and the second speech rate analysis result, judge whether the first audio set is the set of singing audio or the second audio set is the set of singing audio. Perform voiceprint processing on all segments in the set of singing audio to obtain the singing voiceprint result corresponding to the voice of the set of singing audio.

[0094] In this embodiment, the method of judging whether the first audio set is the set of singing audio or the second audio set is the set of singing audio according to the first speech rate analysis result and the second speech rate analysis result can be to set a preset speech rate threshold, which can be set according to actual needs. It can be set by presetting a speech rate threshold that conforms to speech, or by presetting a speech rate threshold that conforms to singing, etc.

[0095] In an alternative implementation, preset a speech rate threshold for judging singing. Compare the first speech rate analysis result and the second speech rate analysis result with the preset speech rate threshold respectively. Specifically, preset a speech rate threshold of 3.6 words per second. If the first speech rate analysis result or the second speech rate analysis result is less than the preset speech rate threshold of 3.6 words per second, then consider this audio segment as a singing segment; otherwise, consider it as a speech segment. When determining whether an audio segment is a singing segment or a speech segment, add a tag with speaker-id as "singing" or "speech". The judgment method of judging singing audio or speech audio according to the speech rate includes but is not limited to this...

[0096] In a more preferred implementation, according to the added tag with speaker-id as "singing" or "speech", determine whether the audio set is singing audio. Specifically, count the number of tags of the audio segments in the audio set after similarity matching respectively. Judge whether it is a singing segment based on the number of tags of the audio segments in the audio set. If the number of "singing" tags in an audio set is greater than the number of "speech" tags, then consider this audio set as singing audio. Only when one of the two paired audio sets is determined to be singing audio and the other is determined to be speech audio, will singing voiceprint extraction be performed and the singing voiceprint be uploaded to the database; otherwise, discard the current video. The judgment and processing method includes but is not limited to...

[0097] Exemplarily, audio segments of an audio set determined to be singing audio are spliced to obtain a spliced audio. Then, the spliced audio is the singing segment extracted from the video. After obtaining a voiceprint vector by performing voiceprint extraction on the spliced audio, it is stored in a temporary vector library together with the identifier to which the video belongs. After processing the audio of multiple videos, the temporary vector library will accumulate singing voiceprint vectors. Finally, clustering is performed on the singing voiceprints, and the vector group with the largest number of clusters is selected and stored in the online machine library.

[0098] As Figure 3 shown, this embodiment also provides a method for processing the voiceprint of live broadcast audio, which may include the following steps:

[0099] S210. Obtain multiple live broadcast audios of the same anchor in different sessions;

[0100] In this embodiment, for the obtaining of multiple live broadcast audios of the same anchor in different sessions, the different sessions may be live broadcast audios with close time periods or different time periods, or a single live broadcast.

[0101] S220. Use the voiceprint processing method of the first aspect to process each of the multiple live broadcast audios respectively to obtain the singing voiceprint result of each live broadcast audio;

[0102] In this embodiment, the audio of the live broadcast video is obtained as the audio to be processed, the audio to be processed is segmented into several audio segments, the several audio segments are clustered to obtain multiple audio sets, similarity pairing is performed on all the audio sets to obtain paired first and second audio sets, the speech rate analysis is respectively performed on the audio segments of the first audio set and the second audio set to obtain corresponding first and second speech rate analysis results, it is determined whether the first audio set is a singing audio set or the second audio set is a singing audio set according to the first speech rate analysis result and the second speech rate analysis result, voiceprint processing is performed on all the segments in the singing audio set to obtain the singing voiceprint result of the person voice to which the singing audio set belongs, and the singing voiceprint result of the live broadcast audio is obtained according to the voiceprint processing method.

[0103] S230. Cluster the singing voiceprint results of all the live broadcast audios to obtain multiple voiceprint result sets corresponding to the same anchor; select the voiceprint result set with the largest number of clusters from the multiple voiceprint result sets as the singing voiceprint of the anchor.

[0104] In this embodiment, the singing voiceprint results of the live audio obtained according to the voiceprint processing method are clustered to obtain a set of voiceprint results. The voiceprint result set with the largest number of clusters is selected from the set of voiceprint results as the singing voiceprint of the host. The live audio is processed according to the voiceprint processing method of the live audio, so as to automatically extract the singing voiceprint results.

[0105] As Figure 4 shown, an embodiment of the present application provides a voiceprint processing device 310. Optionally, the voiceprint processing device 310 may include:

[0106] A segmentation module 311, configured to segment the audio to be processed into a plurality of audio segments;

[0107] In this embodiment, the segmentation module 311 may be used to execute Figure 2 the steps S110 shown. The specific description of the segmentation module 311 may refer to the description of the steps S110.

[0108] A clustering module 312, configured to cluster the plurality of audio segments to obtain a plurality of audio sets;

[0109] In this embodiment, the clustering module 312 may be used to execute Figure 2 the steps S120 shown. The specific description of the clustering module 312 may refer to the description of the steps S120.

[0110] A pairing module 313, configured to perform similarity pairing on all audio sets to obtain a paired first audio set and a second audio set;

[0111] In this embodiment, the pairing module 313 may be used to execute Figure 2 the steps S130 shown. The specific description of the pairing module 313 may refer to the description of the steps S130.

[0112] A speech rate analysis module 314, configured to perform speech rate analysis on the audio segments of the first audio set and the second audio set respectively to obtain corresponding first speech rate analysis results and second speech rate analysis results;

[0113] In this embodiment, the speech rate analysis module 314 may be used to execute Figure 2 the steps S140 shown. The specific description of the speech rate analysis module 314 may refer to the description of the steps S140.

[0114] A judgment processing module 315, configured to determine whether the first audio set is a singing audio set or the second audio set is a singing audio set according to the first speech rate analysis result and the second speech rate analysis result, and perform voiceprint processing on all segments in the singing audio set to obtain a singing voiceprint result corresponding to the voice of the singing audio set.

[0115] In this embodiment, the judgment processing module 315 may be configured to execute Figure 2 the step S150 shown, and the specific description of the judgment processing module 315 may refer to the description of the step S150.

[0116] As Figure 5 shown, an embodiment of the present application further provides a voiceprint processing device 410 for live audio. Optionally, the voiceprint processing device 410 may include:

[0117] A live audio acquisition module 411, configured to acquire multiple live audios of the same anchor in different sessions;

[0118] In this example, the live audio acquisition module 411 may be configured to execute Figure 3 the step S210 shown, and the specific description of the live audio acquisition module 411 may refer to the description of the step S210.

[0119] A live singing voiceprint acquisition module 412, configured to process the multiple live audios respectively by using the method of the first aspect to obtain a singing voiceprint result for each live audio;

[0120] In this example, the live singing voiceprint acquisition module 412 may be configured to execute Figure 3 the step S220 shown, and the specific description of the live singing voiceprint acquisition module 412 may refer to the description of the step S220.

[0121] An anchor voiceprint acquisition module 413, configured to cluster the singing voiceprint results of all the live audios to obtain multiple voiceprint result sets corresponding to the same anchor; select the voiceprint result set with the largest number of clusters from the multiple voiceprint result sets as the singing voiceprint of the anchor.

[0122] In this example, the anchor voiceprint acquisition module 413 may be configured to execute Figure 3 the step S230 shown, and the specific description of the anchor voiceprint acquisition module 413 may refer to the description of the step S230.

[0123] It can be understood that the above device embodiments and the above method embodiments can correspond to each other. Similar descriptions of the device embodiments can refer to the method embodiments. To avoid repetition, they will not be elaborated here. A voiceprint processing device provided by an embodiment of the present application can execute a voiceprint processing method provided by any embodiment of the present application. A voiceprint processing device for live audio provided also can execute a voiceprint processing method for live audio provided by any embodiment of the present application, and has corresponding functional modules and beneficial effects for executing the method. The functional modules of the voiceprint processing device can be implemented in the form of hardware, can be implemented by instructions in the form of software, and can also be implemented by a combination of hardware and software modules.

[0124] Specifically, each step of the method embodiment of the present application can be completed by the integrated logic circuit in the hardware in the processor and / or instructions in the form of software. The steps of the voiceprint processing method in combination with the embodiments of the present application can be directly embodied as being executed and completed by the hardware encoding processor, or executed and completed by a combination of the hardware and software modules in the encoding processor. Optionally, the software module can be located in a random access memory, a read only memory, a programmable read only memory, a flash memory, an electrically erasable programmable memory, a register, and other storage media are all possible. The storage media is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps in the above method embodiments.

[0125] An embodiment of the present application provides an electronic device 510, and its structure is as Figure 6 shown. The electronic device 510 can be the server 100 or the terminal 200 shown in this embodiment Figure 1 shown.

[0126] As Figure 6 shown, the electronic device 510 includes a memory 511, a processor 512, a communication module 513, and an input / output interface 515, etc. Optionally, the memory 511, the processor 512, the communication module 513, and the input / output interface 515 can be connected and communicate through a bus 515.

[0127] The memory 511 is used to store one or more computer programs and transmit the code of the computer programs to the processor 512; when the one or more computer programs are executed by the processor 512, the voiceprint processing method in the embodiments of the present application is implemented.

[0128] Optionally, the electronic device 510 may be connected to a network through the communication module 513 to communicate with other devices, such as terminals or servers, through the network to achieve data interaction. The electronic device 510 may be various forms of digital computers, such as, by way of example, desktop computers, servers, workstations, mainframe computers, or other types of computers. The electronic device 510 may also be various forms of mobile terminals, such as, by way of example, smart phones, tablet computers, wearable devices (such as helmets, glasses, watches, etc.), and other similar mobile terminals.

[0129] Optionally, the electronic device 510 may be connected to the required input / output devices, such as keyboards, display devices, etc., through the input / output interface 515. The electronic device 510 itself may have a display device, and may also externally connect other display devices through the input / output interface 515. Optionally, a storage device, such as a hard disk, etc., may also be connected through the input / output interface 515, so that the data in the electronic device 510 can be stored in the storage device, or the data in the storage device can be read, and the data in the storage device can also be stored in the memory 511. It can be understood that the input / output interface 515 may be a wired interface or a wireless interface. According to different actual application scenarios, the devices connected to the input / output interface 515 may be components of the electronic device 510 or external devices connected to the electronic device 510 when needed.

[0130] Optionally, the memory 511 may be a volatile memory and / or a non-volatile memory. The volatile memory may be a random access memory, etc., and the non-volatile memory may be a read-only memory, a programmable read-only memory, an erasable programmable read-only memory, an electrically erasable programmable read-only memory, or a flash memory, etc.

[0131] Optionally, the computer program stored in the processor 512 may be divided into one or more modules. The one or more modules are stored in the memory 511 and executed by the processor 512 to complete the method provided by the present embodiment itself. The one or more modules may be a series of computer program instruction segments capable of completing specific functions, and the computer program instruction segments are used to describe the execution process of the computer program in the electronic device 510.

[0132] Optionally, the processor 512 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 512 include, but are not limited to, a central processing unit, a graphics processing unit, a digital signal processor, various special artificial intelligence computing chips, various processors running machine learning model algorithms, and can also be any suitable controller, microcontroller, processor, etc. The processor 512 executes the various methods and processes of this embodiment. Exemplarily, it is a voiceprint processing method according to an embodiment of the present application.

[0133] Optionally, the bus 515 can include a path for transmitting information. The bus 515 can be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. According to different functions, the bus 515 can be divided into an address bus, a data bus, a control bus, etc.

[0134] In an alternative implementation, an embodiment of the present application also provides a computer storage medium, on which a computer program is stored. When the computer program is executed by a computer, the computer can execute the methods of the above method embodiments. Part or all of the computer program can be loaded and / or installed on the memory 511 of the electronic device 510. When the computer program is executed by the processor 512, one or more steps of a voiceprint processing method according to an embodiment of the present application can be executed.

[0135] Optionally, the computer-readable storage medium can be a random access memory, a read-only memory, a programmable read-only memory, an erasable programmable read-only memory, an electrically erasable programmable read-only memory, etc.

[0136] Obviously, the above embodiments of the present application are merely examples for clearly illustrating the technical solutions of the present application, rather than limitations on the specific implementation manners of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the claims of the present application shall be included in the protection scope of the claims of the present application.

Claims

1. A voiceprint processing method, characterized in that: The method comprises: Divide the audio to be processed into several audio segments; Clustering the plurality of audio clips to obtain a plurality of audio sets; Perform similarity pairing on all audio sets to obtain a paired first audio set and a second audio set; Performing speech rate analysis on the audio segments of the first audio set and the second audio set respectively to obtain corresponding first speech rate analysis results and second speech rate analysis results; Based on the first speech rate analysis result and the second speech rate analysis result, determine whether the first audio set is a singing audio set or the second audio set is a singing audio set, perform voiceprint processing on all audio clips in the singing audio set, and obtain the singing voiceprint result corresponding to the human voice belonging to the singing audio set.

2. A voiceprint processing method according to claim 1, characterized in that: The step of clustering the plurality of audio clips to obtain a plurality of audio sets includes: Convert the plurality of audio clips into corresponding plurality of audio vectors; The plurality of audio vectors are clustered to obtain a plurality of audio sets.

3. A voiceprint processing method according to claim 1, characterized in that: The method of performing similarity pairing on all audio sets to obtain a paired first audio set and a second audio set includes: Randomly extracting the same number of audio clips from each of the audio sets; Performing voiceprint extraction on the audio segments extracted from each of the audio sets to obtain a group of audio segment voiceprints corresponding to each of the audio sets; Calculate the similarity of each audio segment voiceprint in one group of audio segment voiceprints and each audio segment voiceprint in another group of audio segment voiceprints to obtain a similarity calculation value; The maximum value among all similarity calculation values ​​is obtained, and the two audio sets corresponding to the maximum value are used as the first audio set and the second audio set for pairing.

4. A voiceprint processing method according to claim 1, characterized in that: The method further comprises: Splicing the audio clips in the audio set to obtain a first spliced ​​audio; Performing duration statistics on the first spliced ​​audio to obtain a duration statistics result; According to the duration statistics result, select a number of candidate audio sets from all audio sets; The performing similarity pairing on all audio sets to obtain the paired first audio set and the second audio set specifically includes: performing similarity pairing on all candidate audio sets to obtain the paired first audio set and the second audio set.

5. A voiceprint processing method according to any one of claims 1 to 4, characterized in that: The performing speech rate analysis on the audio segments of the first audio set and the second audio set to obtain corresponding first speech rate analysis results and second speech rate analysis results respectively includes: Performing speech recognition on the audio segments of the first audio set and the second audio set respectively to obtain text recognition results and corresponding text timestamp information; According to the text recognition result and the corresponding text timestamp information, speech rate analysis is performed on each audio segment in the first audio set and the second audio set to obtain corresponding first speech rate analysis results and second speech rate analysis results.

6. A voiceprint processing method according to claim 1, characterized in that: The step of judging whether the first audio set is a singing audio set or the second audio set is a singing audio set according to the first speech rate analysis result and the second speech rate analysis result comprises: Preset speech rate analysis result threshold; Determining whether the corresponding audio segment is a singing segment according to a comparison between the speech rate analysis result and the speech rate analysis result threshold; Respectively counting the number of audio segments determined to be singing segments in the first audio set and the second audio set to obtain the number of first audio segments and the number of second audio segments; The maximum value between the first audio segment quantity and the second audio segment quantity is obtained, and the audio set corresponding to the maximum value is a singing audio set.

7. A voiceprint processing method according to claim 1, characterized in that: The voiceprint processing is performed on all audio clips in the singing audio set to obtain the singing voiceprint result corresponding to the human voice in the singing audio set, including: Splicing all audio clips in the singing audio set to obtain a second spliced ​​audio; Voiceprint extraction is performed on the second spliced ​​audio to obtain a singing voiceprint result corresponding to the human voice belonging to the singing audio set.

8. A voiceprint processing method for live audio, characterized in that: The voiceprint processing method of live audio includes: Get multiple live audios from different sessions of the same host; Using the method described in any one of claims 1 to 7 to process the plurality of live audios respectively, to obtain a singing voiceprint result of each of the live audios; Clustering the singing voiceprint results of all the live audios to obtain multiple voiceprint result sets corresponding to the same host; selecting the voiceprint result set with the largest number of clusters from the multiple voiceprint result sets as the singing voiceprint of the host.

9. A voiceprint processing device, characterized in that: The voiceprint processing device comprises: A segmentation module, used for segmenting the audio to be processed into a number of audio segments; A clustering module, used for clustering the plurality of audio clips to obtain a plurality of audio sets; A pairing module, used for performing similarity pairing on all audio sets to obtain a paired first audio set and a second audio set; A speech rate analysis module, used to perform speech rate analysis on the audio segments of the first audio set and the second audio set respectively to obtain corresponding first speech rate analysis results and second speech rate analysis results; A judgment processing module is used to judge whether the first audio set is a singing audio set or the second audio set is a singing audio set based on the first speech rate analysis result and the second speech rate analysis result, and perform voiceprint processing on all audio clips in the singing audio set to obtain a singing voiceprint result corresponding to the human voice belonging to the singing audio set.

10. An electronic device, characterized in that: include: a memory for storing one or more computer programs; A processor, when the one or more computer programs are executed by the processor, implements a voiceprint processing method as described in any one of claims 1 to 7 or a voiceprint processing method for live audio as described in claim 8.

Citation Information

Cited By

  • Real-time audio acquisition and intelligent analysis system based on intelligent work card

    CN120748455A