An audio extraction method, device, equipment and storage medium
The method enhances audio extraction by segmenting and analyzing audio segments for similarity with a registered sample, addressing the challenge of mixed speaker audio in multi-speaker scenarios.
Patent Information
- Application Number
- CN202111328474.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-10
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2041-11-10
AI Technical Summary
In the scene where multiple people speak one after another, the speech segmentation clustering method cannot accurately extract the audio of the target object, resulting in unsatisfactory clustering effect and may be mixed with the voices of others.
By obtaining the pending audio and registered audio, dividing it into multiple window segments, extracting feature vectors and performing similarity analysis, determining whether the current window segment is the voice audio of the target object, and using the voiceprint model to improve judgment accuracy.
The precise extraction of the target object voice audio is achieved, reducing interference from non-target object audio and improving the purity of the audio extraction.
Smart Images

Figure CN114049898B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of audio processing, and in particular to an audio extraction method, device, equipment and storage medium for audio of a specific speaker based on a voiceprint model. Background Art
[0002] In order to obtain the speech audio of the target object in a speech segment, it is necessary to extract the speech audio of the target object from the speech segment through specific technical means.
[0003] In existing solutions, the speech segmentation and clustering method is usually used to extract the audio information of the target object. This method is basically applied to the scene where multiple people speak in succession. However, the goal of the speech segmentation and clustering method is to distinguish the audio of all speakers and to segment and cluster the original audio into multiple audio segments. However, the number of speakers in the original audio is uncertain. After obtaining the voiceprint information features of multiple audio segments to be processed, the clustering algorithm does not specify the number of clustering categories. Therefore, the clustering effect in practical applications is not ideal. The audio of the conversation between two people may be clustered into multiple categories. Moreover, the clustered audio is not pure and will be mixed with the voices of others.
[0004] How to accurately extract the audio content of the target object from the recording has become one of the technical problems to be solved urgently in this field. Summary of the invention
[0005] In view of this, embodiments of the present invention provide an audio extraction method, apparatus, device and storage medium to extract the speech audio of a target object.
[0006] To achieve the above objectives, the embodiments of the present invention provide the following technical solutions:
[0007] An audio extraction method, comprising:
[0008] Acquire the audio to be processed and the registered audio, wherein the registered audio is a speech audio of a target object in the audio to be processed;
[0009] Segmenting the audio to be processed to obtain multiple window segments;
[0010] Extracting feature vectors of the registered audio and the window segment;
[0011] Performing a similarity analysis on the feature vector of the window segment and the feature vector of the registered audio;
[0012] Based on the similarity between the feature vectors of the current window segment and the window segments adjacent to the current window segment and the registered audio, determining whether the current window segment is the speech audio of the target object;
[0013] Determine the speech audio of the target object as the extracted audio.
[0014] Optionally, in the above audio extraction method, based on the similarity between the current window segment and the window segments adjacent to the current window segment and the feature vectors of the registered audio, determine whether the current window segment is the speech audio of the target object, including:
[0015] Calculate the similarity between the feature vectors of the current window segment and the registered audio, denoted as the first similarity;
[0016] Determine whether the first similarity is greater than a preset score threshold;
[0017] When it is not greater than the preset score threshold, determine that the current window segment is not the speech audio of the target object;
[0018] When it is greater than the preset score threshold, obtain the second similarity and the third similarity. The second similarity is the similarity between a window segment adjacent to the current window segment and before the current window segment and the feature vectors of the registered audio, and the third similarity is the similarity between a window segment adjacent to the current window segment and after the current window segment and the feature vectors of the registered audio;
[0019] Based on the first similarity, the second similarity, and the third similarity, determine whether the current window segment is the speech audio of the target object.
[0020] Optionally, in the above audio extraction method, based on the first similarity, the second similarity, and the third similarity, determine whether the current window segment is the speech audio of the target object, including:
[0021] Determine whether the first similarity is greater than a preset score threshold. When the first similarity is not greater than the preset score threshold, determine that the current window segment is not the speech audio of the target object;
[0022] When the first similarity is greater than the preset score threshold, determine whether the second similarity and the third similarity meet the first preset condition and the second preset condition;
[0023] The first preset condition is that the voiceprint similarity between the current window segment and the adjacent window segments is lower than the set value;
[0024] The second preset condition is that the values of the second similarity and the third similarity are both less than the preset score threshold;
[0025] When any one of the first preset condition and the second preset condition is satisfied, it is determined that the current window segment is not the voice audio of the target object; otherwise, it is determined that the current window segment is the voice audio of the target object.
[0026] Optionally, in the above audio extraction method, the first preset condition is specifically: the differences between the first similarity and the second similarity, and between the first similarity and the third similarity are both greater than a preset difference threshold.
[0027] Optionally, in the above audio extraction method, splitting the audio to be processed includes:
[0028] Performing voice activity detection on the audio to be processed to remove the silent period in the audio to be processed;
[0029] Using a sliding window to split the audio to be processed after removing the silent period.
[0030] Optionally, in the above audio extraction method, extracting the feature vectors of the registered audio and the window segment includes:
[0031] Extracting the feature vector of the registered audio after data augmentation;
[0032] Extracting the feature vector of the window segment after data augmentation.
[0033] An audio extraction device includes:
[0034] An audio acquisition unit, configured to acquire the audio to be processed and the registered audio, where the registered audio is a segment of voice audio of the target object in the audio to be processed;
[0035] An audio processing unit, configured to split the audio to be processed to obtain a plurality of window segments;
[0036] A feature vector extraction unit, configured to extract the feature vectors of the registered audio and the window segment;
[0037] A similarity calculation unit, configured to perform similarity analysis on the feature vector of the window segment and the feature vector of the registered audio;
[0038] A target audio detection unit, configured to determine whether the current window segment is the voice audio of the target object based on the similarity between the current window segment and the window segments adjacent to the current window segment and the feature vector of the registered audio; and determining the voice audio of the target object as the extracted audio.
[0039] Optionally, in the above audio extraction device, when the target audio detection unit determines whether the current window segment is the voice audio of the target object based on the similarity between the current window segment, the window segment adjacent to the current window segment, and the feature vector of the registered audio, it specifically is used for:
[0040] Calculate the similarity between the current window segment and the feature vector of the registered audio, denoted as the first similarity;
[0041] Determine whether the first similarity is greater than a preset score threshold;
[0042] When it is not greater than the preset score threshold, it is determined that the current window segment is not the voice audio of the target object;
[0043] When it is greater than the preset score threshold, obtain the second similarity and the third similarity. The second similarity is the similarity between a window segment adjacent to the current window segment and before the current window segment and the feature vector of the registered audio, and the third similarity is the similarity between a window segment adjacent to the current window segment and after the current window segment and the feature vector of the registered audio;
[0044] Based on the first similarity, the second similarity, and the third similarity, determine whether the current window segment is the voice audio of the target object.
[0045] An audio extraction device, comprising: a memory and a processor;
[0046] The memory is used for storing programs;
[0047] The processor is used for executing the program to implement each step of the audio extraction method described in any one of the above.
[0048] A computer-readable storage medium, on which a computer program is stored, characterized in that when the computer program is executed by a processor, it implements each step of the audio extraction method described in
[0049] any one of the above.
[0050] Based on the above technical solutions, in the above solution provided by the embodiments of the present invention, during the audio extraction process, a voice audio of a target object in the audio to be processed is used as the registered audio, the audio to be processed is segmented to obtain multiple window segments, then the similarity between the window segments and the registered audio is analyzed, and finally, based on the similarity between the current window segment, the window segment adjacent to the current window segment, and the registered audio, it is determined whether the current window segment is the voice audio of the target object, thereby realizing the accurate extraction of the voice audio of the target object. Description of the Drawings
[0051] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to the provided drawings.
[0052] Figure 1 It is a schematic flow chart of the audio extraction method disclosed in the embodiments of the present application;
[0053] Figure 2 It is a schematic flow chart of the audio extraction method disclosed in another embodiment of the present application;
[0054] Figure 3 It is a schematic structural diagram of the audio extraction device disclosed in the embodiments of the present application;
[0055] Figure 4 It is a schematic structural diagram of the audio extraction device disclosed in the embodiments of the present application. Specific embodiments
[0056] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some of the embodiments of the present invention, rather than all of them. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0057] The embodiments of the present application disclose an audio extraction solution. During the extraction process, a segment of the speech audio of a target object in the audio to be processed is used as the registered audio. The audio to be processed is segmented to obtain multiple window segments, and then the similarity between the window segments and the registered audio is analyzed. Finally, based on the similarity between the current window segment and the window segments adjacent to the current window segment and the registered audio, it is determined whether the current window segment is the speech audio of the target object, thereby realizing the accurate extraction of the speech audio of the target object.
[0058] Specifically, refer to Figure 1 , Figure 1 It is a schematic flow chart of an audio extraction method disclosed in the embodiments of the present application. This method may include steps S101-S106.
[0059] Step S101: Obtain the audio to be processed and the registered audio, where the registered audio is a segment of the speech audio of a target object in the audio to be processed.
[0060] In this solution, the audio to be processed is an audio segment containing the speech audio of the target object. In this audio segment, in addition to the speech audio of the target object, there may also be other speech audio of non-target objects or other audio.
[0061] The registered audio is a segment of the speech audio of the target object extracted from the audio to be processed, and the registered audio can be recognized and extracted by the user from the audio to be processed.
[0062] Step S102: Segment the audio to be processed to obtain multiple window segments.
[0063] In this step, the audio to be processed can be divided into multiple window segments, and each window segment is respectively judged whether it is the speech audio corresponding to the target object. When segmenting the audio to be processed, a sliding window can be used to segment the audio to be processed. The window length (e.g., 0.8s) and sliding overlap rate (e.g., 0.5) of each window segment can be flexibly set. The audio data in each window is written into a two-dimensional array. Among them, the audio data to be processed is one-dimensional data, and the additional one-dimensional after segmentation according to the window length is the data of the corresponding window.
[0064] Step S103: Extract the feature vectors of the registered audio and the window segments.
[0065] In this step, when analyzing the similarity between the window segments and the registered audio, specifically, it is the similarity analysis of the feature vectors. Therefore, before performing the similarity analysis, it is necessary to extract the feature vectors of the registered audio and each window segment in advance. Among them, the feature vectors of the registered audio and the window segments can be obtained by processing the registered audio and the window segments by a voiceprint model.
[0066] Step S104: Perform a similarity analysis on the feature vectors of the window segments and the feature vectors of the registered audio.
[0067] In this step, the feature vectors corresponding to each window segment are respectively scored for similarity with the feature vectors of the registered audio, and stored in an array recording the scores of each window segment. This score is used to represent the similarity between the window segment and the registered audio.
[0068] Step S105: Based on the similarity between the current window segment and the window segments adjacent to the current window segment and the feature vectors of the registered audio, judge whether the current window segment is the speech audio of the target object.
[0069] In this step, the current window segment refers to the window segment of the voice audio for which it is being determined whether it is the voice audio of the target object. In this solution, based on the similarity between the current window segment and the window segments adjacent to the current window segment and the feature vector of the registered audio, it is determined whether the current window segment is the voice audio of the target object. When determining whether the current window segment is the voice audio of the target object, the similarity between the spliced window segment and the feature vector of the registered audio is used as a reference factor, which improves the reliability of the determination result.
[0070] Step S106: Determine the voice audio of the target object as the extracted audio.
[0071] In this step, when it is determined that a certain window segment is the voice audio of the target object, the voice audio of the target object is extracted, and the extracted voice audio of the target object is spliced based on the order of the time axis where the voice audio of the target object is located.
[0072] As can be seen from the above solution, in this application, the audio to be processed is segmented to obtain multiple window segments, then the similarity analysis between the window segments and the registered audio is performed, and finally, based on the similarity between the current window segment and the window segments adjacent to the current window segment and the registered audio, it is determined whether the current window segment is the voice audio of the target object, thereby achieving the precise extraction of the voice audio of the target object.
[0073] In the technical solution disclosed in another embodiment of this application, considering that there are silent periods in the audio to be processed, the silent period is the time period in the audio to be processed where there is no voice audio, and the audio to be processed corresponding to these silent periods does not need to be recognized. Therefore, in order to improve the processing efficiency, in this solution, the silent periods in the audio to be processed can be removed first. Specifically, in the above solution of this application, segmenting the audio to be processed specifically includes:
[0074] Perform voice activity detection on the audio to be processed to remove the silent periods in the audio to be processed;
[0075] Use a sliding window to segment the audio to be processed after removing the silent periods.
[0076] Among them, in the above steps, specifically, a VAD module can be used to process the audio to be processed, identify and process the silent periods in the audio to be processed, then splice the audio to be processed after removing the silent periods, and finally use a sliding window to segment the spliced audio to be processed.
[0077] In the technical solution disclosed in another embodiment of the present application, in order to improve the calculation accuracy of similarity, the registered audio and the window segments obtained by segmentation can be subjected to data augmentation processing first. When calculating similarity, in fact, the similarity between the registered audio and the window segments after augmentation processing is calculated. In this regard, in the above solution, extracting the feature vectors of the registered audio and the window segments includes: extracting the feature vector of the registered audio after data augmentation; extracting the feature vector of the window segments after data augmentation.
[0078] In the technical solution disclosed in the embodiment of the present application, in order to accurately determine whether a window segment is the voice audio of a target object, a specific judgment process is disclosed in this solution. Specifically, see Figure 2 In the above method, based on the similarity between the feature vectors of the current window segment and the window segments adjacent to the current window segment and the registered audio, it is determined whether the current window segment is the voice audio of the target object. Specifically, it may include:
[0079] Step S201: Calculate the similarity between the feature vectors of the current window segment and the registered audio, denoted as the first similarity.
[0080] In this step, based on the order of the time axis, it is sequentially determined whether each window segment is the voice audio of the target object. In this process, the window segment being judged is used as the current window segment, and the similarity between the feature vector of the current window segment and the feature vector of the registered audio is denoted as the first similarity.
[0081] Step S202: Determine whether the first similarity is greater than a preset score threshold.
[0082] In this solution, a preset score threshold is set in advance. First, the first similarity is compared with the preset score threshold. Among them, the size of the preset score threshold can be set according to user needs. For example, the preset score threshold can be 0.5. When the first similarity is less than the preset score threshold, it is determined that the acoustic difference between the current window segment and the registered audio is too large, and the current window segment is not the voice audio of the target object. Otherwise, step S204 is continued to make a further judgment.
[0083] Step S203: When it is not greater than the preset score threshold, it is determined that the current window segment is not the voice audio of the target object.
[0084] Step S204: When it is greater than a preset score threshold, obtain a second similarity and a third similarity. The second similarity is the similarity between at least one window segment adjacent to and before the current window segment and the feature vector of the registered audio. The third similarity is the similarity between at least one window segment adjacent to and after the current window segment and the feature vector of the registered audio;
[0085] In this step, when the first similarity is greater than the preset score threshold, other window segments adjacent to the current window segment are used to continue to determine whether the current window segment is the speech audio of the target object. At this time, obtain the similarities between the two window segments adjacent to the current window segment and the feature vector of the registered audio, and denote them as the second similarity and the third similarity respectively. Among them, the second similarity is the similarity between the window segment before the current window segment and the feature vector of the registered audio, and the third similarity is the similarity between the window segment after the current window segment and the feature vector of the registered audio. Of course, this situation is for when the current window segment has two adjacent window segments and it is necessary to obtain the second similarity and the third similarity. If the current window segment has only one adjacent window segment, only one similarity needs to be obtained.
[0086] Step S205: Based on the first similarity, the second similarity, and the third similarity, determine whether the current window segment is the speech audio of the target object.
[0087] When the current window segment has two adjacent window segments, determine whether the current window segment is the speech audio of the target object based on the first similarity, the second similarity, and the third similarity;
[0088] When the current window segment has only 1 adjacent window segment, determine whether the current window segment is the speech audio of the target object based on the first similarity and the second similarity.
[0089] In the technical solution disclosed in the embodiments of the present application, determining whether the current window segment is the speech audio of the target object based on the first similarity, the second similarity, and the third similarity may specifically include:
[0090] Determine whether the first similarity, the second similarity, and the third similarity satisfy a first preset condition and a second preset condition; when any one of the first preset condition and the second preset condition is satisfied, it is determined that the current window segment is not the voice audio of the target object, otherwise, it is determined that the current window segment is the voice audio of the target object, that is, when the first similarity, the second similarity, and the third similarity do not satisfy the first preset condition and the second preset condition at the same time, it is determined that the current window segment is the voice audio of the target object.
[0091] Wherein, the first preset condition is that the voiceprint similarity between the current window segment and its adjacent window segments is lower than a set value; in this solution, when judging the voiceprint similarity between the current window segment and its adjacent window segments, the technical means adopted can be selected according to user needs. For example, the voiceprint features between the current window segment and its adjacent window segments can be directly extracted and compared, and the voiceprint similarity between the current window segment and its adjacent window segments can be judged through the voiceprint features. In the technical solution disclosed in another embodiment of the present application, the first similarity represents the similarity between the current window segment and the registered audio, and the second similarity and the third similarity respectively represent the similarity between the adjacent window segments of the current window segment and the registered audio. Therefore, the voiceprint similarity between the current window segment and its adjacent window segments can be judged by comparing the first similarity, the second similarity, and the third similarity. When the difference between the first similarity and the second similarity is greater than the preset difference threshold, it indicates that the voiceprint similarity between the current window segment and the previous adjacent window segment is lower than the set value. When the difference between the first similarity and the third similarity is greater than the preset difference threshold, it indicates that the voiceprint similarity between the current window segment and the next adjacent window segment is lower than the set value.
[0092] The second preset condition is that the values of the second similarity and the third similarity are both less than the preset score threshold; when the current window segment has only one adjacent window segment, only the corresponding similarity of the adjacent window segment needs to be judged whether it is less than the preset score threshold in the second preset condition.
[0093] When the current window segment does not have a previous window segment, it is directly assumed that the voiceprint similarity between the current window segment and the previous adjacent window segment is higher than the set value. When the current window segment does not have at least one subsequent window segment, it is directly assumed that the voiceprint similarity between the current window segment and the next adjacent window segment is higher than the set value.
[0094] In the technical solution disclosed in the embodiments of the present application, when it is determined that a certain window segment is the voice audio of the target object, it is determined whether the window is the window segment of the first voice audio of the target object determined in the sliding window. If it is, all the audio data in the window segment is written into the new extracted audio. If it is not the window segment of the first voice audio of the target object determined, only the audio data of the forward sliding duration needs to be appended to the new audio. For example, the time point of the window segment of the first voice audio of the target object is 0 to 0.8 seconds, and the time point of the window segment of the second voice audio of the target object is 0.5 to 1.3 seconds. Since there is an overlap between the two window segments, only the unique sliding part of the window segment needs to be appended and written into the new audio, so as to obtain the audio data from 0.8 to 1.3 seconds.
[0095] In this embodiment, an audio extraction device is disclosed. For the specific working content of each unit in the device, please refer to the content of the above method embodiment.
[0096] The audio extraction device provided by the embodiments of the present invention will be described below. The audio extraction device described below can be correspondingly referred to the audio extraction method described above.
[0097] See Figure 3 , the audio extraction device disclosed in the embodiments of the present application may include:
[0098] An audio acquisition unit A, which corresponds to step S101 in the above method, is used to acquire the audio to be processed and the registered audio, and the registered audio is a voice audio of a target object in the audio to be processed;
[0099] An audio processing unit B, which corresponds to step S102 in the above method, is used to segment the audio to be processed to obtain a plurality of window segments;
[0100] A feature vector extraction unit C, which corresponds to step S103 in the above method, is used to extract the feature vectors of the registered audio and the window segments;
[0101] A similarity calculation unit D, which corresponds to step S104 in the above method, is used to perform similarity analysis on the feature vectors of the window segments and the feature vectors of the registered audio;
[0102] A target audio detection unit E, which corresponds to steps S105-S106 in the above method, is used to determine whether the current window segment is the voice audio of the target object based on the similarity between the current window segment and the window segment adjacent to the current window segment and the feature vectors of the registered audio; and determine the voice audio of the target object as the extracted audio.
[0103] Corresponding to the above method, when the target audio detection unit determines whether the current window segment is the voice audio of the target object based on the similarity between the current window segment and the window segments adjacent to the current window segment and the feature vector of the registered audio, it specifically is used for:
[0104] Calculate the similarity between the current window segment and the feature vector of the registered audio, denoted as the first similarity;
[0105] Determine whether the similarity between the current window segment and the feature vector of the registered audio is greater than a preset score threshold;
[0106] When it is not greater than the preset score threshold, it is determined that the current window segment is not the voice audio of the target object;
[0107] When it is greater than the preset score threshold, obtain a second similarity and a third similarity. The second similarity is the similarity between at least one window segment adjacent to the current window segment and before the current window segment and the feature vector of the registered audio, and the third similarity is the similarity between at least one window segment adjacent to the current window segment and after the current window segment and the feature vector of the registered audio;
[0108] Based on the first similarity, the second similarity, and the third similarity, determine whether the current window segment is the voice audio of the target object.
[0109] Corresponding to the above method, when the target audio detection unit determines whether the current window segment is the voice audio of the target object based on the first similarity, the second similarity, and the third similarity, it specifically is used for:
[0110] Determine whether the first similarity, the second similarity, and the third similarity satisfy a first preset condition and a second preset condition;
[0111] The first preset condition is that the first similarity is less than the second similarity, and the difference between the first similarity and the second similarity is greater than a preset difference threshold;
[0112] The second preset condition is that the values of the second similarity and the third similarity are both less than the preset score threshold;
[0113] When the second similarity and the third similarity satisfy any one of the first preset condition and the second preset condition, it is determined that the current window segment is not the voice audio of the target object, otherwise, it is determined that the current window segment is the voice audio of the target object.
[0114] See Figure 4 , Figure 4The following is the hardware structure diagram of the audio extraction device provided by the embodiments of the present invention. Refer to Figure 4 As shown, the device may include: at least one processor 100, at least one communication interface 200, at least one memory 300, and at least one communication bus 400;
[0115] In the embodiments of the present invention, the number of the processor 100, the communication interface 200, the memory 300, and the communication bus 400 is at least one, and the processor 100, the communication interface 200, and the memory 300 complete communication with each other through the communication bus 400; obviously, Figure 4 The communication connection schematic diagram of the processor 100, the communication interface 200, the memory 300, and the communication bus 400 shown is only optional;
[0116] Optionally, the communication interface 200 may be an interface of a communication module, such as an interface of a GSM module;
[0117] The processor 100 may be a central processing unit CPU, or a specific integrated circuit ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present invention.
[0118] The memory 300 may include a high-speed RAM memory, and may also include a non-volatile memory, such as at least one disk memory.
[0119] Among them, the processor 100 is specifically used for:
[0120] Obtain the audio to be processed and the registered audio, where the registered audio is a voice audio of a target object in the audio to be processed;
[0121] Segment the audio to be processed to obtain multiple window segments;
[0122] Extract the feature vectors of the registered audio and the window segments;
[0123] Perform a similarity analysis on the feature vectors of the window segments and the feature vectors of the registered audio;
[0124] Based on the similarity between the current window segment and the window segments adjacent to the current window segment and the feature vectors of the registered audio, determine whether the current window segment is the voice audio of the target object;
[0125] Determine the voice audio of the target object as the extracted audio.
[0126] The processor is further configured to execute other steps in the audio extraction method disclosed in the foregoing embodiments of the present application, which will not be elaborated herein.
[0127] Corresponding to the above method, the present application also discloses a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, each step of the audio extraction method described in any one of the above is implemented.
[0128] For example, when the computer program is executed by a processor, it is configured to:
[0129] Obtain an audio to be processed and a registered audio, where the registered audio is a voice audio of a target object in the audio to be processed;
[0130] Segment the audio to be processed to obtain a plurality of window segments;
[0131] Extract the feature vectors of the registered audio and the window segments;
[0132] Perform a similarity analysis on the feature vectors of the window segments and the feature vectors of the registered audio;
[0133] Based on the similarity between the current window segment and the window segments adjacent to the current window segment and the feature vectors of the registered audio, determine whether the current window segment is a voice audio of the target object;
[0134] Determine the voice audio of the target object as the extracted audio.
[0135] For convenience of description, the above system is described by dividing it into various modules according to functions. Of course, when implementing the present invention, the functions of each module can be implemented in the same or multiple software and / or hardware.
[0136] Each embodiment in this specification is described in a progressive manner. The same or similar parts between the embodiments can be referred to each other, and the key points of each embodiment are the differences from other embodiments. In particular, for the system or system embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method embodiments. The systems and system embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0137] Those skilled in the art may further realize that the units and algorithm steps of each example described in connection with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present invention.
[0138] The steps of the methods or algorithms described in connection with the embodiments disclosed herein can be directly implemented by hardware, software modules executed by a processor, or a combination of the two. The software modules can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the art.
[0139] It should also be noted that, in this document, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising a..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the element.
[0140] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. An audio extraction method, characterized in that, Including: Obtain the audio to be processed and the registered audio, where the registered audio is a voice audio of a target object in the audio to be processed; Segment the audio to be processed to obtain multiple window segments; Extract the feature vectors of the registered audio and the window segments; Perform similarity analysis on the feature vectors of the window segments and the feature vectors of the registered audio; Based on the similarity between the feature vectors of the current window segment and the window segments adjacent to the current window segment and the registered audio, determine whether the current window segment is a voice audio of the target object; the window segments adjacent to the current window segment include one window segment adjacent to and before the current window segment, and one window segment adjacent to and after the current window segment; Determine the voice audio of the target object as the extracted audio.
2. The audio extraction method according to claim 1, wherein Based on the similarity between the feature vectors of the current window segment and the window segments adjacent to the current window segment and the registered audio, determining whether the current window segment is a voice audio of the target object includes: Calculate the similarity between the feature vectors of the current window segment and the registered audio, denoted as the first similarity; Determine whether the first similarity is greater than a preset score threshold; When it is not greater than the preset score threshold, determine that the current window segment is not a voice audio of the target object; When it is greater than the preset score threshold, obtain the second similarity and the third similarity, where the second similarity is the similarity between the feature vectors of one window segment adjacent to and before the current window segment and the registered audio, and the third similarity is the similarity between the feature vectors of one window segment adjacent to and after the current window segment and the registered audio; Based on the first similarity, the second similarity, and the third similarity, determine whether the current window segment is a voice audio of the target object.
3. The audio extraction method according to claim 2, wherein Based on the first similarity, the second similarity, and the third similarity, determining whether the current window segment is a voice audio of the target object includes: Determine whether the first similarity, the second similarity, and the third similarity satisfy the first preset condition and the second preset condition; The first preset condition is that the voiceprint similarity between the current window segment and the window segments adjacent to it is lower than a set value; The second preset condition is that the values of the second similarity and the third similarity are both less than the preset score threshold; When any one of the first preset condition and the second preset condition is satisfied, determine that the current window segment is not a voice audio of the target object, otherwise, determine that the current window segment is a voice audio of the target object.
4. The audio extraction method according to claim 3, wherein The first preset condition is specifically that the difference between the first similarity and the second similarity, and the difference between the first similarity and the third similarity are both greater than a preset difference threshold.
5. The audio extraction method according to any one of claims 1-4, characterized in that, Segmenting the audio to be processed includes: Perform voice activity detection on the audio to be processed to remove the silent period in the audio to be processed; Use a sliding window to segment the audio to be processed after removing the silent period.
6. The audio extraction method according to any one of claims 1-4, characterized in that, Extract the feature vectors of the registered audio and the window segments, including: Extract the feature vector of the registered audio after data augmentation; Extract the feature vector of the window segment after data augmentation.
7. An audio extraction device, characterized in that, Include: An audio acquisition unit for acquiring the audio to be processed and the registered audio, where the registered audio is a voice audio of a target object in the audio to be processed; An audio processing unit for segmenting the audio to be processed to obtain multiple window segments; A feature vector extraction unit for extracting the feature vectors of the registered audio and the window segments; A similarity calculation unit for performing similarity analysis on the feature vectors of the window segments and the feature vectors of the registered audio; A target audio detection unit for judging whether the current window segment is a voice audio of the target object based on the similarity between the current window segment and the feature vectors of the window segments adjacent to the current window segment and the registered audio; determining the voice audio of the target object as the extracted audio; the window segments adjacent to the current window segment include one window segment adjacent to the current window segment and before the current window segment, and one window segment adjacent to the current window segment and after the current window segment.
8. The audio extraction device according to claim 7, characterized in that, When the target audio detection unit judges whether the current window segment is a voice audio of the target object based on the similarity between the current window segment and the feature vectors of the window segments adjacent to the current window segment and the registered audio, it specifically uses: Calculate the similarity between the feature vectors of the current window segment and the registered audio, denoted as the first similarity; Judge whether the first similarity is greater than a preset score threshold; When it is not greater than the preset score threshold, determine that the current window segment is not a voice audio of the target object; When it is greater than the preset score threshold, obtain the second similarity and the third similarity, where the second similarity is the similarity between one window segment adjacent to the current window segment and before the current window segment and the feature vectors of the registered audio, and the third similarity is the similarity between one window segment adjacent to the current window segment and after the current window segment and the feature vectors of the registered audio; Judge whether the current window segment is a voice audio of the target object based on the first similarity, the second similarity, and the third similarity.
9. An audio extraction device, characterized in that, Include: A memory and a processor; The memory is used to store programs; The processor is used to execute the program to implement each step of the audio extraction method described in any one of claims 1-6.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements each step of the audio extraction method described in any one of claims 1-6.
Citation Information
Patent Citations
Method and device for detecting voice turning points of multiple speakers
CN112951212A
Information processing apparatus, control method, and program
US20210287682A1