Speaker-specific speech extraction method and device based on speaker auxiliary information
By framing audio and video data, detecting silence, and implementing speaker classification models, the complexity of extracting the speech of a specific speaker in the existing technology is solved, and efficient speech separation is achieved under unknown speaker numbers and corresponding relationships.
Patent Information
- Application Number
- CN202210610122.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-31
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2042-05-31
AI Technical Summary
In the existing technology, the method of extracting the speech of a specific speaker requires knowing in advance the number of speakers in the mixed speech and the correspondence between speakers and output is not one-to-one. This makes the extraction process complicated and requires the speaker to enter voice information in advance, making it impossible to accurately extract the speech of a specific speaker from the mixed speech of multiple people.
By obtaining the frame segmentation results of the audio and video data to be identified, the sub-speaker activity information is obtained using silence detection, and the target speaker recognition features are generated by combining the logarithmic spectrum amplitude coefficient and the pre-trained speaker classification model. Finally, it is multiplied by the logarithmic spectrum amplitude coefficient to obtain the target speaker spectrum.
It realizes the separation of the speech spectrum of a specific speaker from the mixed speech without inputting the speech of the target speaker, simplifies the extraction process, and improves the accuracy and efficiency of the extraction.
Smart Images

Figure CN114999522B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence speech processing technology, and in particular to a method and device for extracting a specific speaker's speech based on speaker auxiliary information. Background Art
[0002] Currently, systems that use speaker-assisted information to extract a specific speaker's voice are designed to isolate the voice of a specific speaker from a mix of multiple speakers. Traditional methods typically separate the speech to extract individual speaker segments, then perform speaker authentication using the registered target speaker's voice. However, these traditional methods often suffer from the following two issues:
[0003] 1) The number of speakers in a mixed speech must be known or inferred in advance in order to accurately extract the speech of a specific speaker from the mixed speech of multiple speakers;
[0004] 2) The correspondence between speakers and outputs is not one-to-one (speaker 1 may be the first output node or may be another, and it is impossible to splice all the speech segments of speaker 1 for the entire speech). It is also impossible to accurately extract the speech of a specific speaker from the mixed speech of multiple speakers. Summary of the Invention
[0005] The embodiments of the present application provide a method, apparatus, device, and medium for extracting the speech of a specific speaker based on speaker auxiliary information, aiming to solve the problem in the prior art of extracting the speech of a specific speaker using speaker auxiliary information, that the correspondence between the speaker and the output is not a one-to-one correspondence, and that the number of speakers in the mixed speech must be known in order to accurately extract the speech of the specific speaker, resulting in a complicated extraction process and requiring the speaker to enter speech information in advance before speech extraction.
[0006] In a first aspect, an embodiment of the present application provides a method for extracting a specific speaker's speech based on speaker auxiliary information, which includes:
[0007] Acquire audio and video data to be identified, and divide the audio and video data to be identified into frames to obtain a frame division result; wherein the frame division result includes multiple frames of sub-audio and video data;
[0008] Acquire sub-speaker activity information corresponding to each frame of sub-audio and video data in the frame segmentation result through silence detection to form speaker activity information;
[0009] Obtaining the logarithmic spectrum amplitude coefficient of each frame of sub-audio and video data in the framing result;
[0010] Generate input data of corresponding frame sub-audio and video data according to the logarithmic spectrum amplitude coefficient of each frame sub-audio and video data;
[0011] Obtaining a pre-trained speaker classification model, inputting the input data into the speaker classification model for classification, and obtaining target speaker recognition features; and
[0012] The target speaker identification feature is multiplied by the logarithmic spectrum amplitude coefficient of each frame of sub-audio and video data to obtain the target speaker spectrum corresponding to each frame of sub-audio and video data.
[0013] In a second aspect, an embodiment of the present application provides a device for extracting a specific speaker's speech based on speaker auxiliary information, comprising:
[0014] A framing unit is configured to obtain audio and video data to be identified, and to frame the audio and video data to be identified to obtain a framing result; wherein the framing result includes multiple frames of sub-audio and video data;
[0015] an activity information acquisition unit, configured to acquire, by silence detection, sub-speaker activity information corresponding to each frame of sub-audio and video data in the frame segmentation result, to form speaker activity information;
[0016] a logarithmic spectrum amplitude coefficient acquisition unit, configured to acquire the logarithmic spectrum amplitude coefficient of each frame of sub-audio and video data in the framing result;
[0017] An input data generating unit, configured to generate input data of a corresponding frame of sub-audio and video data according to a logarithmic spectrum amplitude coefficient of each frame of sub-audio and video data;
[0018] a speaker classification unit, configured to obtain a pre-trained speaker classification model, input the input data into the speaker classification model for classification, and obtain target speaker recognition features; and
[0019] The target spectrum acquisition unit is configured to multiply the target speaker identification feature by the log spectrum amplitude coefficient of each frame of sub-audio and video data to obtain the target speaker spectrum corresponding to each frame of sub-audio and video data.
[0020] In a third aspect, an embodiment of the present application further provides a computer device, which includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the method for extracting a specific speaker's speech based on speaker auxiliary information as described in the first aspect above is implemented.
[0021] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor executes the speaker-specific speech extraction method based on speaker auxiliary information described in the first aspect above.
[0022] The embodiments of the present application provide a method, apparatus, device, and medium for extracting the speech of a specific speaker based on speaker auxiliary information. The method first obtains the framing result corresponding to the audio and video data to be identified, then obtains the sub-speaker activity information corresponding to each frame of sub-audio and video data in the framing result to form the speaker activity information. Then, based on the logarithmic spectrum amplitude coefficient of each frame of sub-audio and video data in the framing result, input data of the corresponding frame of sub-audio and video data is generated. Finally, the input data is input into the speaker classification model for classification. The obtained target speaker recognition feature is then multiplied by the logarithmic spectrum amplitude coefficient of each frame of sub-audio and video data to obtain the target speaker spectrum corresponding to each frame of sub-audio and video data. This method achieves the separation of the speech spectrum of a specific speaker from mixed speech without inputting the target speaker's speech, thus simplifying the extraction process. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0024] Figure 1 A schematic diagram of an application scenario of the speaker-specific speech extraction method based on speaker-auxiliary information provided in an embodiment of the present application;
[0025] Figure 2 A flowchart of a method for extracting speaker-specific speech based on speaker-aided information provided in an embodiment of the present application;
[0026] Figure 3 A schematic block diagram of a speaker-specific speech extraction device based on speaker auxiliary information provided in an embodiment of the present application;
[0027] Figure 4 A schematic block diagram of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0028] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0029] It will be understood that when used in this specification and the appended claims, the terms “comprises” and “comprising” indicate the presence of described features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.
[0030] It should also be understood that the terms used in this specification are for the purpose of describing specific embodiments only and are not intended to limit the present application. As used in this specification and the appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise.
[0031] It should be further understood that the term "and / or" used in this specification and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.
[0032] See also Figure 1 and Figure 2 , Figure 1 A schematic diagram of an application scenario of the speaker-specific speech extraction method based on speaker-auxiliary information provided in an embodiment of the present application; Figure 2 This is a flow chart of a speaker-specific speech extraction method based on speaker auxiliary information provided in an embodiment of the present application. The speaker-specific speech extraction method based on speaker auxiliary information is applied to a server and is executed by application software installed in the server.
[0033] like Figure 2 As shown, the method includes steps S110 to S160.
[0034] S110 , obtaining audio and video data to be identified, dividing the audio and video data to be identified into frames, and obtaining a frame division result; wherein the frame division result includes multiple frames of sub-audio and video data.
[0035] In this embodiment, the technical solution is described with the server as the execution subject. The server first obtains the audio and video data to be recognized corresponding to the multi-person mixed speaking scene. This audio and video data to be recognized can be collected by the microphone and camera of the user end and uploaded to the server.
[0036] The specific forms of the audio and video data to be identified include the following: one is audio data consisting entirely of audio data (i.e., pure audio data), and the other is audio and video data including both audio and video data. In both forms of audio data to be identified, there are periods of time when users (one or more users) speak, and there are also periods of time when no users speak.
[0037] In order to further extract the target speaker's voice from the audio and video data to be identified, it is necessary to first frame the audio and video data to be identified to obtain a frame result. Specifically, the audio data (i.e., voice signal) in the audio and video data to be identified can be framed based on a preset frame length. If the preset frame length is equal to 50ms, the audio and video data to be identified is divided into multiple 50ms long sub-audio and video data, and each 50ms long sub-audio and video data is recorded as a frame of sub-audio and video data. For example, if the audio and video data to be identified is recorded as s, it can be divided into s1, s2, ..., s n These sub-audio and video data, namely s={s1,s2,……,s n It can be seen that based on the above frame processing, the audio and video data to be recognized can be effectively divided into multiple small segments, and then more refined speaker voice extraction can be performed in each small segment of audio and video data.
[0038] S120 , obtaining sub-speaker activity information corresponding to each frame of sub-audio and video data in the frame segmentation result through silence detection to form speaker activity information.
[0039] In this embodiment, at this time, as long as the time period in which the user is speaking in the audio data to be identified is located, the speaker activity information can be accurately determined. That is, after completing the audio framing, the silence detection technology can be used to determine whether each frame of sub-audio and video data is silent data. For example, if there is a speaker speaking in the sub-audio and video data corresponding to s1, then the sub-speaker activity information p1 corresponding to s1 is 1; for example, if there is no speaker speaking in the sub-audio and video data corresponding to s2 (that is, silent data), then the sub-speaker activity information p2 corresponding to s2 is 0. And so on, the sub-audio and video data s can be determined based on the silence detection technology. t (s t Indicates whether the t-th frame of the audio and video data to be identified is silent data, thereby further determining the sub-speaker activity information corresponding to each frame of the sub-audio and video data. t =1 means that the speaker is speaking in the sub-audio and video data of the tth frame, p t = 0 means that there is no speaker in the t-th frame of sub-audio and video data. It can be seen that based on the silence detection technology, the sub-speaker activity information in each frame of sub-audio and video data can be quickly determined to serve as auxiliary information for subsequent analysis of each frame of sub-audio and video data.
[0040] S130: Obtain the logarithmic spectrum amplitude coefficient of each frame of sub-audio and video data in the frame division result.
[0041] In this embodiment, the audio and video data to be identified is recorded as s={s1, s2, ..., s nAt this point, the logarithmic spectrum amplitude coefficient corresponding to each frame of sub-audio and video data in the audio and video data to be identified, s, can be obtained. Specifically, the 257-dimensional logarithmic spectrum amplitude coefficient corresponding to each frame of sub-audio and video data can be obtained. The relationship between signal frequency and energy is represented by a spectrum, and the logarithmic amplitude spectrum is one type of spectrum graph. The amplitude of each spectral line in the logarithmic amplitude spectrum is calculated logarithmically (20logA) with respect to the original amplitude A, so the unit of its vertical axis is dB (decibel). The purpose of this transformation is to raise the lower-amplitude components relative to the higher-amplitude components, so as to observe periodic signals hidden in low-amplitude noise.
[0042] For example, after obtaining the 257-dimensional logarithmic spectrum amplitude coefficient corresponding to s1, it is recorded as y1; and so on, s2, s3, and so on. n After all the corresponding data are converted into 257-dimensional logarithmic spectrum amplitude coefficients, the logarithmic spectrum amplitude coefficients of each frame of sub-audio and video data in the frame division result are obtained, and are recorded as y1, y2, ..., y n Among them, y1, y2, ..., y n The total is y={y1,y2,……,y n}.
[0043] S140 , generating input data of the corresponding frame of sub-audio and video data according to the logarithmic spectrum amplitude coefficient of each frame of sub-audio and video data.
[0044] In this embodiment, to generate input data for identifying the target speaker's spectrum based on the logarithmic spectral amplitude coefficients of each frame of sub-audio and video data, the logarithmic spectral amplitude coefficients of each frame of sub-audio and video data can be used as the input data alone, or together with other parameters such as the sub-speaker activity information stored in the sub-audio and video data. The input data obtained in this way can be used more accurately as input data for identifying the target speaker's spectrum.
[0045] In one embodiment, as a first implementation method of obtaining input data, step S140 includes:
[0046] The logarithmic spectrum amplitude coefficient of each frame of sub-audio and video data is used as the input data of the corresponding frame of sub-audio and video data.
[0047] In this embodiment, the first method for composing the input data corresponding to each frame of sub-audio and video data is to directly use the logarithmic spectral amplitude coefficients of each frame of sub-audio and video data as the input data for that frame of sub-audio and video data, that is, without combining them with other parameters. This method of generating input data places greater emphasis on the impact of the logarithmic spectral amplitude coefficients themselves on the recognition results.
[0048] In one embodiment, as a second implementation method of obtaining input data, step S140 includes:
[0049] The logarithmic spectrum amplitude coefficient of each frame of sub-audio and video data is connected with the sub-speaker activity information of the corresponding frame of sub-audio and video data through a connection function to obtain the input data of each frame of sub-audio and video data.
[0050] In this embodiment, a second method for assembling the input data corresponding to each frame of sub-audio and video data is to concatenate the logarithmic spectral amplitude coefficients of each frame of sub-audio and video data with the sub-speaker activity information of the corresponding frame of sub-audio and video data using a concatenation function. The concatenation function, also known as the Concatenate function, concatenates multiple strings into a single string. The input data obtained through this method includes comprehensive information from two dimensions: the logarithmic spectral amplitude coefficients and the sub-speaker activity information.
[0051] In one embodiment, as a third implementation method of obtaining input data, step S140 includes:
[0052] Connecting the logarithmic spectrum amplitude coefficient of each frame of sub-audio and video data with the sub-speaker activity information of the corresponding frame of sub-audio and video data through a connection function to obtain the first sub-input data of each frame of sub-audio and video data;
[0053] The logarithmic spectrum amplitude coefficient of each frame of sub-audio and video data and the sub-speaker activity information of the corresponding frame of sub-audio and video data are input into a pre-trained auxiliary model to perform an operation to obtain second sub-input data;
[0054] The first sub-input data of each frame of sub-audio and video data is combined with the second sub-input data to obtain the input data of each frame of sub-audio and video data.
[0055] In this embodiment, the pre-trained auxiliary network is AuxiliaryNet, which has two fully connected layers, each with 50 nodes, one of which uses the ReLU function (i.e., linear rectifier function) as the activation function, and the other fully connected layer serves as the output layer and uses a linear activation function (such as one of the sigmoid function, tanh function, ReLU function, ELU function, and PReLU function). The logarithmic spectrum amplitude coefficient of each frame of sub-audio and video data and the sub-speaker activity information of the corresponding frame of sub-audio and video data are input into the pre-trained auxiliary model for operation to obtain the second sub-input data. For details, please refer to the following operation expression (1):
[0056]
[0057] Among them, p i Indicates the active information of the sub-speaker of the i-th frame of sub-audio and video data, y irepresents the logarithmic spectrum amplitude coefficient of the i-th frame sub-audio and video data, and p i ∈{0,1} and indicates whether a person speaks. If p i =1 means that the speaker is speaking in the i-th frame of sub-audio and video data, p i =0 represents that there is no speaker in the i-th frame of sub-audio and video data.
[0058] After calculating the second sub-input data e based on the above formula (1), it can be combined with the first sub-input data of each frame of sub-audio and video data (the combination here is different from the function of connecting strings of the concatenation function, which can be to form two data into a row vector form) to obtain the input data of each frame of sub-audio and video data. For example, the first sub-input data of the i-th frame of sub-audio and video data is used to Concatenate(y i , p i ) means that, at this time, it is combined with e to obtain (Concatenate(y i , p i ), e). The input data obtained in the above manner includes comprehensive information in three dimensions.
[0059] S150: Obtain a pre-trained speaker classification model, input the input data into the speaker classification model for classification, and obtain target speaker recognition features.
[0060] In this embodiment, the pre-trained speech classification network is an extraction network structure (ExtractionNet), and the input of the extraction network structure is the input data, specifically including a 3-layer Bi-LSTM (1200 units) mixed with a 2-layer fully connected layer (ReLU as the activation function).
[0061] Corresponding to the three ways of obtaining input data, the above three types of input data can all be input into the speaker classification model. The input data is input into the speaker classification model for classification to obtain target speaker recognition features.
[0062] For example, taking the first form of input data as input to the speaker classification model, one of the first form of input data is y i To express, with y i The corresponding target speaker recognition features are used Indicates that based on y i Get Please refer to the following formula (2):
[0063]
[0064] Also coming soon iInput it into the extraction network structure (ExtractionNet) for operation to obtain the same value as y i Corresponding target speaker recognition features
[0065] For example, taking the second form of input data as input to the speaker classification model, one of the second form of input data is Concatenate(y i , p i ) to express, and Concatenate(y i , p i ) The corresponding target speaker identification feature is used Indicates that, based on Concatenate(y i , p i )Get Please refer to the following formula (3):
[0066]
[0067] Also about to Concatenate(y i , p i ) is input into the extraction network structure (ExtractionNet) for operation to obtain the same value as y i Corresponding target speaker recognition features
[0068] For example, taking the third form of input data as input to the speaker classification model, one of the third form of input data is (Concatenate(y i , p i ), e) to express, and (Concatenate(y i , p i ), e) the corresponding target speaker identification features are used Indicates that, based on (Concatenate(y i , p i ), e) obtain Please refer to the following formula (4):
[0069]
[0070] Also about (Concatenate (y i , p i ), e) input into the extraction network structure (ExtractionNet) for operation to obtain the same value as y i Corresponding target speaker recognition features
[0071] The target speaker recognition features corresponding to each frame of sub-audio and video data can be extracted through the above three methods, which facilitates the subsequent extraction of the target speaker spectrum based on the target speaker recognition features.
[0072] In one embodiment, before step S150, the method further includes:
[0073] A target speaker speech set is obtained, and a speaker classification model to be trained is trained using the target speaker speech set as a training set to obtain a speaker classification model.
[0074] In this embodiment, the ExtractionNet structure, or speaker classification model, is trained based on the target speaker's speech set. Its purpose is to extract only the specific speaker identification features from the audio and video data to be recognized in mixed speech scenes. Thus, when training the speaker classification model, the target speaker's speech set (i.e., the speech set of the specific speaker) is input for model training, ultimately resulting in a speaker classification model specifically designed to extract specific speaker identification features from the audio and video data to be recognized.
[0075] In order to allow the ExtractionNet structure to better learn the characteristics of a specific speaker, only non-overlapping speech is used for training. That is, each target speaker speech in the target speaker speech set is different.
[0076] In one embodiment, in order to improve the robustness of the speaker classification model, the target speaker speech set may be preprocessed. Therefore, before obtaining the target speaker speech set, the following steps are further included:
[0077] An initial target speaker speech set is obtained, and a duration of each initial target speaker speech data in the initial target speaker speech set is randomly increased or decreased by a random duration to update each initial target speaker speech data to form a target speaker speech set; wherein the random duration has a value range of [0, 1s].
[0078] In this embodiment, taking one of the initial target speaker speech data in the initial target speaker speech set as an example, if its duration is 6 seconds, the initial target speaker speech data is updated by extending it forward by 1 second from its starting time point (if this 1 second is silent audio data) to obtain a 7-second duration initial target speaker speech data. Alternatively, the initial target speaker speech data is updated by extending it backward by 1 second from its ending time point (if this 1 second is silent audio data) to obtain a 7-second duration initial target speaker speech data. The target speaker speech set obtained in this way is more suitable for training a more robust speaker classification model.
[0079] S160: Multiply the target speaker identification feature by the logarithmic spectrum amplitude coefficient of each frame of sub-audio and video data to obtain the target speaker spectrum corresponding to each frame of sub-audio and video data.
[0080] In this embodiment, for example, after obtaining the sub-audio and video data s i Corresponding target speaker recognition features At this time, the target speaker identification feature With sub-audio and video data i The corresponding logarithmic spectrum amplitude coefficient y i Multiply them together to get the sub-audio and video data s i Corresponding target speaker spectrum The above calculation can refer to the following formula (5):
[0081]
[0082] That is, through the above operation, the frequency spectrum of the target speaker corresponding to each frame of sub-audio and video data can be obtained.
[0083] In one embodiment, after step S160, the method further includes:
[0084] The target speaker's spectrum corresponding to each frame of sub-audio and video data is sequentially subjected to inverse Fourier transform and spliced to obtain the target speaker's speech data.
[0085] In this embodiment, the target speaker's speech data is obtained by performing an inverse Fourier transform on the target speaker's spectrum corresponding to each frame of the sub-audio video data, and then concatenating the frames to restore the speech. This method allows the target speaker's speech data to be extracted from mixed speech without pre-recording the target speaker's speech.
[0086] The embodiments of the present application can acquire and process data from related servers based on artificial intelligence technology. Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to achieve optimal results.
[0087] Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.
[0088] This method can separate the speech spectrum of a specific speaker from mixed speech without inputting the speech of the target speaker, thus simplifying the extraction process.
[0089] The embodiment of the present application also provides a speaker-specific voice extraction device based on speaker auxiliary information, which is used to perform any embodiment of the aforementioned speaker-specific voice extraction method based on speaker auxiliary information. Figure 3 , Figure 3 1 is a schematic block diagram of a speaker-specific speech extraction device 100 based on speaker auxiliary information provided in an embodiment of the present application.
[0090] Among them, Figure 3 As shown, the speaker-specific speech extraction device 100 based on speaker auxiliary information includes a framing unit 110, an active information acquisition unit 120, a logarithmic spectrum amplitude coefficient acquisition unit 130, an input data generation unit 140, a speaker classification unit 150 and a target spectrum acquisition unit 160.
[0091] The framing unit 110 is configured to obtain audio and video data to be identified, and to frame the audio and video data to be identified to obtain a framing result; wherein the framing result includes multiple frames of sub-audio and video data.
[0092] In this embodiment, the technical solution is described with the server as the execution subject. The server first obtains the audio and video data to be recognized corresponding to the multi-person mixed speaking scene. This audio and video data to be recognized can be collected by the microphone and camera of the user end and uploaded to the server.
[0093] The specific forms of the audio and video data to be identified include the following: one is audio data consisting entirely of audio data (i.e., pure audio data), and the other is audio and video data including both audio and video data. In both forms of audio data to be identified, there are periods of time when users (one or more users) speak, and there are also periods of time when no users speak.
[0094] In order to further extract the target speaker's voice from the audio and video data to be identified, it is necessary to first frame the audio and video data to be identified to obtain a frame result. Specifically, the audio data (i.e., voice signal) in the audio and video data to be identified can be framed based on a preset frame length. If the preset frame length is equal to 50ms, the audio and video data to be identified is divided into multiple 50ms long sub-audio and video data, and each 50ms long sub-audio and video data is recorded as a frame of sub-audio and video data. For example, if the audio and video data to be identified is recorded as s, it can be divided into s1, s2, ..., s nThese sub-audio and video data, namely s={s1,s2,……,s n It can be seen that based on the above frame processing, the audio and video data to be recognized can be effectively divided into multiple small segments, and then more refined speaker voice extraction can be performed in each small segment of audio and video data.
[0095] The activity information acquisition unit 120 is configured to acquire the sub-speaker activity information corresponding to each frame of sub-audio and video data in the frame segmentation result through silence detection to form speaker activity information.
[0096] In this embodiment, at this time, as long as the time period in which the user is speaking in the audio data to be identified is located, the speaker activity information can be accurately determined. That is, after completing the audio framing, the silence detection technology can be used to determine whether each frame of sub-audio and video data is silent data. For example, if there is a speaker speaking in the sub-audio and video data corresponding to s1, then the sub-speaker activity information p1 corresponding to s1 is 1; for example, if there is no speaker speaking in the sub-audio and video data corresponding to s2 (that is, silent data), then the sub-speaker activity information p2 corresponding to s2 is 0. And so on, the sub-audio and video data s can be determined based on the silence detection technology. t (s t Indicates whether the t-th frame of the audio and video data to be identified is silent data, thereby further determining the sub-speaker activity information corresponding to each frame of the sub-audio and video data. t =1 means that the speaker is speaking in the sub-audio and video data of the tth frame, p t = 0 means that there is no speaker in the t-th frame of sub-audio and video data. It can be seen that based on the silence detection technology, the sub-speaker activity information in each frame of sub-audio and video data can be quickly determined to serve as auxiliary information for subsequent analysis of each frame of sub-audio and video data.
[0097] The logarithmic spectrum amplitude coefficient obtaining unit 130 is configured to obtain the logarithmic spectrum amplitude coefficient of each frame of the sub-audio and video data in the frame division result.
[0098] In this embodiment, the audio and video data to be identified is recorded as s={s1, s2, ..., s n At this point, the logarithmic spectrum amplitude coefficient corresponding to each frame of sub-audio and video data in the audio and video data to be identified, s, can be obtained. Specifically, the 257-dimensional logarithmic spectrum amplitude coefficient corresponding to each frame of sub-audio and video data can be obtained. The relationship between signal frequency and energy is represented by a spectrum, and the logarithmic amplitude spectrum is one type of spectrum graph. The amplitude of each spectral line in the logarithmic amplitude spectrum is calculated logarithmically (20logA) with respect to the original amplitude A, so the unit of its vertical axis is dB (decibel). The purpose of this transformation is to raise the lower-amplitude components relative to the higher-amplitude components, so as to observe periodic signals hidden in low-amplitude noise.
[0099] For example, after obtaining the 257-dimensional logarithmic spectrum amplitude coefficient corresponding to s1, it is recorded as y1; and so on, s2, s3, and so on. n After all the corresponding data are converted into 257-dimensional logarithmic spectrum amplitude coefficients, the logarithmic spectrum amplitude coefficients of each frame of sub-audio and video data in the frame division result are obtained, and are recorded as y1, y2, ..., y n Among them, y1, y2, ..., y n The total is y={y1,y2,……,y n}.
[0100] The input data generating unit 140 is configured to generate input data of a corresponding frame of sub-audio and video data according to the logarithmic spectrum amplitude coefficient of each frame of sub-audio and video data.
[0101] In this embodiment, to generate input data for identifying the target speaker's spectrum based on the logarithmic spectral amplitude coefficients of each frame of sub-audio and video data, the logarithmic spectral amplitude coefficients of each frame of sub-audio and video data can be used as the input data alone, or together with other parameters such as the sub-speaker activity information stored in the sub-audio and video data. The input data obtained in this way can be used more accurately as input data for identifying the target speaker's spectrum.
[0102] In one embodiment, as a first implementation method for obtaining input data, the input data generating unit 140 is specifically configured to:
[0103] The logarithmic spectrum amplitude coefficient of each frame of sub-audio and video data is used as the input data of the corresponding frame of sub-audio and video data.
[0104] In this embodiment, the first method for composing the input data corresponding to each frame of sub-audio and video data is to directly use the logarithmic spectral amplitude coefficients of each frame of sub-audio and video data as the input data for that frame of sub-audio and video data, that is, without combining them with other parameters. This method of generating input data places greater emphasis on the impact of the logarithmic spectral amplitude coefficients themselves on the recognition results.
[0105] In one embodiment, as a second implementation method for obtaining input data, the input data generating unit 140 is specifically configured to:
[0106] The logarithmic spectrum amplitude coefficient of each frame of sub-audio and video data is connected with the sub-speaker activity information of the corresponding frame of sub-audio and video data through a connection function to obtain the input data of each frame of sub-audio and video data.
[0107] In this embodiment, a second method for assembling the input data corresponding to each frame of sub-audio and video data is to concatenate the logarithmic spectral amplitude coefficients of each frame of sub-audio and video data with the sub-speaker activity information of the corresponding frame of sub-audio and video data using a concatenation function. The concatenation function, also known as the Concatenate function, concatenates multiple strings into a single string. The input data obtained through this method includes comprehensive information from two dimensions: the logarithmic spectral amplitude coefficients and the sub-speaker activity information.
[0108] In one embodiment, as a third implementation method for obtaining input data, the input data generating unit 140 is specifically configured to:
[0109] Connecting the logarithmic spectrum amplitude coefficient of each frame of sub-audio and video data with the sub-speaker activity information of the corresponding frame of sub-audio and video data through a connection function to obtain the first sub-input data of each frame of sub-audio and video data;
[0110] The logarithmic spectrum amplitude coefficient of each frame of sub-audio and video data and the sub-speaker activity information of the corresponding frame of sub-audio and video data are input into a pre-trained auxiliary model to perform an operation to obtain second sub-input data;
[0111] The first sub-input data of each frame of sub-audio and video data is combined with the second sub-input data to obtain the input data of each frame of sub-audio and video data.
[0112] In this embodiment, the pre-trained auxiliary network is AuxiliaryNet, which has two fully connected layers, each with 50 nodes. One of the fully connected layers uses the ReLU function (i.e., linear rectifier function) as the activation function, and the other fully connected layer serves as the output layer and uses a linear activation function (e.g., one of the sigmoid function, tanh function, ReLU function, ELU function, and PReLU function). Based on the logarithmic spectrum amplitude coefficient of each frame of sub-audio and video data and the sub-speaker activity information of the corresponding frame of sub-audio and video data, a pre-trained auxiliary model is input for operation to obtain the second sub-input data. For details, please refer to the above operation expression (1).
[0113] After calculating the second sub-input data e based on the above formula (1), it can be combined with the first sub-input data of each frame of sub-audio and video data (the combination here is different from the function of connecting strings of the concatenation function, which can be to form two data into a row vector form) to obtain the input data of each frame of sub-audio and video data. For example, the first sub-input data of the i-th frame of sub-audio and video data is used to Concatenate(y i , p i ) means that, at this time, it is combined with e to obtain (Concatenate(y i , pi ), e). The input data obtained in the above manner includes comprehensive information in three dimensions.
[0114] The speaker classification unit 150 is configured to obtain a pre-trained speaker classification model, input the input data into the speaker classification model for classification, and obtain target speaker recognition features.
[0115] In this embodiment, the pre-trained speech classification network is an extraction network structure (ExtractionNet), and the input of the extraction network structure is the input data, specifically including a 3-layer Bi-LSTM (1200 units) mixed with a 2-layer fully connected layer (ReLU as the activation function).
[0116] Corresponding to the three ways of obtaining input data, the above three types of input data can all be input into the speaker classification model. The input data is input into the speaker classification model for classification to obtain target speaker recognition features.
[0117] For example, taking the first form of input data as input to the speaker classification model, one of the first form of input data is y i To express, with y i The corresponding target speaker recognition features are used Indicates that based on y i Get Please refer to the above formula (2). i Input it into the extraction network structure (ExtractionNet) for operation to obtain the same value as y i Corresponding target speaker recognition features
[0118] For example, taking the second form of input data as input to the speaker classification model, one of the second form of input data is Concatenate(y i , p i ) to express, and Concatenate(y i , p i ) The corresponding target speaker identification feature is used Indicates that, based on Concatenate(y i , p i )Get Please refer to the above formula (3). i , p i ) is input into the extraction network structure (ExtractionNet) for operation to obtain the same value as y i Corresponding target speaker recognition features
[0119] For example, taking the third form of input data as input to the speaker classification model, one of the third form of input data is (Concatenate(y i , p i ), e) to express, and (Concatenate(y i , p i ), e) the corresponding target speaker identification features are used Indicates that, based on (Concatenate(y i , p i ), e) obtain Please refer to the above formula (4). i , p i ), e) input into the extraction network structure (ExtractionNet) for operation to obtain the same value as y i Corresponding target speaker recognition features
[0120] The target speaker recognition features corresponding to each frame of sub-audio and video data can be extracted through the above three methods, which facilitates the subsequent extraction of the target speaker spectrum based on the target speaker recognition features.
[0121] In one embodiment, the apparatus 100 for extracting a specific speaker's speech based on speaker auxiliary information further includes:
[0122] The speaker classification model training unit is used to obtain a target speaker speech set, and perform model training on a speaker classification model to be trained using the target speaker speech set as a training set to obtain a speaker classification model.
[0123] In this embodiment, the ExtractionNet structure, or speaker classification model, is trained based on the target speaker's speech set. Its purpose is to extract only the specific speaker identification features from the audio and video data to be recognized in mixed speech scenes. Thus, when training the speaker classification model, the target speaker's speech set (i.e., the speech set of the specific speaker) is input for model training, ultimately resulting in a speaker classification model specifically designed to extract specific speaker identification features from the audio and video data to be recognized.
[0124] In order to allow the ExtractionNet structure to better learn the characteristics of a specific speaker, only non-overlapping speech is used for training. That is, each target speaker speech in the target speaker speech set is different.
[0125] In one embodiment, in order to improve the robustness of the speaker classification model, the target speaker speech set may be preprocessed. Therefore, the speaker-specific speech extraction apparatus 100 based on speaker auxiliary information further includes:
[0126] The initial speech set acquisition unit is used to obtain an initial target speaker speech set, and randomly increase or decrease the duration of each initial target speaker speech data in the initial target speaker speech set by a random duration to update each initial target speaker speech data to form a target speaker speech set; wherein the value range of the random duration is [0, 1s].
[0127] In this embodiment, taking one of the initial target speaker speech data in the initial target speaker speech set as an example, if its duration is 6 seconds, the initial target speaker speech data is updated by extending it forward by 1 second from its starting time point (if this 1 second is silent audio data) to obtain a 7-second duration initial target speaker speech data. Alternatively, the initial target speaker speech data is updated by extending it backward by 1 second from its ending time point (if this 1 second is silent audio data) to obtain a 7-second duration initial target speaker speech data. The target speaker speech set obtained in this way is more suitable for training a more robust speaker classification model.
[0128] The target spectrum acquisition unit 160 is configured to multiply the target speaker identification feature by the log spectrum amplitude coefficient of each frame of sub-audio and video data to obtain a target speaker spectrum corresponding to each frame of sub-audio and video data.
[0129] In this embodiment, for example, after obtaining the sub-audio and video data s i Corresponding target speaker recognition features At this time, the target speaker identification feature With sub-audio and video data i The corresponding logarithmic spectrum amplitude coefficient y i Multiply them together to get the sub-audio and video data s i Corresponding target speaker spectrum The above operation can refer to the above formula (5). That is, through the above operation, the target speaker's frequency spectrum corresponding to each frame of sub-audio and video data can be obtained.
[0130] In one embodiment, the apparatus 100 for extracting a specific speaker's speech based on speaker auxiliary information further includes:
[0131] The target speaker voice data restoration unit is used to sequentially perform inverse Fourier transformation and splicing on the target speaker spectrum corresponding to each frame of sub-audio and video data to obtain the target speaker voice data.
[0132] In this embodiment, the target speaker's speech data is obtained by performing an inverse Fourier transform on the target speaker's spectrum corresponding to each frame of the sub-audio video data, and then concatenating the frames to restore the speech. This method allows the target speaker's speech data to be extracted from mixed speech without pre-recording the target speaker's speech.
[0133] The device can separate the speech spectrum of a specific speaker from mixed speech without inputting the speech of the target speaker, thus simplifying the extraction process.
[0134] The above-mentioned speaker-specific speech extraction device based on speaker auxiliary information can be implemented in the form of a computer program. The computer program can be used in Figure 4 Runs on the computer equipment shown.
[0135] See also Figure 4 , Figure 4 1 is a schematic block diagram of a computer device provided in an embodiment of the present application. The computer device 500 is a server or a server cluster. The server can be a standalone server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.
[0136] See Figure 4 The computer device 500 includes a processor 502 , a memory, and a network interface 505 connected via a device bus 501 , wherein the memory may include a storage medium 503 and an internal memory 504 .
[0137] The storage medium 503 may store an operating system 5031 and a computer program 5032. When the computer program 5032 is executed, the processor 502 may execute a speaker-specific speech extraction method based on speaker auxiliary information.
[0138] The processor 502 is used to provide computing and control capabilities to support the operation of the entire computer device 500.
[0139] The internal memory 504 provides an environment for running the computer program 5032 in the storage medium 503 . When the computer program 5032 is executed by the processor 502 , the processor 502 can execute a speaker-specific speech extraction method based on speaker auxiliary information.
[0140] The network interface 505 is used for network communication, such as providing data information transmission. Those skilled in the art will understand that Figure 4The structure shown in the figure is merely a block diagram of a portion of the structure related to the solution of the present application, and does not constitute a limitation on the computer device 500 to which the solution of the present application is applied. The specific computer device 500 may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0141] The processor 502 is configured to run a computer program 5032 stored in the memory to implement the speaker-specific speech extraction method based on speaker auxiliary information disclosed in the embodiment of the present application.
[0142] Those skilled in the art will understand that Figure 4 The embodiment of the computer device shown in the figure does not constitute a limitation on the specific composition of the computer device. In other embodiments, the computer device may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently. For example, in some embodiments, the computer device may only include a memory and a processor. In such an embodiment, the structure and function of the memory and processor are the same as those in the figure. Figure 4 The embodiments shown are consistent and will not be described again here.
[0143] It should be understood that in the embodiment of the present application, the processor 502 may be a central processing unit (CPU), and the processor 502 may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.
[0144] In another embodiment of the present application, a computer-readable storage medium is provided. The computer-readable storage medium may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the speaker-specific speech extraction method based on speaker auxiliary information disclosed in an embodiment of the present application.
[0145] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described equipment, devices and units can refer to the corresponding processes in the aforementioned method embodiments, and will not be repeated here. Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented with electronic hardware, computer software or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0146] In the several embodiments provided in this application, it should be understood that the disclosed devices, apparatuses and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, or units with the same function may be combined into one unit. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection through some interfaces, devices or units, or may be an electrical, mechanical or other form of connection.
[0147] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the embodiments of the present application.
[0148] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0149] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a background server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk or an optical disk.
[0150] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present application, and such modifications or substitutions should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A speaker-specific speech extraction method based on speaker auxiliary information, characterized in that: include: Acquire audio and video data to be identified, and divide the audio and video data to be identified into frames to obtain a frame division result; wherein the frame division result includes multiple frames of sub-audio and video data; Acquire sub-speaker activity information corresponding to each frame of sub-audio and video data in the frame segmentation result through silence detection to form speaker activity information; Obtaining the logarithmic spectrum amplitude coefficient of each frame of sub-audio and video data in the framing result; Generate input data of corresponding frame sub-audio and video data according to the logarithmic spectrum amplitude coefficient of each frame sub-audio and video data; Obtaining a target speaker speech set, and using the target speaker speech set as a training set to perform model training on a speaker classification model to obtain a speaker classification model; wherein each target speaker speech in the target speaker speech set is different; Obtaining a pre-trained speaker classification model, inputting the input data into the speaker classification model for classification, and obtaining target speaker recognition features; and Multiplying the target speaker identification feature by the logarithmic spectrum amplitude coefficient of each frame of sub-audio and video data to obtain the target speaker spectrum corresponding to each frame of sub-audio and video data; The step of generating input data of the corresponding frame of sub-audio and video data according to the logarithmic spectrum amplitude coefficient of each frame of sub-audio and video data includes: The logarithmic spectrum amplitude coefficient of each frame of sub-audio and video data is connected with the sub-speaker activity information of the corresponding frame of sub-audio and video data through a connection function to obtain the input data of each frame of sub-audio and video data.
2. The speaker-specific speech extraction method based on speaker auxiliary information according to claim 1, characterized in that: The step of generating input data of the corresponding frame of sub-audio and video data according to the logarithmic spectrum amplitude coefficient of each frame of sub-audio and video data includes: The logarithmic spectrum amplitude coefficient of each frame of sub-audio and video data is used as the input data of the corresponding frame of sub-audio and video data.
3. The speaker-specific speech extraction method based on speaker auxiliary information according to claim 1, characterized in that: The step of generating input data of the corresponding frame of sub-audio and video data according to the logarithmic spectrum amplitude coefficient of each frame of sub-audio and video data includes: Connecting the logarithmic spectrum amplitude coefficient of each frame of sub-audio and video data with the sub-speaker activity information of the corresponding frame of sub-audio and video data through a connection function to obtain the first sub-input data of each frame of sub-audio and video data; The logarithmic spectrum amplitude coefficient of each frame of sub-audio and video data and the sub-speaker activity information of the corresponding frame of sub-audio and video data are input into a pre-trained auxiliary model to perform an operation to obtain second sub-input data; The first sub-input data of each frame of sub-audio and video data is combined with the second sub-input data to obtain the input data of each frame of sub-audio and video data.
4. The speaker-specific speech extraction method based on speaker auxiliary information according to claim 1, characterized in that: Before acquiring the target speaker's speech set, the method further includes: An initial target speaker speech set is obtained, and a duration of each initial target speaker speech data in the initial target speaker speech set is randomly increased or decreased by a random duration to update each initial target speaker speech data to form a target speaker speech set; wherein the random duration has a value range of [0, 1s].
5. The speaker-specific speech extraction method based on speaker auxiliary information according to claim 1, characterized in that: After multiplying the target speaker identification feature by the logarithmic spectrum amplitude coefficient of each frame of sub-audio and video data to obtain the target speaker spectrum corresponding to each frame of sub-audio and video data, the method further includes: The target speaker's spectrum corresponding to each frame of sub-audio and video data is sequentially subjected to inverse Fourier transform and spliced to obtain the target speaker's speech data.
6. A speaker-specific speech extraction device based on speaker auxiliary information, characterized in that: include: A framing unit is configured to obtain audio and video data to be identified, and to frame the audio and video data to be identified to obtain a framing result; wherein the framing result includes multiple frames of sub-audio and video data; an activity information acquisition unit, configured to acquire, by silence detection, sub-speaker activity information corresponding to each frame of sub-audio and video data in the frame segmentation result, to form speaker activity information; a logarithmic spectrum amplitude coefficient acquisition unit, configured to acquire the logarithmic spectrum amplitude coefficient of each frame of sub-audio and video data in the framing result; An input data generating unit, configured to generate input data of a corresponding frame of sub-audio and video data according to a logarithmic spectrum amplitude coefficient of each frame of sub-audio and video data; a speaker classification model training unit, configured to obtain a target speaker speech set, and perform model training on a speaker classification model to be trained using the target speaker speech set as a training set, thereby obtaining a speaker classification model; wherein each target speaker speech in the target speaker speech set is different; a speaker classification unit, configured to obtain a pre-trained speaker classification model, input the input data into the speaker classification model for classification, and obtain target speaker recognition features; and a target spectrum acquisition unit, configured to multiply the target speaker identification feature by the logarithmic spectrum amplitude coefficient of each frame of sub-audio and video data to obtain a target speaker spectrum corresponding to each frame of sub-audio and video data; The input data generating unit is specifically used for: The logarithmic spectrum amplitude coefficient of each frame of sub-audio and video data is connected with the sub-speaker activity information of the corresponding frame of sub-audio and video data through a connection function to obtain the input data of each frame of sub-audio and video data.
7. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the method for extracting a specific speaker's speech based on speaker auxiliary information according to any one of claims 1 to 5 is implemented.
8. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, causes the processor to perform the speaker-specific speech extraction method based on speaker auxiliary information according to any one of claims 1 to 5.
Citation Information
Patent Citations
Multi-mode voice separation method, training method and related device
CN113782048A
Training method of speaker separation model, speaker separation method and related device
CN114360573A
Single-channel and multi-channel source separation enhanced by lip motion
US20210217182A1