Voice processing method, device, electronic device and medium
By converting the frame-level acoustic features of the speech data into character-level semantic and vocalprint features, the speaker conversion in the speech data is detected, and the problem of speech recognition accuracy decreases when multiple people speak alternately or simultaneously is solved, and a simplified speech recognition process and improved recognition accuracy are achieved.
Patent Information
- Application Number
- CN202210939799.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-05
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2042-08-05
AI Technical Summary
Existing voice recognition technology is difficult to effectively deal with the situation where multiple people speak alternately or speak simultaneously, resulting in a decrease in speech recognition accuracy.
The speaker conversion is detected at the character level by converting the frame-level acoustic features of the target speech data into character-level semantic features and vocalprint features and determining the speaker-transformed characters based on these features.
It realizes that the speech recognition results based on the speaker are directly output without post-processing, simplifying the speech recognition process and improving the accuracy and efficiency of speech recognition.
Smart Images

Figure CN115273862B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the field of speech processing technology, and more specifically, to a speech processing method, an apparatus, an electronic device, a computer-readable storage medium, and a computer program product. Background Art
[0002] With the rapid development of the Internet and artificial intelligence (AI) technology, automatic speech recognition (ASR) has brought great convenience to people's lives. In some scenarios (for example, remote meetings, remote teaching), there is a need to collect and organize voice content, and it is hoped that speech recognition will be performed according to the speaker role. However, there may be times when multiple people speak alternately or simultaneously, which brings challenges to speech recognition.
[0003] Speech Conversion Detect (SCD) aims to locate the time when different speakers start speaking. SCD systems are usually used as a submodule of speaker segmentation and clustering, or as the front end of speech recognition tasks to segment long speech. The performance of the SCD system will greatly affect subsequent processing tasks. Summary of the invention
[0004] In view of this, an embodiment of the present disclosure proposes a technical solution for speech processing.
[0005] According to a first aspect of the present disclosure, a method for speech processing is provided. The method comprises: generating character-level semantic features of target speech data based on frame-level acoustic features of target speech data; generating character-level voiceprint features of target speech data based on frame-level acoustic features; and determining characters in the target speech data where speaker switching occurs based on character-level semantic features and character-level voiceprint features.
[0006] According to the embodiments of the present disclosure, the speaker conversion in the speech data is detected at the character level in combination with the speaker's acoustic features and speech content, and the speaker-based speech recognition result can be directly output without post-processing, thereby simplifying the speech recognition process.
[0007] According to a second aspect of the present disclosure, a speech processing device is provided. The device includes a semantic feature generation unit, a voiceprint feature generation unit, and a detection unit. The semantic feature generation unit is configured to generate character-level semantic features of the target speech data based on the frame-level acoustic features of the target speech data. The voiceprint feature generation unit is configured to generate character-level voiceprint features of the target speech data based on the frame-level acoustic features. The detection unit is configured to determine the characters in the target speech data where speaker switching occurs based on the character-level semantic features and the character-level voiceprint features.
[0008] According to a third aspect of the present disclosure, an electronic device is provided, comprising: at least one processing unit; and at least one memory, wherein the at least one memory is coupled to the at least one processing unit and stores instructions for execution by the at least one processing unit, and when the instructions are executed by the at least one processing unit, the device executes the method according to the first aspect of the present disclosure.
[0009] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided, comprising machine-executable instructions, which, when executed by a device, cause the device to perform the method according to the first aspect of the present disclosure.
[0010] According to a fifth aspect of the present disclosure, there is provided a computer program product, comprising machine executable instructions, which, when executed by a device, cause the device to perform the method according to the first aspect of the present disclosure.
[0011] This summary is provided to introduce a selection of concepts in a simplified form, which will be further described in the following detailed description. This summary is not intended to identify key features or essential features of the present disclosure, nor is it intended to limit the scope of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The above and other objects, features and advantages of the present disclosure will become more apparent through a more detailed description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, wherein like reference numerals generally represent like components throughout the exemplary embodiments of the present disclosure.
[0013] Figure 1 A schematic diagram illustrating an example environment in which various embodiments of the present disclosure can be implemented;
[0014] Figure 2 An overall schematic diagram of a process of detecting speaker switching in target speech data according to an embodiment of the present disclosure is shown;
[0015] Figure 3 A schematic flow chart of a speech processing method according to an embodiment of the present disclosure is shown;
[0016] Figure 4A schematic diagram showing the structure of a semantic feature model according to an embodiment of the present disclosure;
[0017] Figure 5 A schematic diagram showing the structure of a voiceprint feature model according to an embodiment of the present disclosure;
[0018] Figure 6 A schematic diagram showing the structure of a speaker switch detection model according to an embodiment of the present disclosure;
[0019] Figure 7 A schematic block diagram showing a speech processing device according to an embodiment of the present disclosure; and
[0020] Figure 8 A schematic block diagram of an example device that may be used to implement embodiments of the present disclosure is shown. DETAILED DESCRIPTION
[0021] It is understandable that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, scope of use, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0022] For example, in response to receiving an active request from a user, a prompt message is sent to the user to clearly prompt the user that the operation requested to be performed will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, application, server, or storage medium that performs the operation of the technical solution of the present disclosure according to the prompt message.
[0023] As an optional but non-limiting implementation, in response to receiving an active request from the user, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. In addition, the pop-up window may also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0024] It is understandable that the above notification and the process of obtaining user authorization are merely illustrative and do not constitute a limitation on the implementation of the present disclosure. Other methods that meet the relevant laws and regulations may also be applied to the implementation of the present disclosure.
[0025] The preferred embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although the preferred embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided to make the present disclosure more thorough and complete, and to fully convey the scope of the present disclosure to those skilled in the art.
[0026] As used herein, the term "including" and its variations mean open inclusion, i.e., "including but not limited to". Unless otherwise stated, the term "or" means "and / or". The term "based on" means "based at least in part on". The terms "an example embodiment" and "an embodiment" mean "at least one example embodiment". The term "another embodiment" means "at least one additional embodiment". The terms "first", "second", etc. may refer to different or the same objects. Other explicit and implicit definitions may also be included below.
[0027] It should be noted that any numerical values or numbers used in the present disclosure are exemplary and are not intended to limit the scope of the present disclosure.
[0028] As mentioned above, the performance of the speaker detection (SCD) system will greatly affect the subsequent processing tasks of speech processing. Some traditional methods use distance-based methods. This type of method divides long speech into fixed lengths, and then calculates the distance between the voiceprint features in adjacent segments. Once the distance exceeds the threshold, it is determined that the speaker has switched between the two segments. However, the detection accuracy of this method is limited by the segmentation length of the speech segment, and it is impossible to detect the switch when the speaker switches quickly. There are also some end-to-end methods that use neural network models to directly predict speaker switching without relying on distance metrics. However, this method predicts speaker switching at the speech frame level, has a strong reliance on the annotation of speech data, and requires subsequent speech-to-text recognition processing, and the processing process is complicated.
[0029] In view of this, an embodiment of the present disclosure provides a method for speech processing. In the method, the frame-level acoustic features of the target speech data are converted into the character (token) level semantic features of the target speech data. The frame-level acoustic features may be in the form of an acoustic feature sequence, wherein each acoustic feature corresponds to a speech frame in the speech data, and the character-level semantic features may be in the form of a semantic feature sequence, wherein each semantic feature corresponds to a character in the speech data. In this article, multiple speech frames may be aggregated together to correspond to a character. In the method, character-level voiceprint features of the target speech data are also generated based on the frame-level acoustic features. The character-level voiceprint features may be in the form of a voiceprint feature sequence, wherein each voiceprint feature corresponds to a character in the speech data. In the method, the characters in the target speech data where the speaker conversion occurs are determined based on the character-level semantic features and the character-level voiceprint features. According to the embodiment of the present disclosure, the speaker conversion in the speech data is detected at the character level in combination with the acoustic features and speech content of the speaker, and the speech recognition result based on the speaker can be directly output without post-processing, thereby simplifying the speech recognition process.
[0030] The following references Figures 1 to 8 Detailed description of implementation details of the embodiments of the present disclosure.
[0031] Figure 1 1 is a schematic diagram of an example environment 100 in which various embodiments of the present disclosure can be implemented. Figure 1 As shown, the system architecture 100 may include terminal devices 1011, 1012, 1013, a network 102, and a server 103. The network 102 is used to provide a medium for communication links between the terminal devices 1011, 1012, 1013 and the server 103. The network 102 may include various connection types, such as wired, wireless communication links, or optical fiber cables.
[0032] Terminal devices 1011, 1012, 1013 may be hardware, software, or a combination of hardware and software. When terminal devices 1011, 1012, 1013 are hardware, they may be electronic devices with computing capabilities, including but not limited to smart phones, tablet computers, laptop computers, etc., desktop computers, etc. When terminal devices 1011, 1012, 1013 are software, they may be installed in the electronic devices listed above. They may be implemented as multiple software or software modules (e.g., multiple software or software modules with speech processing or speech recognition capabilities), or they may be implemented as a single software or software module, which is not limited in the present disclosure.
[0033] The terminal devices 1011, 1012, and 1013 may also obtain a speech processing model from the server 103. The terminal devices 1011, 1012, and 1013 may collect speech data in real time via a microphone, receive speech data from other devices, or read stored speech data as target speech data, and then perform a speech processing process according to an embodiment of the present disclosure on the target speech data.
[0034] Alternatively, terminal devices 1011, 1012, 1013 can interact with server 103 through network 102 to send or receive data, etc. For example, server 103 can receive real-time or pre-collected voice data sent by terminal devices 1011, 1012, 1013, and terminal devices 1011, 1012, 1013 can also receive voice processing results output by server 103.
[0035] The server 103 may be hardware, software, or a combination of hardware and software. When the server 103 is hardware, it may be implemented as a distributed server cluster consisting of multiple servers, or it may be implemented as a single server. When the server 103 is software, it may be implemented as multiple software or software modules (e.g., multiple software or software modules with speech processing or speech recognition capabilities), or it may be implemented as a single software or software module, and the present disclosure does not limit this.
[0036] The server 103 may be a server that processes the target voice data to be recognized received from the terminal devices 1011, 1012, 1013. The server 103 may receive the target voice data from the terminal devices 1011, 1012, 1013, and then perform the voice processing process according to the embodiment of the present disclosure on the target voice data. For example, the server 103 may output the voice processing result to the terminal devices 1011, 1012, 1013.
[0037] It should be noted that, when the speech processing process provided in the embodiments of the present disclosure is performed by the terminal devices 1011, 1012, and 1013, if the terminal devices 1011, 1012, and 1013 have pre-trained speech processing models stored locally, the exemplary system architecture 100 may not have the network 102 and the server 103. The speech processing model may be trained by the server 103 and sent to the terminal devices 1011, 1012, and 1013, or it may be trained locally by any of the terminal devices 1011, 1012, and 1013.
[0038] in addition, Figure 1 The number of terminal devices, networks and servers in the embodiment is only for illustration. Any number of terminal devices, networks and servers may be provided according to implementation requirements.
[0039] References Figure 1 An exemplary environment in which embodiments of the present disclosure can be implemented is described. It should be understood that Figure 1 This is merely illustrative, and the environment may include more modules or components, or some modules or components may be omitted, or the modules or components shown may be recombined. Figure 1 The present disclosure is not limited to the different environments shown.
[0040] Figure 2 A general schematic diagram of a process of detecting speaker switching in target speech data according to an embodiment of the present disclosure is shown.
[0041] In this article, the interactive voice data with multiple speakers for speaker conversion detection using the embodiments of the present disclosure is referred to as target voice data. In addition, the present disclosure does not limit the language type of the target voice data, and the target voice data can be Chinese voice data, English voice data, or other types of voice data. The present disclosure also does not limit the source of the target voice data, and the target voice data can be pre-collected voice data, or voice data collected in real time by a terminal device, or voice data received via a network.
[0042] As shown, the process 200 of detecting speaker switching involves a speech processing model 210 and an acoustic feature extraction model 250, both of which can be implemented or deployed in a system such as Figure 1 Any of the terminal devices 1011, 1012, 1013, or the server 103 shown.
[0043] The target speech data 201 is provided as input to the acoustic feature extraction unit 250. The acoustic feature extraction unit 250 may output the frame-level acoustic features x=(x 1 ,x 2 ,…,x T ), where T is the number of speech frames, x i (i=1, 2, ...T) is the acoustic feature of any speech frame. Specifically, when extracting the acoustic features of the target speech data, the acoustic features can be extracted in the following manner. First, the target speech data needs to be framed to obtain the corresponding speech frame sequence, and then the framed speech frame sequence is pre-emphasized, and then the acoustic features of each speech frame are extracted in turn. The acoustic features include feature data used to characterize the acoustic information of the corresponding speech frame, for example, it can be Fbank features (Filter Bank), Mel scale Frequency Cepstral Coefficients (MFCC) features, or Perceptual Linear Predictive (PLP) features, etc. The acoustic feature x of each speech frame i It can be expressed in the form of a multi-dimensional vector, and thus the frame-level acoustic feature x output from the acoustic feature extraction unit 230 can be expressed in the form of an acoustic feature sequence (eg, a matrix).
[0044] The frame-level acoustic feature x is provided to the speech processing model 210 according to an embodiment of the present disclosure. The speech processing model 210 can be implemented as a neural network model, which is used to perform speaker switching detection on the frame-level acoustic feature x after training, and output a character-level processing result p=(p 1 ,p 2 ,…,p s ), where S is the number of characters, p i (i=1, 2, ...S) is the detection result for any character in the target speech data. In this article, a character may include at least one of the following: a word, a subword, a letter, a syllable, or a phoneme. Multiple speech frames can be aggregated together to correspond to one character. As an example but not a limitation, the speech processing model 210 performs character-level binary classification on the target speech data 201, and the detection result p for each character is iIt can be, for example, "0" or "1", where "0" indicates that no speaker switch occurs at the current character, and "1" indicates that a speaker switch occurs starting from the current character. It should be noted that, in this article, speaker switch includes a switch from one speaker to another speaker, a switch from one speaker to multiple speakers (i.e., speakers speak at the same time), a switch from multiple speakers to one speaker, a switch from multiple speakers to different multiple speakers (at least one speaker is different), etc.
[0045] In general, the speech processing model 210 uses two clues, the speaker's voiceprint features (also called speaker features or speaker representation) and the speech content, to detect speaker switching. As shown in the figure, the speech processing model 210 includes a semantic feature model 220, a voiceprint feature model 230, and a speaker detection model 240. The speech feature model 220 receives the frame-level acoustic feature x = (x 1 ,x 2 ,…,x T ) as input, and outputs character-level semantic features (not shown) of the target speech data 201. The voiceprint feature model 230 also receives the frame-level acoustic features x=(x 1 ,x 2 ,…,x T ) as input, and outputs the character-level voiceprint features (not shown) of the target speech data 201. However, the character-level semantic features and the character-level voiceprint features are provided to the speaker switch detection model 240, and the speaker switch detection model 240 generates a detection result p=(p 1 ,p 2 ,…,p s ).
[0046] Next reference Figures 3 to 6 The speech processing process according to the embodiment of the present disclosure is described in detail.
[0047] Figure 3 300 is a schematic flow chart of a speech processing method according to an embodiment of the present disclosure. Figure 1 The method 300 may be implemented by any one of the terminal devices 1011, 1012, 1013 or the server 103. Figure 2 It should be understood that the method 300 may also include additional actions not shown and / or may omit the actions shown, and the scope of the present disclosure is not limited in this respect. Figure 2 The method 300 is described in detail. Figure 2 Give a description.
[0048] In block 310, character-level semantic features of the target speech data are generated based on the frame-level acoustic features of the target speech data. The actions shown in block 310 can be performed by, for example, Figure 2 It is implemented by the semantic feature model 220 in .
[0049] Figure 4 FIG. 2 is a schematic diagram showing the structure of the semantic feature model 220 according to an embodiment of the present disclosure. Figure 4 Describes the details of generating character-level semantic features.
[0050] exist Figure 4 In the example, the semantic feature model 220 receives the frame-level acoustic feature x=(x 1 ,x 2 ,…,x T ). Frame-level acoustic feature x = (x 1 ,x 2 ,…,x T ) is provided to the semantic encoder 221. The semantic encoder 221 may be, for example, a decoder based on a stacked-conformer encoder or other suitable structures, which extracts the input frame-level acoustic features x in high dimensions to obtain a frame-level semantic coding feature h = (h 1 ,h 2 ...h U ), where U is the number of extracted speech frames, h i (i=1, 2, ...U) is the semantic coding feature for a speech frame. The frame-level semantic coding feature h also has the form of a feature sequence.
[0051] The frame-level semantic coding feature h is provided as input to the weight estimator 222. The weight estimator 222 can generate a set of weights α=(α 1 ,α 2 ...α U ), where each weight corresponds frame by frame to the semantic coding feature h for one frame in the frame-level semantic coding feature h i In some embodiments, the weight estimator 222 may be implemented as a network model and may be trained to generate a different set of weights α for different h. As an example, the weight estimator 222 may include a convolutional neural network (CNN) and a fully connected network. The weight estimator 222 may also include other neural networks with time series modeling capabilities, such as a recurrent neural network (RNN).
[0052] The weight α is provided to a continuous integration-and-fire (CIF) model 223. CIF 223 is used to integrate and fire the frame-level semantic coding features h accumulated frame by frame in the target speech data to determine the character-level semantic coding features with characters as boundaries in the target speech data. Using CIF 223, a character-level semantic coding feature c=(c 1 ,c 2 ...c S ), where S represents the number of characters, c i (i=1, 2, ...S) is the semantic coding feature of each character. Character level semantic coding feature c=(c 1 ,c 2 ...c S ) has the form of a feature sequence and can be further used to generate the output of the semantic feature model 220.
[0053] Specifically, the frame-level semantic coding features can be divided based on the comparison between the accumulated value of the continuous weights in the weight α and the threshold β. When the accumulated value is greater than the threshold β, it is determined that byte decomposition exists, and the frame-level semantic coding features are divided here.
[0054] As an example, assuming α = (0.1, 0.5, 0.6, 0.3, 0.6, 0.5, 0.2, 0.1, 0.4, 0.5, 0.2, ...), β = 1.0, the above weights correspond to the frame-level semantic coding features h i (i=1, 2, ...). It can be seen that the cumulative value of the first three weights (0.1+0.5+0.6=1.2) is greater than the threshold 1.0. Therefore, the semantic coding feature h in the third frame can be 3 After that, the character boundary is determined. Furthermore, the part of the weight accumulation value that exceeds the threshold can be retained and used to define the next character. For example, the part of the accumulated value of the first three weights 1.2 that exceeds the threshold 1.0 is 0.2, and 0.2 can be accumulated with the candidate weight to determine the boundary of the second character. It can be seen that the accumulated value of the excess part and the fourth and fifth weights (0.2+0.3+0.6=1.1) is greater than the threshold 1.0. Therefore, in the semantic coding feature h of the fifth frame, 3 After that, the character boundaries are determined, and so on. Thus, the frame-level semantic coding feature h is divided into the character level. Figure 4 As shown, the weight α is also provided to the voiceprint feature model 230 for dividing the frame-level voiceprint features, that is, the weight α is shared between the semantic feature model 220 and the voiceprint feature model 230. For details, please refer to Figure 5 describe.
[0055] Continue to refer Figure 4 , CIF 223 is based on the divided frame level semantic coding feature h = (h 1 ,h 2 ...h U ) and weight α=(α 1 ,α 2 ...α U ), generate character-level semantic encoding features c = (c 1 ,c 2 ...c S ), where S is the number of characters, c i is the semantic coding feature for a character of the target speech data. In some embodiments, the frame-level semantic coding feature h and the weight α=(α 1 ,α 2 ...α U ) is weighted summed to obtain the character-level semantic coding feature c = (c 1 ,c 2 ...c S ). Continuing with the above example, c 1 =α 1 *h1+α 2 *h2+α 3 *h 3 , c 2 =α 4 *h4+α 5 *h 5 , ..., and so on. In some embodiments, the portion exceeding the threshold β may be passed to the next character, for example, c 2 =(α 1 +α 2 +α 3 +α 4 -β)*h4+α 5 *h 5 , and so on.
[0056] Character level semantic encoding feature c = (c 1 ,c 2 ...c S ) can be provided to the semantic decoder 224. The semantic decoder 224 can be, for example, a decoder based on a stacked transformer (Stacked-Transformer Decoder) or other suitable structures. The semantic decoder 224 encodes the character-level semantic features c=(c 1 ,c 2 ...c S ) recursively decodes each character to obtain the character-level semantic decoding feature o = (o1 ,o 2 ...o S In some embodiments, the character-level semantic decoding feature o=(o 1 ,o 2 ...o S ) and character-level semantic encoding feature c=(c 1 ,c 2 ...c S ) are concatenated 226 to generate character-level semantic features [c; o] as outputs of the semantic feature model 220. The character-level semantic features [c; o] may be in the form of a feature sequence.
[0057] like Figure 4 As shown, the character level semantic decoding feature o = (o 1 ,o 2 ...o S ) can be provided to a Softmax layer (e.g., a dense softmax or fully connected layer) to obtain a speech-to-text recognition result y = (y 1 ,y 2 ...y S It should be noted that the recognition result y = (y 1 ,y 2 ...y S ) does not participate in the process of generating character-level semantic features [c; o], but is used in the training process to adjust the parameters of each model in the entire semantic feature model 220. According to the embodiment of the present disclosure, as the frame-level acoustic feature x is input into the semantic feature model 210, the speech-to-text recognition result y=(y 1 ,y 2 ...y S ) and character-level semantic features [c; o], therefore, speaker switching detection and speech-to-text recognition can be achieved simultaneously without additional subsequent processing. It should be noted that the character-level semantic features output by the semantic feature model 210 may also have other forms, for example, it may be a character-level semantic decoding feature o=(o 1 ,o 2 ...o S ), or character-level semantic encoding feature c=(c 1 ,c 2 ...c S ), or any combination thereof.
[0058] Continue to refer Figure 3 In block 320, character-level voiceprint features of the target speech data are generated based on the frame-level acoustic features. The actions shown in block 320 can be performed by, for example, Figure 2It is implemented by the semantic feature model 220 in .
[0059] Figure 5 FIG. 2 is a schematic diagram showing the structure of the voiceprint feature model 230 according to an embodiment of the present disclosure. Figure 5 Describes the details of generating character-level voiceprint features.
[0060] As shown in the figure, the voiceprint feature model 230 receives the frame-level acoustic feature x=(x 1 ,x 2 ,…,x T ). Frame-level acoustic feature x = (x 1 ,x 2 ,…,x T ) is provided to the voiceprint encoder 231. The voiceprint encoder 231 may have, for example, a ResNet18 structure or other suitable structures, which performs high-dimensional extraction on the input frame-level acoustic feature x to obtain a frame-level voiceprint encoding feature z=(z 1 ,z 2 ...z U ), where U is the number of extracted speech frames, z i (i=1, 2, ...U) is the voiceprint coding feature for a speech frame. The frame-level voiceprint coding feature z also has the form of a feature sequence and corresponds to the frame-level semantic coding feature h, and the number of frames U of the two is the same.
[0061] The frame-level voiceprint coded feature z is provided to another CIF 233. The CIF 233 also receives a set of weights α from the weight estimator 222 of the semantic recognition model 220. As mentioned above, the voiceprint feature model 230 and the semantic feature model 220 share a set of weights α. Using the CIF 233, the frame-level voiceprint coded feature z=(z 1 ,z 2 ...z U ) and a set of weights α to generate character-level voiceprint features. Specifically, similar to CIF 223 in semantic feature model 220, CIF 233 can divide frame-level voiceprint coding features based on the comparison between the accumulated value of continuous weights in a set of weights α and the threshold. Then, based on the divided frame-level voiceprint coding features z=(z 1 ,z 2 ...z U ) and weight α to generate character-level voiceprint feature e = (e 1 ,e 2 ...e S ), where S represents the number of characters, and each component e i(i=1, 2, ...S) represents the voiceprint coding feature for a character. For example, it can be done by weighted summation, which is similar to CIF 223 in the semantic feature model 220 and will not be repeated here. Thus, the conversion of the frame-level voiceprint representation z to the character-level voiceprint feature e is completed. Based on this approach, sharing the weights from the semantic feature model 220 can ensure that the length of the output e sequence is strictly consistent with the length of c, and the features in each sequence correspond to each other and correspond to the characters in the target speech data.
[0062] Character level voiceprint feature e=(e 1 ,e 2 ...e S ) is then provided to the voiceprint decoder 234. The voiceprint decoder 234 can be implemented as a fully connected structure for character-level voiceprint features e=(e 1 ,e 2 ...e S ) is decoded for each character-level voiceprint feature, thereby obtaining the classification result v = (v 1 ,v 2 ...v S ), where each component v i (i=1, 2, ...S) represents the probability of speaker classification for a character. In some embodiments, the hidden layer output m in the voiceprint decoder 234 can be retained and used as a character-level voiceprint feature provided to the speaker detection model 240. The character-level voiceprint feature can be in the form of a feature sequence. It should be noted that the classification result v=(v 1 ,v 2 ...v S ) does not participate in the process of generating character-level voiceprint features m, but it is used in the training process to adjust the parameters of each model in the entire speech processing model 210.
[0063] Continue to refer Figure 3 In block 330, based on the character-level semantic features and the character-level voiceprint features, the characters in the target speech data where the speaker switch occurs are determined. The actions shown in block 330 can be performed by, for example, Figure 2 The speaker switch detection model 240 in FIG.
[0064] Figure 6 FIG. 2 is a schematic diagram showing the structure of the speaker switch detection model 240 according to an embodiment of the present disclosure. Figure 6 Describes the details of the character on which the detected speaker switch occurred.
[0065] As shown in the figure, the speaker switch detection model 240 includes a speech content extraction model 241 for receiving and processing character-level speech features [c; o]. The speech content extraction model 241 includes a fully connected structure 242 and a Transformer structure 243. After the character-level speech features [c; o] are processed, the speech content representation l of the target speech data 201 is obtained. The speech content representation l is a feature sequence at the character level.
[0066] The speaker switch detection model 240 also includes a voiceprint difference extraction model 245 for receiving and processing character-level acoustic features m. The voiceprint difference extraction model 245 includes a series of convolutional layers 246 and a feed-forward network (FFN) 247. The voiceprint difference extraction model 245 is used to capture the difference d between each character-level voiceprint feature and its adjacent voiceprint features before and after it. Depending on the size of the convolution kernel of the convolutional layer 246, the meaning of "adjacent" includes closely adjacent and separated by a certain distance (for example, separated by one, two or more character positions).
[0067] Then, the voiceprint difference representation d and the speech content representation l of the voiceprint feature can be concatenated 248 and provided to the combiner 249. The combiner 249 can adopt a fully connected structure or any other suitable structure to provide a binary classification detection result for each character in the concatenated character level representation [l; d]. The combiner 249 can output a character level prediction result p = (p 1 ,p 2 ,…,p S ), where S is the number of characters in the target speech data, p i (i=1, 2, . . . S) is the detection result for any character in the target voice data.
[0068] The structure of the speech processing model 210 and the corresponding speech processing process 300 according to the embodiment of the present disclosure have been described. After the speech processing model 210 can be trained, the target speech data 201 is received and the above process 300 is implemented. The embodiment of the present disclosure also provides an exemplary method for training the speech processing model 210.
[0069] In some embodiments, the training of the speech processing model 210 can be performed in two stages. In the first stage, the parameters of the semantic feature model 220 and the voiceprint feature model 230 are pre-trained independently. In the second stage, the semantic feature model 220 and the voiceprint feature model 230 are pre-loaded with the parameters obtained by training in the first stage, and then the parameters of the speaker switching detection model 240 are randomly initialized, and then the parameters of the ASR part are fixed, and the parameters of the voiceprint feature model 230 and the speaker switching detection model 240 are jointly optimized.
[0070] Reference above Figures 1 to 6A speech processing method or process according to an embodiment of the present disclosure is described. Compared with existing solutions, the embodiment of the present disclosure can combine the acoustic features and speech content of the speaker to detect speaker conversion in the speech data at the character level, and can directly output the speaker-based speech recognition results without post-processing, thereby simplifying the speech recognition process. In some embodiments, the frame-level acoustic features are respectively integrated into character-level semantic features carrying speech content and voiceprint features carrying speaker information, and the two are aligned with each other. Therefore, through a relatively simple model structure, the two clues of speaker features and speech content can be used to detect speaker conversion. On the other hand, the process of speaker conversion detection and the speech recognition process can be carried out simultaneously, thereby simplifying the process of sorting out the speech recognition results in the later stage.
[0071] Figure 7 FIG. 7 is a schematic block diagram of a speech processing apparatus 700 according to an embodiment of the present disclosure. The apparatus 700 may be implemented in Figure 1 At any terminal device 1011, 1012, 1013 or server 103 shown.
[0072] As shown in the figure, the apparatus 700 includes a semantic feature generation unit 710, a voiceprint feature generation unit 720, and a detection unit 730. The semantic feature generation unit 710 is configured to generate character-level semantic features of the target speech data based on the frame-level acoustic features of the target speech data. The voiceprint feature generation unit 720 is configured to generate character-level voiceprint features of the target speech data based on the frame-level acoustic features. The detection unit 730 is configured to determine the characters in the target speech data where speaker switching occurs based on the character-level semantic features and the character-level voiceprint features.
[0073] In some embodiments, the semantic feature generation unit 710 can also be configured to: semantically encode the frame-level acoustic features to obtain frame-level semantic coding features; based on the frame-level semantic coding features, generate a set of weights, wherein the weights in the set of weights correspond to the semantic coding features for the frame in the frame-level semantic coding features frame by frame; and generate character-level semantic features based on a set of weights and the frame-level semantic coding features.
[0074] In some embodiments, the semantic feature generation unit 710 can also be configured to: divide the frame level semantic coding features based on the comparison of the accumulated value of continuous weights in a set of weights with a threshold; and generate character level semantic features based on the divided frame level semantic coding features and a set of weights.
[0075] In some embodiments, the semantic feature generation unit 710 can also be configured to: generate character-level semantic coding features of the target speech data based on the divided frame-level semantic coding features and a set of weights; and generate character-level semantic features based on the character-level semantic coding features.
[0076] In some embodiments, the semantic feature generation unit 710 can also be configured to: semantically decode the character-level semantic encoding features to obtain character-level semantic decoding features; and concatenate the character-level semantic decoding features and the character-level semantic encoding features to generate character-level semantic features.
[0077] In some embodiments, the voiceprint feature generation unit 720 may also be configured to: perform voiceprint encoding on the frame-level acoustic features to obtain frame-level voiceprint encoding features; and generate character-level voiceprint features based on the frame-level voiceprint encoding features and a set of weights.
[0078] In some embodiments, the voiceprint feature generation unit 720 can also be configured to: divide the frame-level voiceprint coding features based on the comparison between the accumulated value of continuous weights in a set of weights and a threshold; generate character-level voiceprint coding features based on the divided frame-level voiceprint coding features and a set of weights; and perform voiceprint decoding on the character-level voiceprint coding features to obtain character-level voiceprint features.
[0079] In some embodiments, the detection unit 730 can also be configured to: generate a speech content representation of the target speech data based on character-level semantic features; generate a speaker voiceprint difference representation of the target speech data based on character-level voiceprint features; and determine the characters in the target speech data where speaker switching occurs based on the speech content representation and the speaker voiceprint difference representation.
[0080] In some embodiments, the characters of the target voice data include at least one of the following: a character, a word, a subword, a letter, a syllable, or a phoneme.
[0081] Figure 8A schematic block diagram of an example device 800 that can be used to implement an embodiment of the present disclosure is shown. For example, a terminal device 1011, 1012, 1013 or a server 103 according to an embodiment of the present disclosure can be implemented by a device 800. As shown in the figure, the device 800 includes a central processing unit (CPU) or a graphics processing unit (GPU) 801, which can perform various appropriate actions and processes according to computer program instructions stored in a read-only memory (ROM) 802 or computer program instructions loaded from a storage unit 808 into a random access memory (RAM) 803. In RAM 803, various programs and data required for the operation of the device 800 can also be stored. CPU / GPU 801, ROM 802 and RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0082] A number of components in the device 800 are connected to the I / O interface 805, including: an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a disk, an optical disk, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows the device 800 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0083] Each process, processing, model or device described above may be executed or implemented by the CPU / GPU 801. For example, in some embodiments, the method 300 may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 808. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the CPU / GPU 801, one or more actions of the method 300 described above may be performed, the method 300 may be implemented, and the method 300 may be implemented. Figure 2 , Figures 4 to 6 Any one or more of the speech processing model 210, the semantic feature model 220, the voiceprint feature 230 model, the speaker switching detection model 240, and the acoustic feature extraction model 250 shown in the figure, or the implementation Figure 7 The device 700 is shown.
[0084] The present disclosure may be a method, an apparatus, a system and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for executing various aspects of the present disclosure.
[0085] A computer-readable storage medium may be a tangible device that can hold and store instructions used by an instruction execution device. A computer-readable storage medium may be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples of computer-readable storage media (a non-exhaustive list) include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disk read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punch card or a raised structure in a groove on which instructions are stored, and any suitable combination of the foregoing. As used herein, a computer-readable storage medium is not to be interpreted as a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., a light pulse through a fiber optic cable), or an electrical signal transmitted through a wire.
[0086] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical fiber transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in the computer-readable storage medium in each computing / processing device.
[0087] The computer program instructions for performing the operation of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages, such as Smalltalk, C++, etc., and conventional procedural programming languages, such as "C" language or similar programming languages. Computer-readable program instructions may be executed entirely on a user's computer, partially on a user's computer, as an independent software package, partially on a user's computer, partially on a remote computer, or completely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., using an Internet service provider to connect via the Internet). In some embodiments, an electronic circuit, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), may be customized by utilizing the state information of the computer-readable program instructions, and the electronic circuit may execute the computer-readable program instructions, thereby realizing various aspects of the present disclosure.
[0088] Various aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of the methods, devices (systems) and computer program products according to the embodiments of the present disclosure. It should be understood that each box in the flowchart and / or block diagram and the combination of each box in the flowchart and / or block diagram can be implemented by computer-readable program instructions.
[0089] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, so that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device that implements the functions / actions specified in one or more boxes in the flowchart and / or block diagram is generated. These computer-readable program instructions can also be stored in a computer-readable storage medium, and these instructions cause the computer, programmable data processing device, and / or other equipment to work in a specific manner, so that the computer-readable medium storing the instructions includes a manufactured product, which includes instructions for implementing various aspects of the functions / actions specified in one or more boxes in the flowchart and / or block diagram.
[0090] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operating steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.
[0091] The flow chart and block diagram in the accompanying drawings show the possible architecture, function and operation of the system, method and computer program product according to multiple embodiments of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a part of a module, program segment or instruction, and a part of a module, program segment or instruction includes one or more executable instructions for realizing the specified logical function. In some alternative implementations, the function marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous square boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs the specified function or action, or can be implemented with a combination of special hardware and computer instructions.
[0092] The embodiments of the present disclosure have been described above, and the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and changes will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The selection of terms used herein is intended to best explain the principles of the embodiments, practical applications, or improvements to the technology in the market, or to enable other persons of ordinary skill in the art to understand the embodiments disclosed herein.
Claims
1. A method for speech processing, include: Based on the frame-level acoustic features of the target speech data, generating a character-level semantic feature sequence of the target speech data, wherein the semantic features in the character-level semantic feature sequence correspond to the characters in the target speech data; Based on the frame-level acoustic features, generating a character-level voiceprint feature sequence of the target speech data, wherein the voiceprint features in the character-level voiceprint feature sequence correspond to the characters in the target speech data; and Based on the character-level semantic feature sequence and the character-level voiceprint feature sequence, characters in the target speech data where speaker switching occurs are determined.
2. The method according to claim 1, in, Generating a character-level semantic feature sequence of the target speech data based on the frame-level acoustic features of the target speech data includes: Performing semantic encoding on the frame-level acoustic features to obtain frame-level semantic encoding features; Based on the frame-level semantic coding features, generating a set of weights, wherein the weights in the set of weights correspond frame by frame to the semantic coding features for the frame in the frame-level semantic coding features; and The character-level semantic feature sequence is generated based on the set of weights and the frame-level semantic encoding features.
3. The method according to claim 2, in, Generating the character-level semantic feature sequence based on the set of weights and the frame-level semantic encoding features comprises: Based on the comparison between the accumulated value of the continuous weights in the set of weights and a threshold, the frame-level semantic coding features are divided; and The character-level semantic feature sequence is generated based on the divided frame-level semantic coding features and the set of weights.
4. The method according to claim 3, in, Generating the character-level semantic feature sequence based on the divided frame-level semantic coding features and the set of weights includes: Generate character-level semantic coding features of the target speech data based on the divided frame-level semantic coding features and the set of weights; and Based on the character-level semantic encoding features, the character-level semantic feature sequence is generated.
5. The method according to claim 4, in, Generating the character-level semantic feature sequence based on the character-level semantic encoding feature comprises: semantically decoding the character-level semantic encoding features to obtain character-level semantic decoding features; and The character-level semantic decoding features and the character-level semantic encoding features are concatenated to generate the character-level semantic feature sequence.
6. The method according to claim 2, in, Generating the character-level voiceprint feature of the target speech data based on the frame-level acoustic feature includes: Performing voiceprint encoding on the frame-level acoustic features to obtain frame-level voiceprint encoding features; and The character-level voiceprint feature is generated based on the frame-level voiceprint encoding feature and the set of weights.
7. The method according to claim 6, in, Generating the character-level voiceprint feature based on the frame-level voiceprint encoding feature and the set of weights includes: Based on the comparison between the accumulated value of the continuous weights in the set of weights and the threshold, dividing the frame-level voiceprint coding features; Generate character-level voiceprint coding features based on the divided frame-level voiceprint coding features and the set of weights; and The character-level voiceprint encoding feature is voiceprint decoded to obtain the character-level voiceprint feature.
8. The method according to claim 1, in, Determining characters in the target speech data where speaker switching occurs based on the character-level semantic feature sequence and the character-level voiceprint feature sequence includes: generating a speech content representation of the target speech data based on the character-level semantic feature sequence; generating a speaker voiceprint difference representation of the target speech data based on the character-level voiceprint feature sequence; and Based on the speech content representation and the speaker voiceprint difference representation, characters in the target speech data where speaker switching occurs are determined.
9. The method according to any one of claims 1 to 8, in, The characters of the target voice data include at least one of the following: a word, a phrase, a subword, a letter, a syllable, or a phoneme.
10. A device for speech processing, include: A semantic feature generating unit configured to generate a character-level semantic feature sequence of the target speech data based on the frame-level acoustic features of the target speech data, wherein the semantic features in the character-level semantic feature sequence correspond to the characters in the target speech data; a voiceprint feature generating unit configured to generate a character-level voiceprint feature sequence of the target voice data based on the frame-level acoustic features, wherein the voiceprint features in the character-level voiceprint feature sequence correspond to the characters in the target voice data; as well as The detection unit is configured to determine the characters in the target speech data where speaker switching occurs based on the character-level semantic feature sequence and the character-level voiceprint feature sequence.
11. An electronic device, include: at least one processing unit; At least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, causing the device to perform the method according to any one of claims 1 to 9.
12. A computer-readable storage medium comprising machine-executable instructions which, when executed by a device, cause the device to perform the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Text-independent voiceprint recognition method, device and equipment based on phoneme assistance
CN111785284A
Semantic and sound groove information combined speaking person identity system
CN1547191A
Cited By
Method, apparatus, electronic device, and medium for speech processing
US12614553B2