A speech separation recognition method, device, storage medium and electronic equipment
By processing and recognizing real-time voice streams, extracting and separating the voice of the target object, the problem of low efficiency and low accuracy of voice data recognition in intelligent customer service is solved, achieving efficient and accurate voice separation and improving user experience.
Patent Information
- Application Number
- CN202211508807.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-29
- Publication Date
- 2026-02-27
- Estimated Expiration
- 2042-11-29
AI Technical Summary
Existing intelligent customer service systems suffer from human voice interference and noise interference when receiving customer voice data, resulting in low voice data recognition efficiency and accuracy, which affects user experience.
By processing real-time audio streams, the speech of the target object is extracted and recognized, including acquiring speech waveforms and frame signals, reconstructing and calculating using audio feature vectors and basis functions, and separating the speech text of the target object.
It achieves efficient and accurate separation of target speech text from real-time speech streams with interference, improving recognition efficiency and accuracy, and enhancing user experience.
Smart Images

Figure CN115831118B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of speech processing, in particular to a speech separation and recognition method and device, a storage medium and an electronic device. BACKGROUND
[0002] With the development of artificial intelligence technology, intelligent customer service is widely used in various industries. Intelligent customer service uses artificial intelligence and natural language understanding technology, which can analyze the questions raised by customers and respond according to the background and conversation content of the other party, just like a human.
[0003] Currently, when providing services for customers, intelligent customer service needs to analyze the voice data of customers in real time. Due to the high uncertainty of the external environment of customers, there may be human voice interference when the intelligent customer service receives the voice of the customer. Currently, after the intelligent customer receives the voice data of the customer, it usually identifies all the voice data. However, due to the related interference noise in the voice data, the efficiency and accuracy of voice data recognition are low.
[0004] Therefore, how to provide a technical solution of an efficient speech separation and recognition method has become a technical problem to be solved. SUMMARY
[0005] Some embodiments of the present application aim to provide a speech separation and recognition method, device, storage medium and electronic device. Through the technical solution of the embodiments of the present application, the voice of the target object can be accurately separated from the real-time voice stream with interference, which has high accuracy and high efficiency.
[0006] In a first aspect, some embodiments of the present application provide a speech separation and recognition method, comprising: extracting target object voice in a real-time voice stream; and identifying the target object voice to obtain voice text of the target object.
[0007] Some embodiments of the present application first extract the target object voice from the real-time voice stream, and then identify the target object voice to obtain the voice text, which can accurately separate the voice text of the target object from the real-time voice stream with interference, and has high accuracy and high efficiency.
[0008] In some embodiments, the extraction of the target object voice in the real-time voice stream comprises: processing the real-time voice stream to obtain a voice waveform and acquire waveform signals of each voice frame in the real-time voice stream; and generating the target object voice corresponding to the voice waveform and the waveform signals of each voice frame.
[0009] Some embodiments of the present application can obtain a speech waveform and a waveform signal by processing and analyzing a real-time speech stream, and then obtain target object speech, so as to realize accurate extraction of target object speech with high efficiency.
[0010] In some embodiments, the processing of the real-time speech stream to obtain a speech waveform comprises: extracting an audio feature vector in the real-time speech stream; and reconstructing the audio feature vector to obtain the speech waveform.
[0011] Some embodiments of the present application can provide effective data for quickly obtaining target object speech by reconstructing an audio feature vector in a real-time speech stream to obtain a speech waveform.
[0012] In some embodiments, the reconstructing of the audio feature vector to obtain the speech waveform comprises: multiplying the audio feature vector by a basis function to obtain the speech waveform.
[0013] Some embodiments of the present application can obtain a speech waveform by an audio feature vector and a basis function, which is simple and efficient.
[0014] In some embodiments, the obtaining of the waveform signal of each speech frame in the real-time speech stream comprises: separating the real-time speech stream to obtain a speech frame vector; performing operation on the speech frame vector and the audio feature vector to obtain a source feature vector; and multiplying the source feature vector by the basis function to obtain the waveform signal of each speech frame.
[0015] Some embodiments of the present application can provide effective data for quickly obtaining target object speech by separating a real-time speech stream to obtain a speech frame vector, and then performing operation to obtain a speech frame waveform signal.
[0016] In a second aspect, some embodiments of the present application provide a speech separation and recognition device, comprising: an extraction module configured to extract target object speech in a real-time speech stream; and a recognition module configured to recognize the target object speech to obtain speech text of a target object.
[0017] In some embodiments, the extraction module is configured to process the real-time speech stream to obtain a speech waveform, and obtain a waveform signal of each speech frame in the real-time speech stream; and generate the target object speech corresponding to the speech waveform and the waveform signal of each speech frame.
[0018] In some embodiments, the extraction module is configured to extract an audio feature vector in the real-time speech stream; and reconstruct the audio feature vector to obtain the speech waveform.
[0019] In some embodiments, the extraction module is configured to multiply the audio feature vector by a basis function to obtain the speech waveform.
[0020] In some embodiments, the extraction module is configured to separate the real-time speech stream to obtain speech frame vectors; perform operations on the speech frame vectors and the audio feature vector to obtain source feature vectors; and multiply the source feature vectors by the basis function to obtain waveform signals of the speech frames.
[0021] In a third aspect, some embodiments of the present application provide a computer readable storage medium having stored thereon a computer program, which, when executed by a processor, can implement the method according to any of the embodiments of the first aspect.
[0022] In a fourth aspect, some embodiments of the present application provide an electronic device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, can implement the method according to any of the embodiments of the first aspect.
[0023] In a fifth aspect, some embodiments of the present application provide a computer program product, which includes a computer program, wherein the computer program, when executed by a processor, can implement the method according to any of the embodiments of the first aspect. BRIEF DESCRIPTION OF DRAWINGS
[0024] In order to more clearly illustrate the technical solutions of some embodiments of the present application, the following will briefly introduce the drawings needed to be used in some embodiments of the present application. It should be understood that the following drawings only show some embodiments of the present application, and therefore should not be regarded as a limitation on the scope, and for those skilled in the art, other related drawings can also be obtained without creative labor on the basis of these drawings.
[0025] Figure 1 A speech separation system diagram provided by some embodiments of the present application;
[0026] Figure 2 One of the speech separation identification method flowcharts provided by some embodiments of the present application;
[0027] Figure 3 A speech separation structure composition schematic diagram provided by some embodiments of the present application;
[0028] Figure 4 The second speech separation identification method flowchart provided by some embodiments of the present application;
[0029] Figure 5 The device composition block diagram of speech separation identification provided by some embodiments of the present application;
[0030] Figure 6 An electronic device schematic diagram is provided for some embodiments of the present application. DETAILED DESCRIPTION
[0031] The technical solutions in some embodiments of the present application will be described below in conjunction with the drawings in some embodiments of the present application.
[0032] It should be noted that similar reference numerals and letters refer to like items in the following drawings, and therefore, once an item is defined in one drawing, it need not be further defined and explained in subsequent drawings. Meanwhile, in the description of the present application, the terms "first", "second", and the like are only used to distinguish description, and cannot be understood as indicating or implying relative importance.
[0033] In the related art, intelligent customer service is developed on the basis of large-scale knowledge processing and is an industry application-oriented technology with industry universality. The intelligent customer service not only provides fine-grained knowledge management technology for enterprises, but also establishes a fast and effective technical means based on natural language for communication between enterprises and a large number of users. At the same time, it can also provide statistical analysis information required for fine management for enterprises. The main challenge in intelligent customer service is to automatically recognize speech and obtain speech data of users with high accuracy. In the prior art, a model adaptive method is used to recognize all input speech streams to obtain speech text of users. However, due to the complexity of the external environment of the user, there may be noise interference (for example, other voices or environmental noise) in the speech data of the user. At this time, the model adaptive method in the prior art is still used to recognize the speech data of the user, which results in low recognition efficiency and cannot guarantee the recognition accuracy, thereby reducing the user experience.
[0034] In view of this, some embodiments of the present application provide a voice separation and recognition method, which extracts target object speech from a real-time voice stream, and then recognizes the target object speech to obtain speech text of the target object. Some embodiments of the present application can quickly separate the voice stream to obtain the speech text of the target object, which is efficient and has high accuracy, thereby improving the user experience.
[0035] The technical solutions in some embodiments of the present application will be described below in conjunction with the drawings in some embodiments of the present application. Figure 1 The system composition structure of the voice separation provided by some embodiments of the present application is described by way of example.
[0036] As Figure 1As shown, some embodiments of the present application provide a voice separation system, which comprises a terminal 100 and a voice separation server 200. A user 300 (as a specific example of a target object) can communicate with the voice separation server 200 through the terminal 100. The voice separation server 200 can receive a voice stream from the user in real time through the terminal 100. The voice separation server 200 first extracts the user voice from the voice stream, and then recognizes the user voice to finally obtain the voice text of the user 300. Then the voice separation server 200 can also reply to the terminal 100 with the corresponding voice based on the voice text of the user 300, and the terminal 100 can broadcast the voice to the user 300.
[0037] In some embodiments of the present application, the terminal 100 can be a mobile terminal or a non-mobile terminal. For example, a mobile phone, an iPad, a smart communication watch, or a fixed telephone, etc., and the present application is not limited thereto.
[0038] The following will be described in conjunction with the accompanying drawings Figure 2 The following will be described in conjunction with the accompanying drawings
[0039] Please refer to the accompanying drawings Figure 2 , Figure 2 A flowchart of a voice separation and recognition method provided by some embodiments of the present application is shown in FIG. 6. The method comprises the following steps:
[0040] S210, extracting the target object voice in the real-time voice stream.
[0041] For example, in some embodiments of the present application, in the field of ASR (Automatic Speech Recognition), there may be human voice interference or other noise interference in the real-time voice stream generated by the user 300. After receiving the real-time voice stream, the first thing the voice separation server 200 does is to extract the user voice (as a specific example of the target object voice) so as to facilitate subsequent voice recognition. Unlike the method of filtering noise interference in the prior art, the purpose of some embodiments of the present application is to extract the user voice and improve the efficiency of voice processing, regardless of how much noise interference is contained in the real-time voice stream.
[0042] In some embodiments of the present application, S210 can comprise:
[0043] S211, processing the real-time voice stream to obtain a voice waveform and acquiring a waveform signal of each voice frame in the real-time voice stream; S212, generating the target object voice corresponding to the voice waveform and the waveform signal of each voice frame.
[0044] For example, to extract user speech from a multi-voice or mixed audio stream (as a specific example of a real-time audio stream), in some embodiments of this application, it is first necessary to process the real-time audio stream to obtain the user speech waveform (as a specific example of a speech waveform). Then, the real-time audio stream needs to be segmented into frames to obtain the waveform signal of each audio frame. Finally, the user speech can be obtained based on the user speech waveform and the waveform signals of each frame.
[0045] The following is in conjunction with the appendix Figure 3 The specific implementation process of S210 is illustrated by example. (The appendix...) Figure 3 The diagram illustrates a speech separation structure for some embodiments of this application. It is understood that the speech separation structure is deployed in a speech separation server 200 to achieve the purpose of speech separation.
[0046] In some embodiments of this application, S211 may include: extracting audio feature vectors from the real-time speech stream; reconstructing the audio feature vectors to obtain the speech waveform. Specifically, the audio feature vectors are multiplied by a basis function to obtain the speech waveform.
[0047] For example, in some embodiments of this application, a multi-voice speech stream 310 (as a specific example of a real-time speech stream) is input into an encoder 320, where the encoder 320 can be a 1-D convolutional layer. By inputting the multi-voice speech stream 310 into the encoder 320, the extraction and reconstruction of audio feature vectors can be achieved.
[0048] Specifically, the audio feature vector w = H(xU) is obtained using the following formula. T ), where x is a multi-person voice stream 310, U∈R NxL , where N is the number of speech vectors, and L is the length of the speech stream corresponding to each speech vector (L is the overlapping segment obtained by dividing the input x, and the overlapping part refers to how many different human voice parts are in x). H(·) is an optional nonlinear function.
[0049] Then, for example, decoder 330 processes w to obtain the speech waveform. Specifically, decoder 330 uses a one-dimensional transpose convolution operation to modify the representation and reconstruct the waveform. That is, the user speech stream corresponding to the reconstructed speech vector's speech waveform is obtained using the following formula: in, Let V ∈ R be the user speech stream corresponding to a reconstructed speech vector of length L. NxL These are the basis functions of decoder 330. Finally, the user speech streams obtained by reconstructing the speech vectors of each speech stream of length L are reassembled to obtain the speech waveform corresponding to the entire segment x.
[0050] In some embodiments of the present application, the input mixed signal (as a specific example of real-time speech stream) can be divided into overlapping segments (e.g., there are L overlapping segments) with x k ∈R 1xL , k = 1,..., k represents the segment index. The encoder 320 converts x k into N-dimensional representation w by 1-D convolution operation, where w ∈ R 1xN .
[0051] In some embodiments of the present application, S212 can include: separating the real-time speech stream to obtain speech frame vectors; operating the speech frame vectors and the audio feature vectors to obtain source feature vectors; and multiplying the source feature vectors and the basis functions to obtain waveform signals of the speech frames.
[0052] For example, in some embodiments of the present application, the separator 340 is used to realize the separation of each frame in x by estimating C vector masks to obtain speech frame vectors: m i ∈R N×L , i = 1, 2,..., C, C is the number of human voices in x, and m i ∈ [0, 1]. Operating with w to obtain the corresponding source representation (as a specific example of source feature vectors): d i = w ⊙ m i . Then the decoder 330 is used to estimate the waveform signal of each source: Finally, the target human voice 350 (i.e., user speech) matching the waveform signal of each source and speech waveform is generated.
[0053] S220, recognizing the target object speech to obtain speech text of the target object.
[0054] For example, in some embodiments of the present application, after obtaining the user speech by the above-mentioned embodiments, the speech recognition algorithm is used to recognize the text of the user speech to obtain the corresponding speech text.
[0055] The specific process of speech separation provided by some embodiments of the present application will be described below with reference to the accompanying drawings. Figure 4
[0056] Please refer to the accompanying drawings Figure 4 , Figure 4 for a method flowchart of speech separation and recognition provided by some embodiments of the present application. The specific process of the above-mentioned speech separation will be described below.
[0057] S410, acquire a real-time voice stream.
[0058] For example, as a specific example of the present application, it is assumed that the acquired real-time voice stream of the user is a multi-voice voice stream containing human voice interference. After the multi-voice voice stream is acquired by the voice separation server 200, it is input into the encoder 320 in the voice separation server 200.
[0059] S420, extract an audio feature vector in the real-time voice stream.
[0060] For example, as a specific example of the present application, the encoder 320 uses a convolution layer to divide the multi-voice voice stream into a voice vector with a length of L and extract it to obtain an audio feature vector. The specific acquisition of the audio feature vector can refer to Figure 2 The method embodiment, to avoid repetition hereinafter.
[0061] S430, reconstruct the audio feature vector to acquire a voice waveform.
[0062] For example, as a specific example of the present application, the decoder 330 uses a one-dimensional transpose convolution operation to reconstruct the waveform in a modified representation form to obtain Figure 2 The corresponding representation in the method embodiment, to avoid repetition hereinafter. After reconstructing each voice vector with a length of L, the reconstructed segment addition generates a voice waveform.
[0063] S440, separate the real-time voice stream to obtain a voice frame vector.
[0064] For example, as a specific example of the present application, the separator 340 realizes the separation of each frame by estimating the number of human voices in the multi-voice voice stream. One person corresponds to one voice vector.
[0065] S450, operate each voice frame vector with the audio feature vector to obtain a source feature vector.
[0066] For example, as a specific example of the present application, the separator 340 mixes each voice frame vector onto the audio feature vector to obtain the corresponding representation d of the source i .
[0067] S460, multiply the source feature vector with a basis function to obtain a waveform signal of each voice frame.
[0068] For example, as a specific example of the present application, the decoder 330 multiplies each source with a basis function to obtain a waveform signal of each source. The specific formula can refer to Figure 2 The method embodiment, to avoid repetition hereinafter.
[0069] S470, generating target object voice corresponding to the speech waveform and the waveform signal of each speech frame.
[0070] For example, as a specific example of the present application, the decoder 330 generates the corresponding target human voice (as a specific example of the target object voice) according to the speech waveform and the waveform signal of each source.
[0071] S480, identifying the target object voice to obtain the speech text of the target object.
[0072] For example, as a specific example of the present application, the speech separation server 200 converts the target human voice into a speech text form by using a speech recognition algorithm, so as to achieve the purpose of speech recognition of the target human voice.
[0073] As can be seen from the above some embodiments of the present application, in the real-time speech recognition system, the terminal 100 can transmit the real-time speech stream through the network, the speech separation server 200 first extracts the target human voice by separating the real-time speech stream, and then performs speech recognition on the target human voice. The whole process is simple, efficient and has high reliability and good practicability.
[0074] Please refer to Figure 5 , Figure 5 The composition block diagram of the speech separation and recognition device provided by some embodiments of the present application is shown. It should be understood that the speech separation and recognition device corresponds to the above-mentioned method embodiments, and can perform each step involved in the above-mentioned method embodiments. The specific functions of the speech separation and recognition device can be referred to the description in the above, and the detailed description is appropriately omitted here to avoid repetition.
[0075] Figure 5 The speech separation and recognition device includes at least one software function module which can be stored in the memory in the form of software or firmware or solidified in the speech separation and recognition device. The speech separation and recognition device includes: an extraction module 510 for extracting target object voice in a real-time speech stream; an identification module 520 for identifying the target object voice to obtain the speech text of the target object.
[0076] In some embodiments of the present application, the extraction module 510 is configured to process the real-time speech stream to obtain a speech waveform, and obtain a waveform signal of each speech frame in the real-time speech stream; and generate the target object voice corresponding to the speech waveform and the waveform signal of each speech frame.
[0077] In some embodiments of the present application, the extraction module 510 is configured to extract an audio feature vector in the real-time speech stream; and reconstruct the audio feature vector to obtain the speech waveform.
[0078] In some embodiments of the present application, the extraction module 510 is configured to multiply the audio feature vector by a basis function to obtain the speech waveform.
[0079] In some embodiments of the present application, the extraction module 510 is configured to separate the real-time speech stream to obtain speech frame vectors, perform operations on the speech frame vectors and the audio feature vector to obtain source feature vectors, and multiply the source feature vectors by the basis function to obtain waveform signals of the speech frames.
[0080] Some embodiments of the present application also provide a computer readable storage medium having a computer program stored thereon, where the program, when executed by a processor, can implement the operations of the method corresponding to any of the above embodiments.
[0081] Some embodiments of the present application also provide a computer program product, which includes a computer program, where the computer program, when executed by a processor, can implement the operations of the method corresponding to any of the above embodiments.
[0082] As shown in Figure 6 Some embodiments of the present application provide an electronic device 600, which includes a memory 610, a processor 620, and a computer program stored in the memory 610 and executable on the processor 620, where the processor 620 reads the program from the memory 610 through a bus 630 and executes the program to implement the method of any of the above embodiments.
[0083] The processor 620 can process digital signals and can include various computing structures, such as a complex instruction set computer structure, a reduced instruction set computer structure, or a structure implementing a combination of multiple instruction sets. In some examples, the processor 620 can be a microprocessor.
[0084] The memory 610 can be used to store instructions executed by the processor 620 or data related to the execution of the instructions. The instructions and / or data can include code for implementing some or all of the functions of one or more modules described in the embodiments of the present application. The processor 620 of the embodiments of the present disclosure can be configured to execute the instructions in the memory 610 to implement the method shown above. The memory 610 includes a dynamic random access memory, a static random access memory, a flash memory, an optical memory, or other memories well known to those skilled in the art.
[0085] The above merely provides an example of the present application and is not intended to limit the protection scope of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application. It should be noted that similar reference numerals and letters represent similar items in the following drawings, and thus, once an item is defined in one drawing, it need not be further defined and explained in subsequent drawings.
[0086] The above merely provides an example of the present application and is not intended to limit the protection scope of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application. It should be noted that similar reference numerals and letters represent similar items in the following drawings, and thus, once an item is defined in one drawing, it need not be further defined and explained in subsequent drawings.
[0087] It should be noted that the relational terms herein such as first and second and the like are used solely to distinguish one from another entity or action without necessarily requiring or implying any actual relationship or order between such entities or actions. Moreover, the terms "comprises", "comprising", or any other variation thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article, or apparatus. An element proceeded by "comprises... a" does not, without more constraints, exclude the existence of additional identical elements in the process, method, article, or apparatus that comprises the element.
Claims
1. A method of speech separation recognition, characterized by, The method comprises the following steps: extracting target object speech in a real-time speech stream; wherein the target object speech is generated based on a speech waveform output by a decoder and a waveform signal of each speech frame; the waveform signal of each speech frame is obtained by the following method: a separator separates the real-time speech stream to obtain a speech frame vector; the decoder operates the speech frame vector and an audio feature vector to obtain a source feature vector; wherein the audio feature vector is obtained by an encoder extracting the real-time speech stream; the speech waveform is obtained by the decoder processing the audio feature vector; the decoder multiplies the source feature vector by a basis function to obtain the waveform signal of each speech frame; recognizing the target object speech to obtain speech text of the target object.
2. The method of claim 1, wherein, The method for extracting target object speech in a real-time speech stream comprises the following steps: processing the real-time speech stream to obtain a speech waveform and a waveform signal of each speech frame in the real-time speech stream; generating the target object speech corresponding to the speech waveform and the waveform signal of each speech frame.
3. The method of claim 2, wherein, The method for processing the real-time speech stream to obtain a speech waveform comprises the following steps: extracting an audio feature vector in the real-time speech stream; reconstructing the audio feature vector to obtain the speech waveform.
4. The method of claim 3, wherein, The method for reconstructing the audio feature vector to obtain the speech waveform comprises the following steps: multiplying the audio feature vector by a basis function to obtain the speech waveform.
5. An apparatus for speech separation recognition, the apparatus comprising: The device is used to execute the method of claim 1, comprising: an extraction module for extracting target object speech in a real-time speech stream; an identification module for identifying the target object speech to obtain speech text of the target object.
6. The apparatus of claim 5, wherein, The extraction module is used to: process the real-time speech stream to obtain a speech waveform and a waveform signal of each speech frame in the real-time speech stream; generate the target object speech corresponding to the speech waveform and the waveform signal of each speech frame.
7. The apparatus of claim 6, wherein, The extraction module is used to: extract an audio feature vector in the real-time speech stream; reconstruct the audio feature vector to obtain the speech waveform.
8. The apparatus of claim 7, wherein, The extraction module is used to: multiply the audio feature vector by a basis function to obtain the speech waveform.
9. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, wherein the computer program is run by a processor to execute the method of any one of claims 1-4.
10. An electronic device, comprising: The device comprises a memory, a processor and a computer program stored in the memory and run on the processor, wherein the computer program is run by the processor to execute the method of any one of claims 1-4.
Citation Information
Patent Citations
Voice signal recognition method and device, electronic equipment and storage medium
CN113571063A
Voice separation method and device, electronic equipment and storage medium
CN114783459A