Voice interaction method, apparatus, device, medium and program product

By determining the confidence levels of the original signal and the speech enhancement signal and performing a weighted sum, the distortion problem caused by speech enhancement is solved, the accuracy of voice interaction is improved, model training is simplified, and costs are reduced.

CN119943071BActive Publication Date: 2025-12-12IFLYTEK CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510063415.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-12-12
Estimated Expiration
2045-01-15

AI Technical Summary

Technical Problem

Speech enhancement techniques may cause distortion of the target speech under low signal-to-noise ratio conditions, affecting the accuracy of voice interaction. Although existing front-end and back-end integrated joint training solutions have improved this, they cannot completely eliminate distortion and increase the cost of model training.

Method used

By determining the confidence levels of the original signal and the speech enhancement signal, and using confidence model training and weighted summation methods, the target signal is determined by comprehensively considering the confidence levels of the original signal and the speech enhancement signal to reduce speech distortion.

Benefits of technology

By reducing distortion caused by speech enhancement, the accuracy of voice interaction is improved, the model training process is simplified, and the technical cost is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119943071B_ABST
    Figure CN119943071B_ABST
Patent Text Reader

Abstract

The application provides a voice interaction method, device, equipment, medium and program product. The voice interaction method comprises the following steps: determining a first confidence degree of a first original signal and a second confidence degree of a first voice enhanced signal; the first voice enhanced signal is a signal obtained by performing voice enhancement on the first original signal; determining a target signal based on the first confidence degree, the second confidence degree, a second original signal and a second voice enhanced signal; the second voice enhanced signal is a signal obtained by performing voice enhancement on the second original signal; and performing voice interaction with a target device based on the target signal. The application can reduce voice distortion caused by voice enhancement.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of voice processing, and in particular to a voice interaction method, device, equipment, medium and program product. BACKGROUND

[0002] With the development of artificial intelligence, human-computer interaction through voice has been more and more widely used in people's daily life. For example, in the field of intelligent automobile vehicle, with the landing application of voice wake-up and voice recognition technology, people can give instructions to the vehicle equipment through voice to make the vehicle equipment complete specific tasks, so that people can free their hands.

[0003] In the voice human-computer interaction technology, the machine picks up the instructions issued by people through the microphone, and then sends them to the downstream voice interaction module (such as voice wake-up, voice recognition, etc.) after processing by the acoustic front-end module to complete the specified task. Because the environment where the machine is located can be various, the microphone signal can contain various noises, which will increase the difficulty of the downstream voice interaction task, so that the machine can not correctly understand and complete various instructions. In order to solve this problem, by adding voice enhancement technology in the acoustic front-end module, through various linear or nonlinear methods, the noise components in the microphone signal are eliminated, and the target voice components are retained, so that the downstream voice interaction task can better understand and complete various instructions.

[0004] However, voice enhancement sometimes also brings distortion of the target voice, especially when the signal-to-noise ratio is low, the distortion of the target voice introduced by voice enhancement can be more serious, and even the signal after voice enhancement is more difficult to understand and complete than the original microphone signal. SUMMARY

[0005] The present application provides a voice interaction method, device, equipment, medium and program product to reduce the voice distortion caused by voice enhancement.

[0006] According to a first aspect of an embodiment of the present application, a voice interaction method is provided, comprising:

[0007] determining a first confidence of a first original signal and a second confidence of a first voice enhanced signal; wherein the first voice enhanced signal is a signal obtained by voice enhancement on the first original signal;

[0008] determining a target signal based on the first confidence, the second confidence, a second original signal and a second voice enhanced signal; wherein the second voice enhanced signal is a signal obtained by voice enhancement on the second original signal;

[0009] based on the target signal, voice interaction with a target device.

[0010] Optionally, the first original signal and the second original signal are extracted from the same original speech signal.

[0011] Optionally, the determining the target signal based on the first confidence, the second confidence, the second original signal and the second speech enhanced signal comprises:

[0012] determining a first weight corresponding to the second original signal based on the first confidence;

[0013] determining a second weight corresponding to the second speech enhanced signal based on the second confidence;

[0014] performing weighted summation based on the second original signal, the first weight, the second speech enhanced signal and the second weight to obtain the target signal.

[0015] Optionally, the determining the first confidence of the first original signal and the second confidence of the first speech enhanced signal comprises:

[0016] extracting a first feature of the first original signal and a second feature of the first speech enhanced signal;

[0017] inputting the first feature and the second feature into a confidence model to obtain the first confidence of the first original signal and the second confidence of the first speech enhanced signal.

[0018] Optionally, the training process of the confidence model comprises:

[0019] extracting a first sample feature of a first sample original signal and a second sample feature of a first sample speech enhanced signal;

[0020] inputting the first sample feature and the second sample feature into a confidence model to obtain a first sample confidence of the first sample original signal and a second sample confidence of the first sample speech enhanced signal;

[0021] determining a sample speech interaction signal based on the first sample confidence, the second sample confidence, the second sample original signal and the second sample speech enhanced signal;

[0022] obtaining a predicted speech interaction label based on the sample speech interaction signal;

[0023] training the confidence model based on the predicted speech interaction label and a real speech interaction label corresponding to the second sample original signal to obtain a trained confidence model;

[0024] The first sample voice enhanced signal is a signal obtained by performing voice enhancement on the first sample original signal.

[0025] Optionally, the first original signal and the second original signal are the same, and the first voice enhanced signal and the second voice enhanced signal are the same.

[0026] Optionally, the first original signal and the second original signal are different, and the first voice enhanced signal and the second voice enhanced signal are different; the first original signal is used to wake up the target device; and the second original signal is used to control the target device to perform a target operation.

[0027] According to a second aspect of an embodiment of the present application, a voice interaction device is provided, comprising:

[0028] a first processing unit configured to determine a first confidence of a first original signal and a second confidence of a first voice enhanced signal, wherein the first voice enhanced signal is a signal obtained by performing voice enhancement on the first original signal;

[0029] a second processing unit configured to determine a target signal based on the first confidence, the second confidence, a second original signal, and a second voice enhanced signal, wherein the second voice enhanced signal is a signal obtained by performing voice enhancement on the second original signal;

[0030] a voice interaction unit configured to perform voice interaction with a target device based on the target signal.

[0031] According to a third aspect of an embodiment of the present application, an electronic device is provided, comprising a memory and a processor;

[0032] The memory is connected with the processor, and is configured to store a program;

[0033] The processor is configured to realize the voice interaction method according to the first aspect by running the program in the memory.

[0034] According to a fourth aspect of an embodiment of the present application, a storage medium is provided, and the storage medium stores a computer program. When the computer program is run by a processor, the voice interaction method according to the first aspect is realized.

[0035] According to a fifth aspect of an embodiment of the present application, a computer program product is provided, comprising computer program instructions. When the computer program instructions are run by a processor, the processor performs the voice interaction method according to the first aspect.

[0036] In the present application, the first confidence of the first original signal and the second confidence of the first speech enhanced signal are determined, wherein the first speech enhanced signal is a signal obtained by performing speech enhancement on the first original signal, the confidence of the signal before speech enhancement and the confidence of the signal after speech enhancement can be obtained, since the second speech enhanced signal is a signal obtained by performing speech enhancement on the second original signal, the first confidence and the second confidence can also represent the confidence of the second original signal and the second speech enhanced signal, based on the first confidence, the second confidence, the second original signal and the second speech enhanced signal, the target signal is determined, the second speech enhanced signal is considered in the process of determining the target signal, the advantages brought by speech enhancement can be obtained to a certain extent, moreover, the second original signal and the second speech enhanced signal and the first confidence and the second confidence are comprehensively considered in the process of determining the target signal, the speech distortion brought by speech enhancement can be reduced, and based on the target signal, the voice interaction with the target device can further improve the accuracy of voice interaction on the basis of reducing the speech distortion brought by speech enhancement. BRIEF DESCRIPTION OF DRAWINGS

[0037] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or prior art description will be briefly introduced below. Obviously, the drawings in the following description are only embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor on the basis of the provided drawings.

[0038] FIG. 1 A flowchart of a voice interaction method provided in an embodiment of the present application;

[0039] FIG. 2 A flowchart of step 101 provided in an embodiment of the present application;

[0040] FIG. 3 A flowchart of the training process of a confidence model provided in an embodiment of the present application;

[0041] FIG. 4 A flowchart of step 102 provided in an embodiment of the present application;

[0042] FIG. 5 A flowchart of a voice interaction method provided in an embodiment of the present application;

[0043] FIG. 6 A flowchart of a voice interaction method provided in an embodiment of the present application;

[0044] FIG. 7 A structural diagram of a voice interaction device provided in an embodiment of the present application;

[0045] FIG. 8 Fig. 1 is a structural schematic diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0046] In the voice human-computer interaction technology, a machine picks up an instruction issued by a person through a microphone, and then sends the instruction to a downstream voice interaction module (for example, voice wake-up, voice recognition, etc.) after processing by an acoustic front-end module to complete a specified task. By adding a voice enhancement technology in the acoustic front-end module, through various linear or nonlinear methods, noise components in the microphone signal are eliminated, and target voice components are retained, so that the downstream voice interaction task can better understand and complete various instructions.

[0047] However, voice enhancement sometimes also causes distortion of the target voice, especially when the signal-to-noise ratio is low, the target voice distortion introduced by voice enhancement can be more serious, and even the signal after voice enhancement is more difficult to understand and complete than the original microphone signal.

[0048] At present, a commonly used scheme for reducing voice distortion caused by voice enhancement is as follows: front-end and back-end integration technology, through joint training of an acoustic front-end network (voice enhancement) and a back-end network (voice wake-up, voice recognition, etc.), to reduce voice distortion caused by voice enhancement.

[0049] The front-end and back-end integration joint training method can alleviate the voice distortion problem caused by voice enhancement to a certain extent, but cannot completely eliminate this problem, and the front-end and back-end integration joint training significantly increases the model training time, resulting in an increase in technical cost.

[0050] In order to reduce voice distortion caused by voice enhancement, the present application provides a voice interaction method, device, equipment, medium and program product.

[0051] The technical solutions in the embodiments of the present application will be described clearly and completely in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor are within the scope of protection of the present application.

[0052] Example Implementation Environment

[0053] The voice interaction method according to the embodiments of the present application can be executed by an electronic device such as a terminal device or a server. The terminal device can be a user device, a mobile device, a computing device, a wearable device, etc. The server can be a physical server, a server cluster composed of multiple physical servers, or a cloud server capable of cloud computing. The method can be realized by a processor invoking computer readable program instructions stored in a memory.

[0054] Example Method

[0055] Referring to FIG. 1 In an exemplary embodiment, a voice interaction method is provided. As shown in FIG. 1 The flow of the voice interaction method mainly includes:

[0056] Step 101, determining a first confidence of a first original signal and a second confidence of a first speech enhanced signal.

[0057] The first speech enhanced signal is a signal obtained by performing speech enhancement on the first original signal.

[0058] The first confidence of the first original signal and the second confidence of the first speech enhanced signal are determined, wherein the first speech enhanced signal is a signal obtained by performing speech enhancement on the first original signal, and the confidence of the signal before speech enhancement and the confidence of the signal after speech enhancement can be obtained.

[0059] In an exemplary embodiment, the first original signal can refer to an original speech signal collected by a microphone, or can refer to a part of the original speech signal extracted from the original speech signal collected by the microphone.

[0060] In an exemplary embodiment, the first confidence is used to indicate the confidence of the first original signal for a downstream task; and the second confidence is used to indicate the confidence of the first speech enhanced signal for the downstream task. For example, the first original signal is used to wake up a target device, the first confidence is used to indicate the confidence of the first original signal for the task of waking up the target device, and the second confidence is used to indicate the confidence of the first speech enhanced signal for the task of waking up the target device; for another example, the first original signal is used to control a target device to perform a target operation, the first confidence is used to indicate the confidence of the first original signal for the task of controlling the target device to perform the target operation, and the second confidence is used to indicate the confidence of the first speech enhanced signal for the task of controlling the target device to perform the target operation.

[0061] In some embodiments, step 101 comprises: inputting the first original signal and the first speech enhanced signal into the speech interaction model to obtain a predicted label corresponding to the first original signal and a predicted label corresponding to the first speech enhanced signal; and determining a first confidence of the first original signal and a second confidence of the first speech enhanced signal based on the predicted label corresponding to the first original signal, the predicted label corresponding to the first speech enhanced signal, and a true label corresponding to the first original signal.

[0062] In some other embodiments, as shown in FIG. 1, step 101 comprises: FIG. 2

[0063] Step 201: extracting a first feature of the first original signal and a second feature of the first speech enhanced signal.

[0064] In exemplary embodiments, step 201 can comprise: using a feature extraction method corresponding to a feature such as MFCC (Mel-scale Frequency Cepstral Coefficients), filterbank, Bark domain feature (Bark domain feature refers to an audio feature extracted based on Bark frequency scale, and Bark frequency scale is a kind of nonlinear frequency scale designed based on the hearing characteristics of human ear), and the like, to extract the first feature of the first original signal and the second feature of the first speech enhanced signal from the first original signal and the first speech enhanced signal. Of course, other feature extraction methods can also be used, and the present application does not limit this.

[0065] Step 202: inputting the first feature and the second feature into the confidence model to obtain a first confidence of the first original signal and a second confidence of the first speech enhanced signal.

[0066] ​In the example embodiment, the confidence model can include, but is not limited to, a deep neural network (DNN), a recurrent neural network (RNN), a convolutional neural network (CNN), etc. Considering the limited computing resources of the embedded chip, the confidence model can also be designed as a lightweight network structure for deployment, such as a lightweight CLDNN (Convolutional, Long Short-Term Memory, Deep Neural Network, which is a network structure combining convolutional neural network (CNN), long short-term memory network (LSTM), and deep neural network (DNN). It combines these networks in series to perform frequency domain conversion, time domain association, and feature abstraction in turn, thereby improving the performance of the model), ResNet (Residual Network), UNet (a deep learning network model for image segmentation, named after its U-shaped network structure), etc.

[0067] The first feature of the first original signal and the second feature of the first speech enhanced signal are extracted, and the first feature and the second feature are input to the confidence model to obtain the first confidence of the first original signal and the second confidence of the first speech enhanced signal. The first confidence of the first original signal and the second confidence of the first speech enhanced signal are obtained through the confidence model, which can simply and conveniently obtain the first confidence of the first original signal and the second confidence of the first speech enhanced signal. Moreover, the accuracy of the first confidence of the first original signal and the second confidence of the first speech enhanced signal can be improved by training the confidence model, thereby reducing the speech distortion caused by speech enhancement.

[0068] In some embodiments, as shown in FIG. 3 The training process of the confidence model includes:

[0069] Step 301: Extracting a first sample feature of a first sample original signal and a second sample feature of a first sample speech enhanced signal.

[0070] The first sample speech enhanced signal is a signal obtained by performing speech enhancement on the first sample original signal.

[0071] Step 302: Inputting the first sample feature and the second sample feature to the confidence model to obtain a first sample confidence of the first sample original signal and a second sample confidence of the first sample speech enhanced signal.

[0072] The second sample speech enhanced signal is a signal obtained by performing speech enhancement on the second sample original signal.

[0073] In an example embodiment, the first sample original signal and the second sample original signal are extracted from the same sample original speech signal.

[0074] At step 303, a sample speech interaction signal is determined based on the first sample confidence, the second sample confidence, the second sample original signal, and the second sample speech enhanced signal.

[0075] At step 304, a predicted speech interaction label is obtained based on the sample speech interaction signal.

[0076] In an example embodiment, the predicted speech interaction label can include waking up the target device and not waking up the target device; the predicted speech interaction label can also include instruction A, instruction B, instruction C, and the like, where instruction A is used to control the target device to perform operation a, instruction B is used to control the target device to perform operation b, and instruction C is used to control the target device to perform operation c.

[0077] At step 305, the confidence model is trained based on the predicted speech interaction label and a true speech interaction label corresponding to the second sample original signal, to obtain a trained confidence model.

[0078] In an example embodiment, the predicted speech interaction label is denoted as Pred_asr, and the true speech interaction label corresponding to the second sample original signal is denoted as Label_asr. Step 305 can include determining a loss value Loss based on the predicted speech interaction label Pred_asr and the true speech interaction label Label_asr corresponding to the second sample original signal, and training the confidence model based on the loss value Loss to obtain the trained confidence model.

[0079] In an example embodiment, the loss value calculation can use CE loss (Cross Entropy Loss), CTC loss (Connectionist Temporal Classification Loss), RNNT loss (Recurrent Neural Network Transducer Loss), and the like.

[0080] In an example embodiment, training the confidence model based on the loss value Loss to obtain the trained confidence model can be updating parameters of the confidence model based on the loss value Loss through gradient backpropagation to obtain the trained confidence model.

[0081] The first sample feature of the first sample original signal and the second sample feature of the first sample voice enhanced signal are extracted, the first sample feature and the second sample feature are input into the confidence model, the first sample confidence of the first sample original signal and the second sample confidence of the first sample voice enhanced signal are obtained, the sample voice interaction signal is determined based on the first sample confidence, the second sample confidence, the second sample original signal and the second sample voice enhanced signal, the predicted voice interaction label is obtained based on the sample voice interaction signal, and the confidence model is trained based on the predicted voice interaction label and the real voice interaction label corresponding to the second sample original signal, so as to obtain the trained confidence model. By training the confidence model based on the predicted voice interaction label and the real voice interaction label corresponding to the second sample original signal, the accuracy of the first confidence and the second confidence output by the confidence model can be improved, and the voice distortion caused by voice enhancement can be reduced, and the accuracy of voice interaction can be further improved.

[0082] In step 102, the target signal is determined based on the first confidence, the second confidence, the second original signal and the second voice enhanced signal.

[0083] The second voice enhanced signal is a signal obtained by performing voice enhancement on the second original signal.

[0084] In some embodiments, the first original signal and the second original signal are extracted from the same original voice signal.

[0085] In an exemplary embodiment, the second original signal can refer to the original voice signal collected by the microphone, or can refer to part of the original voice signal collected by the microphone.

[0086] In an exemplary embodiment, the original voice signal is "Xiao X, Xiao X, please help me open the window", the first original signal can be "Xiao X, Xiao X", and the second original signal can be "please help me open the window". The first original signal and the second original signal can also be the same, both being "Xiao X, Xiao X", or the first original signal and the second original signal can also be the same, both being "please help me open the window".

[0087] Since the first original signal and the second original signal are extracted from the same original voice signal, the second voice enhanced signal is a signal obtained by performing voice enhancement on the second original signal, and the first confidence and the second confidence can also represent the confidence of the second original signal and the second voice enhanced signal.

[0088] In an example embodiment, the first confidence is used to indicate a confidence of the second original signal for a downstream task, and the second confidence is used to indicate a confidence of the second speech enhanced signal for the downstream task. For example, the second original signal is used to wake up a target device, the first confidence is used to indicate a confidence of the second original signal for the task of waking up the target device, and the second confidence is used to indicate a confidence of the second speech enhanced signal for the task of waking up the target device. For another example, the second original signal is used to control a target device to perform a target operation, the first confidence is used to indicate a confidence of the second original signal for the task of controlling the target device to perform the target operation, and the second confidence is used to indicate a confidence of the second speech enhanced signal for the task of controlling the target device to perform the target operation.

[0089] Based on the first confidence, the second confidence, the second original signal and the second speech enhanced signal, the target signal is determined, the second speech enhanced signal is considered in the process of determining the target signal, the advantage brought by speech enhancement can be obtained to a certain extent, and the second original signal and the second speech enhanced signal and the first confidence and the second confidence are comprehensively considered in the process of determining the target signal, and the speech distortion brought by speech enhancement can be reduced.

[0090] In some embodiments, as shown in FIG. 1, FIG. 4 Step 102 includes:

[0091] Step 401, based on the first confidence, a first weight corresponding to the second original signal is determined.

[0092] Step 402, based on the second confidence, a second weight corresponding to the second speech enhanced signal is determined.

[0093] In an example embodiment, the first confidence is between 0 and 1, the second confidence is between 0 and 1, and the sum of the first confidence and the second confidence is equal to 1. The first confidence can be directly determined as the first weight corresponding to the second original signal, and the second confidence can be directly determined as the second weight corresponding to the second speech enhanced signal.

[0094] In an example embodiment, the first confidence can be any value, and the second confidence can be any value. The sum of the first confidence and the second confidence can be determined as a total confidence, the first confidence divided by the total confidence can be determined as the first weight corresponding to the second original signal, and the second confidence divided by the total confidence can be determined as the second weight corresponding to the second speech enhanced signal.

[0095] Step 403, based on the second original signal, the first weight, the second speech enhanced signal and the second weight, weighted summation is performed to obtain the target signal.

[0096] In the example embodiment, the second original signal is denoted as X1, the second speech enhanced signal is denoted as X2, the first weight is denoted as G1, the second weight is denoted as G2, and the target signal is denoted as X3. The specific calculation formula of the target signal X3 is as follows:

[0097] X3 = X1 * G1 + X2 * G2

[0098] Based on the first confidence, the first weight corresponding to the second original signal is determined, based on the second confidence, the second weight corresponding to the second speech enhanced signal is determined, and based on the second original signal, the first weight, the second speech enhanced signal and the second weight, weighted summation is performed to obtain the target signal. In the process of determining the target signal, the second speech enhanced signal is considered, which can to some extent obtain the advantage brought by speech enhancement. Moreover, in the process of determining the target signal, based on the second original signal, the first weight, the second speech enhanced signal and the second weight, weighted summation is performed to obtain the target signal, which can reduce the speech distortion brought by speech enhancement.

[0099] In other embodiments, step 102 includes: multiplying the first confidence by the second original signal to obtain a product, adding the product to a product obtained by multiplying the second confidence by the second speech enhanced signal to obtain an intermediate signal; and determining the target signal based on the intermediate signal.

[0100] In the example embodiment, determining the target signal based on the intermediate signal can be determining the target signal as a result obtained by multiplying the intermediate signal by a preset value, or as a result obtained by dividing the intermediate signal by a preset value, or as a result obtained by adding a preset value to the intermediate signal, or as a result obtained by subtracting a preset value from the intermediate signal. The application does not limit this.

[0101] Step 103, based on the target signal, performing voice interaction with the target device.

[0102] Based on the target signal, voice interaction with the target device can further improve the accuracy of voice interaction on the basis of reducing the speech distortion brought by speech enhancement.

[0103] In the example embodiment, the target device can be a vehicle-mounted device, a smart speaker, a smart air conditioner or other smart home devices, or other devices capable of being controlled by voice. In the following embodiments, the target device is taken as a vehicle-mounted device for explanation and illustration, but the application is not limited thereto.

[0104] In the example embodiment, the voice interaction includes at least one of voice wake-up and voice recognition.

[0105] In an example embodiment, in a case where the voice interaction includes voice wake-up, the step 103 comprises: in a case where the target signal indicates to wake up the target device, waking up the target device; in a case where the target signal indicates not to wake up the target device, not waking up the target device. In a case where the voice interaction includes voice recognition, the step 103 comprises: in a case where the target signal indicates to control the target device to perform a target operation, recognizing a target instruction corresponding to the target signal, and controlling the target device to perform a target operation corresponding to the target instruction. In a case where the voice interaction includes voice wake-up and voice recognition, the step 103 comprises: based on the target signal, waking up the target device; recognizing a target instruction corresponding to the target signal, and controlling the target device to perform a target operation corresponding to the target instruction after waking up the target device.

[0106] In some embodiments, the first original signal and the second original signal are the same, and the first speech enhanced signal and the second speech enhanced signal are the same.

[0107] In an example embodiment, the first original signal and the second original signal are the same, and are used to wake up the target device, for example, both the first original signal and the second original signal are "Xiao X, Xiao X"; the first original signal and the second original signal are the same, and are used to control the target device to perform a target operation, for example, both the first original signal and the second original signal are "please help me open the window"; the first original signal and the second original signal are the same, and are used to wake up the target device, and after waking up the target device, the target device is controlled to perform a target operation, for example, both the first original signal and the second original signal are "Xiao X, Xiao X, please help me open the window".

[0108] In some embodiments, the first original signal and the second original signal are the same, and the first speech enhanced signal and the second speech enhanced signal are the same.

[0109] As shown in FIG. 5 , the flow of the voice interaction method mainly includes:

[0110] Step 501, determining a first confidence of a first original signal and a second confidence of a first speech enhanced signal.

[0111] The first speech enhanced signal is a signal obtained by performing speech enhancement on the first original signal.

[0112] In the example embodiment, the first confidence degree is used to indicate a confidence degree of the first original signal for the voice interaction task, and the second confidence degree is used to indicate a confidence degree of the first voice enhanced signal for the voice interaction task. For example, the first original signal is used to wake up a target device, the first confidence degree is used to indicate a confidence degree of the first original signal for the task of waking up the target device, and the second confidence degree is used to indicate a confidence degree of the first voice enhanced signal for the task of waking up the target device; for another example, the first original signal is used to control the target device to perform a target operation, the first confidence degree is used to indicate a confidence degree of the first original signal for the task of controlling the target device to perform the target operation, and the second confidence degree is used to indicate a confidence degree of the first voice enhanced signal for the task of controlling the target device to perform the target operation.

[0113] In step 502, a target signal is determined based on the first confidence degree, the second confidence degree, the first original signal, and the first voice enhanced signal.

[0114] In step 503, a voice interaction with the target device is performed based on the target signal.

[0115] The first confidence degree of the first original signal and the second confidence degree of the first voice enhanced signal are determined, wherein the first voice enhanced signal is a signal obtained by performing voice enhancement on the first original signal, the confidence degree of the signal before voice enhancement and the confidence degree of the signal after voice enhancement can be obtained, the target signal is determined based on the first confidence degree, the second confidence degree, the first original signal, and the first voice enhanced signal, the first voice enhanced signal is considered in the process of determining the target signal, the advantage brought by voice enhancement can be obtained to a certain extent, moreover, the first original signal and the first voice enhanced signal and the first confidence degree and the second confidence degree are comprehensively considered in the process of determining the target signal, the voice distortion brought by voice enhancement can be reduced, and the voice interaction with the target device is performed based on the target signal, which can further improve the accuracy of voice interaction on the basis of reducing the voice distortion brought by voice enhancement.

[0116] In some embodiments, the first original signal and the second original signal are different, and the first voice enhanced signal and the second voice enhanced signal are different.

[0117] In an example embodiment, the first original signal and the second original signal are different, the first original signal is used to wake up the target device, and the second original signal is used to control the target device to perform a target operation. For example, the first original signal is "Xiaoxia, Xiaoxia", and the second original signal is "Please help me open the window". The first original signal and the second original signal are different, the first original signal is used to wake up the target device, and the second original signal is used to wake up the target device. For example, the first original signal is "Xiaoxia", and the second original signal is "Xiaoxia ya". The first original signal and the second original signal are different, the first original signal is used to control the target device to perform a first operation, and the second original signal is used to control the target device to perform a second operation. For example, the first original signal is "Please help me open the window", and the second original signal is "Please help me play music".

[0118] In some embodiments, the first original signal and the second original signal are different, and the first voice enhanced signal and the second voice enhanced signal are different. The first original signal is used to wake up the target device, and the second original signal is used to control the target device to perform a target operation.

[0119] As shown in FIG. 6 , the flow of the voice interaction method mainly includes:

[0120] Step 601, after waking up the target device, determining a first confidence of a first original signal and a second confidence of a first voice enhanced signal, wherein the first original signal is used to wake up the target device.

[0121] The first voice enhanced signal is a signal obtained by performing voice enhancement on the first original signal.

[0122] In an example embodiment, the target device can be woken up based on the first original signal, or the target device can be woken up based on the first voice enhanced signal.

[0123] Step 602, determining a target signal based on the first confidence, the second confidence, a second original signal, and a second voice enhanced signal, wherein the second original signal is used to control the target device to perform a target operation.

[0124] The second voice enhanced signal is a signal obtained by performing voice enhancement on the second original signal.

[0125] The first original signal and the second original signal are extracted from the same original voice signal.

[0126] In an example embodiment, the first confidence is used to indicate the confidence of the second original signal for the task of identifying a target instruction, and the second confidence is used to indicate the confidence of the second voice enhanced signal for the task of identifying a target instruction.

[0127] In step 603, the target signal corresponding to the target instruction is identified, and the target device is controlled to perform the target operation corresponding to the target instruction.

[0128] After the target device is woken up, a first confidence of the first original signal and a second confidence of the first voice enhanced signal are determined, wherein the first original signal is used to wake up the target device, and the first voice enhanced signal is a signal obtained by performing voice enhancement on the first original signal. The confidence of the signal before voice enhancement and the confidence of the signal after voice enhancement can be obtained. Since the first original signal and the second original signal are extracted from the same original voice signal, the second voice enhanced signal is a signal obtained by performing voice enhancement on the second original signal, the first confidence and the second confidence can also represent the confidence of the second original signal and the second voice enhanced signal. Based on the first confidence, the second confidence, the second original signal and the second voice enhanced signal, the target signal is determined. In the process of determining the target signal, the second voice enhanced signal is considered, which can to some extent obtain the advantage brought by voice enhancement. Moreover, in the process of determining the target signal, the second original signal and the second voice enhanced signal and the first confidence and the second confidence are comprehensively considered, which can reduce the voice distortion brought by voice enhancement. The target instruction corresponding to the target signal is identified, and the target device is controlled to perform the target operation corresponding to the target instruction, which can further improve the accuracy of voice recognition on the basis of reducing the voice distortion brought by voice enhancement.

[0129] In some embodiments, the flow of the voice interaction method mainly includes: inputting the first original signal into the voice enhancement model to obtain the first voice enhanced signal; inputting at least one of the first original signal and the first voice enhanced signal into the wake-up model to obtain a voice wake-up result; in the case that the voice wake-up result is to wake up the target device, waking up the target device; after waking up the target device, extracting a first feature of the first original signal and a second feature of the first voice enhanced signal, inputting the first feature and the second feature into the confidence model to obtain a first confidence of the first original signal and a second confidence of the first voice enhanced signal; inputting the second original signal into the voice enhancement model to obtain the second voice enhanced signal; determining the target signal based on the first confidence, the second confidence, the second original signal and the second voice enhanced signal; inputting the target signal into the recognition model to obtain the target instruction corresponding to the target signal, and controlling the target device to perform the target operation corresponding to the target instruction.

[0130] The first original signal is used to wake up the target device; the second original signal is used to control the target device to perform the target operation; and the first original signal and the second original signal are extracted from the same original voice signal.

[0131] In some embodiments, the training process of the speech enhancement model, the wake-up model, the confidence model and the recognition model includes: pre-training the speech enhancement model, the wake-up model and the recognition model; after the speech enhancement model, the wake-up model and the recognition model complete pre-training, training the confidence model, and fixing the parameters of the speech enhancement model, the wake-up model and the recognition model when training the confidence model. The training process of the confidence model can be referred to steps 301 to 305, which will not be described here.

[0132] In summary, in the present application, the first confidence of the first original signal and the second confidence of the first speech enhanced signal are determined, wherein the first speech enhanced signal is a signal obtained by performing speech enhancement on the first original signal, the confidence of the signal before speech enhancement and the confidence of the signal after speech enhancement can be obtained. Since the second speech enhanced signal is a signal obtained by performing speech enhancement on the second original signal, the first confidence and the second confidence can also represent the confidence of the second original signal and the second speech enhanced signal. Based on the first confidence, the second confidence, the second original signal and the second speech enhanced signal, the target signal is determined. In the process of determining the target signal, the second speech enhanced signal is considered, which can to some extent obtain the advantage brought by speech enhancement. Moreover, in the process of determining the target signal, the second original signal and the second speech enhanced signal and the first confidence and the second confidence are comprehensively considered, which can reduce the speech distortion brought by speech enhancement. Based on the target signal, the voice interaction with the target device can further improve the accuracy of voice interaction on the basis of reducing the speech distortion brought by speech enhancement.

[0133] Example Device

[0134] Correspondingly, the present application also provides a voice interaction device, as shown in the following FIG. 7 The voice interaction device includes:

[0135] The first processing unit 701 is configured to determine the first confidence of the first original signal and the second confidence of the first speech enhanced signal; wherein the first speech enhanced signal is a signal obtained by performing speech enhancement on the first original signal.

[0136] The second processing unit 702 is configured to determine the target signal based on the first confidence, the second confidence, the second original signal and the second speech enhanced signal; wherein the second speech enhanced signal is a signal obtained by performing speech enhancement on the second original signal.

[0137] The voice interaction unit 703 is configured to perform voice interaction with the target device based on the target signal.

[0138] Optionally, the first original signal and the second original signal are extracted from the same original speech signal.

[0139] Optionally, the second processing unit 702 is specifically configured to:

[0140] determine a first weight corresponding to the second original signal based on the first confidence;

[0141] determine a second weight corresponding to the second speech enhanced signal based on the second confidence;

[0142] perform weighted summation based on the second original signal, the first weight, the second speech enhanced signal and the second weight to obtain a target signal.

[0143] Optionally, the first processing unit 701 is specifically configured to:

[0144] extract a first feature of the first original signal and a second feature of the first speech enhanced signal;

[0145] input the first feature and the second feature into a confidence model to obtain a first confidence of the first original signal and a second confidence of the first speech enhanced signal.

[0146] Optionally, the speech interaction device further comprises a confidence model training unit.

[0147] The confidence model training unit is configured to:

[0148] extract a first sample feature of a first sample original signal and a second sample feature of a first sample speech enhanced signal;

[0149] input the first sample feature and the second sample feature into a confidence model to obtain a first sample confidence of the first sample original signal and a second sample confidence of the first sample speech enhanced signal;

[0150] determine a sample speech interaction signal based on the first sample confidence, the second sample confidence, the second sample original signal and the second sample speech enhanced signal;

[0151] obtain a predicted speech interaction label based on the sample speech interaction signal;

[0152] train the confidence model based on the predicted speech interaction label and a real speech interaction label corresponding to the second sample original signal to obtain a trained confidence model;

[0153] The first sample voice enhanced signal is a signal obtained by performing voice enhancement on the first sample original signal; and the second sample voice enhanced signal is a signal obtained by performing voice enhancement on the second sample original signal.

[0154] Optionally, the first original signal and the second original signal are the same, and the first voice enhanced signal and the second voice enhanced signal are the same.

[0155] Optionally, the first original signal and the second original signal are different, and the first voice enhanced signal and the second voice enhanced signal are different; the first original signal is used to wake up the target device; and the second original signal is used to control the target device to perform a target operation.

[0156] The voice interaction apparatus provided in the embodiment belongs to the same application concept as the voice interaction method provided in the above-mentioned embodiments of the application, can execute the voice interaction method provided in any of the above-mentioned embodiments of the application, and has the corresponding function modules and beneficial effects of executing the voice interaction method. Technical details not described in detail in the embodiment can be referred to the specific processing content of the voice interaction method provided in the above-mentioned embodiments of the application, which will not be described here again.

[0157] The functions implemented by the first processing unit 701, the second processing unit 702 and the voice interaction unit 703 described above can be implemented by the same or different processors, and the embodiments of the application are not limited.

[0158] It should be understood that the units in the above apparatus can be implemented in the form of processor calling software. For example, the apparatus includes a processor connected with a memory, the memory stores instructions, and the processor calls the instructions stored in the memory to implement any of the above methods or to implement the functions of the units of the apparatus, wherein the processor can be a general processor, such as a CPU or a microprocessor, and the memory can be an internal memory of the apparatus or an external memory of the apparatus. Alternatively, the units in the apparatus can be implemented in the form of hardware circuit. The functions of part or all of the units can be implemented by designing the hardware circuit. The hardware circuit can be understood as one or more processors. For example, in one implementation, the hardware circuit is an ASIC, and the functions of part or all of the units are implemented by designing the logical relationship of elements in the circuit. For another example, in another implementation, the hardware circuit can be implemented by a PLD. Taking an FPGA as an example, it can include a large number of logic gate circuits, and the connection relationship between the logic gate circuits is configured by a configuration file, so as to implement the functions of part or all of the units. All units of the above apparatus can be implemented in the form of processor calling software, or implemented in the form of hardware circuit, or part of them are implemented in the form of processor calling software, and the remaining part is implemented in the form of hardware circuit.

[0159] In embodiments of the present application, the processor is a circuit with signal processing capability. In one implementation, the processor can be a circuit with instruction reading and running capability, such as a CPU, a microprocessor, a GPU, or a DSP, etc. In another implementation, the processor can implement certain functions through a logic relationship of a hardware circuit, which is fixed or can be reconfigured. For example, the processor is an ASIC or a PLD implemented hardware circuit, such as an FPGA, etc. In the reconfigurable hardware circuit, the processor loads the configuration document to implement the hardware circuit configuration process, which can be understood as the process of the processor loading instructions to implement the functions of the above part or all units. In addition, it can also be a hardware circuit designed for artificial intelligence, which can be understood as an ASIC, such as NPU, TPU, DPU, etc.

[0160] It can be seen that each unit in the above apparatus can be one or more processors (or processing circuits) configured to implement the above methods, such as: CPU, GPU, NPU, TPU, DPU, microprocessor, DSP, ASIC, FPGA, or a combination of at least two of these processor forms.

[0161] In addition, each unit in the above apparatus can be integrated together or can be independently implemented. In one implementation, these units are integrated together to implement in the form of SOC. The SOC can include at least one processor for implementing any of the above methods or the functions of each unit of the apparatus, and the at least one processor can be different, such as including a CPU and an FPGA, a CPU and an artificial intelligence processor, a CPU and a GPU, etc.

[0162] Example Electronic Device

[0163] An embodiment of the present application provides an electronic device, as shown in FIG. 8 The device includes:

[0164] a memory 200 and a processor 210;

[0165] The memory 200 is connected with the processor 210, and is configured to store programs.

[0166] The processor 210 is configured to implement the voice interaction method disclosed in any of the above embodiments by running the programs stored in the memory 200.

[0167] Specifically, the above electronic device can further include a bus, a communication interface 220, an input device 230, and an output device 240.

[0168] The processor 210, the memory 200, the communication interface 220, the input device 230 and the output device 240 are connected with each other through a bus. Among them:

[0169] The bus can include a path for transmitting information between various components of the computer system.

[0170] The processor 210 can be a general-purpose processor, such as a general-purpose central processing unit (CPU), a microprocessor, etc., or can be an application-specific integrated circuit (ASIC), or one or more integrated circuits for controlling the execution of programs of the present application. It can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a ready-to-use programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component.

[0171] The processor 210 can include a main processor and can also include a baseband chip, a modem, etc.

[0172] The memory 200 stores programs for executing the technical solutions of the present application, and can also store operating systems and other key services. Specifically, the program can include program code, and the program code includes computer operation instructions. More specifically, the memory 200 can include read-only memory (ROM), other types of static storage devices that can store static information and instructions, random access memory (RAM), other types of dynamic storage devices that can store information and instructions, disk storage, flash, etc.

[0173] The input device 230 can include a device that receives data and information input by a user, such as a keyboard, a mouse, a camera, a scanner, a light pen, a voice input device, a touch screen, a pedometer or a gravity sensor, etc.

[0174] The output device 240 can include a device that allows information to be output to a user, such as a display screen, a printer, a speaker, etc.

[0175] The communication interface 220 can include a device using any transceiver, such as an Ethernet, a radio access network (RAN), a wireless local area network (WLAN), etc., to communicate with other devices or communication networks.

[0176] The processor 210 executes the program stored in the memory 200 and calls other devices, which can be used to implement each step of any one of the voice interaction methods provided by the embodiments of the present application.

[0177] Example Computer Program Product and Storage Medium

[0178] In addition to the method and device described above, the embodiments of the present application can also be a computer program product, which includes computer program instructions, which, when executed by a processor, cause the processor to perform the steps in the voice interaction method according to various embodiments of the present application described in any of the embodiments in the specification.

[0179] The computer program product can be written in any combination of one or more programming languages to perform the operations of the embodiments of the present application, including object-oriented programming languages, such as Java, C++, and conventional procedural programming languages, such as "C" language or similar programming languages. The program code can be executed entirely on a user computing device, partially on a user device, as an independent software package, partially on a user computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0180] In addition, the embodiments of the present application can also be a storage medium having a computer program stored thereon, which is executed by a processor to perform the steps in the voice interaction method according to various embodiments of the present application described in any of the embodiments in the specification, which can specifically implement the following steps:

[0181] Step 101, determining a first confidence of a first original signal and a second confidence of a first speech enhanced signal.

[0182] Wherein the first speech enhanced signal is a signal obtained by performing speech enhancement on the first original signal.

[0183] Step 102, determining a target signal based on the first confidence, the second confidence, a second original signal and a second speech enhanced signal.

[0184] Wherein the second speech enhanced signal is a signal obtained by performing speech enhancement on the second original signal.

[0185] Step 103, performing voice interaction with a target device based on the target signal.

[0186] For each of the foregoing method embodiments, in order to simply describe, it is expressed as a combination of a series of actions, but those skilled in the art should know that the present application is not limited by the order of the described actions, because according to the present application, certain steps can be performed in other order or simultaneously. Secondly, those skilled in the art should know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present application.

[0187] It should be noted that the various embodiments of the present specification are described in progressive manner, and each embodiment focuses on the differences from other embodiments, and the same and similar parts between embodiments can be mutually referred to.

[0188] The steps in the methods of the embodiments of the present application can be adjusted in sequence, combined and deleted according to actual needs, and the technical features recorded in each embodiment can be replaced or combined.

[0189] The modules and sub-modules in the devices and terminals in the embodiments of the present application can be combined, divided and deleted according to actual needs.

[0190] In several embodiments of the present application, it should be understood that the disclosed terminals, devices and methods can be implemented in other ways. For example, the terminal embodiments described above are only schematic, for example, the division of modules or sub-modules is only a logical function division, and actual implementation can have another division manner, for example, a plurality of sub-modules or modules can be combined or integrated into another module, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the displayed or discussed each other can be indirect coupling or communication connection between some interfaces, devices or modules, which can be electrical, mechanical or other forms.

[0191] The modules or sub-modules described as separate components can or can not be physically separated, and the components of the modules or sub-modules can or can not be physical modules or sub-modules, that is, they can be located in one place, or can be distributed to multiple network modules or sub-modules. Some or all of the modules or sub-modules can be selected according to actual needs to achieve the purpose of the present embodiment.

[0192] In addition, each functional module or sub-module in each embodiment of the present application can be integrated in one processing module, or each module or sub-module can exist physically, or two or more modules or sub-modules can be integrated in one module. The above integrated module or sub-module can be realized in the form of hardware or software functional module or sub-module.

[0193] Those skilled in the art will further appreciate that the units and algorithm steps of the various examples described in connection with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or combinations of both. To clearly illustrate this interchangeability of hardware and software, various examples have been described herein in terms of their functionality, their composition, and their manner of operation. Whether such functionality is implemented in hardware or software depends on the particular application and design constraints imposed on the overall system. Skilled persons can implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present application.

[0194] The steps of a method or algorithm described in connection with the embodiments disclosed herein can be embodied directly in hardware, in a software module executed by a processor, or in a combination of the two. A software module can reside in random access memory (RAM), flash memory, read-only memory (ROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), registers, hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.

[0195] Finally, it should be noted that the terms "first", "second", and the like, herein do not denote any order, quantity, combination, or importance, but rather are used to distinguish one element from another, and are more especially used for the purpose of identification in claims. In addition, the terms "comprises", "comprising", or any other variation thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article, or apparatus. An element proceeded by "comprises... a" does not, without more constraints, exclude the existence of additional identical elements in the process, method, article, or apparatus that comprises the element.

[0196] The above description of disclosed embodiments provides enabling disclosure sufficient for others to practice the application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and generic principles defined herein can be applied to other embodiments without departing from the spirit or scope of the application. Thus, the present application is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A voice interaction method, characterized in that, The method comprises: determining a first confidence of a first original signal and a second confidence of a first speech enhanced signal; wherein the first speech enhanced signal is a signal obtained by performing speech enhancement on the first original signal; determining a target signal based on the first confidence, the second confidence, a second original signal and a second speech enhanced signal; wherein the second speech enhanced signal is a signal obtained by performing speech enhancement on the second original signal; the first original signal and the second original signal are extracted from the same original speech signal; the first original signal and the second original signal are different, and the first speech enhanced signal and the second speech enhanced signal are different; the first original signal is used to wake up a target device; and the second original signal is used to control the target device to perform a target operation; based on the target signal, performing voice interaction with the target device.

2. The voice interaction method of claim 1, wherein, The method of determining a target signal based on the first confidence, the second confidence, a second original signal and a second speech enhanced signal comprises: determining a first weight corresponding to the second original signal based on the first confidence; determining a second weight corresponding to the second speech enhanced signal based on the second confidence; performing weighted summation based on the second original signal, the first weight, the second speech enhanced signal and the second weight to obtain a target signal.

3. The voice interaction method of claim 1, wherein, The method of determining a first confidence of a first original signal and a second confidence of a first speech enhanced signal comprises: extracting a first feature of the first original signal and a second feature of the first speech enhanced signal; inputting the first feature and the second feature into a confidence model to obtain the first confidence of the first original signal and the second confidence of the first speech enhanced signal.

4. The voice interaction method of claim 3, wherein, The training process of the confidence model comprises: extracting a first sample feature of a first sample original signal and a second sample feature of a first sample speech enhanced signal; inputting the first sample feature and the second sample feature into a confidence model to obtain a first sample confidence of the first sample original signal and a second sample confidence of the first sample speech enhanced signal; determining a sample voice interaction signal based on the first sample confidence, the second sample confidence, a second sample original signal and a second sample speech enhanced signal; obtaining a predicted voice interaction label based on the sample voice interaction signal; training the confidence model based on the predicted voice interaction label and a real voice interaction label corresponding to the second sample original signal to obtain a trained confidence model; wherein the first sample speech enhanced signal is a signal obtained by performing speech enhancement on the first sample original signal; and the second sample speech enhanced signal is a signal obtained by performing speech enhancement on the second sample original signal.

5. A voice interaction device, characterized by The method comprises: a first processing unit configured to determine a first confidence of a first original signal and a second confidence of a first speech enhanced signal; wherein the first speech enhanced signal is a signal obtained by performing speech enhancement on the first original signal; A second processing unit is configured to determine a target signal based on the first confidence, the second confidence, a second original signal, and a second speech enhanced signal, wherein the second speech enhanced signal is a signal obtained by performing speech enhancement on the second original signal, the first original signal and the second original signal are extracted from a same original speech signal, the first original signal and the second original signal are different, the first speech enhanced signal and the second speech enhanced signal are different, the first original signal is used to wake up a target device, and the second original signal is used to control the target device to perform a target operation. A voice interaction unit is configured to perform voice interaction with the target device based on the target signal.

6. An electronic device, comprising: comprise a memory and a processor; the memory is connected with the processor and is configured to store a program; the processor is configured to realize the voice interaction method in any one of claims 1 to 4 by running the program in the memory.

7. A storage medium, characterized by The storage medium has a computer program stored thereon, and the computer program, when being run by a processor, realizes the voice interaction method in any one of claims 1 to 4.

8. A computer program product, characterised in that, comprise computer program instructions, and the computer program instructions, when being run by a processor, cause the processor to perform the voice interaction method in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Voice processing method and device, electronic equipment and readable storage medium

    CN118918906A

  • Distributed voice wake-up method and apparatus, storage medium, and electronic apparatus

    WO2023231552A1