Voice processing method and device, electronic equipment and storage medium

By using multiple voice acquisition devices to separate the voice signals to be processed in the voice wake-up mode, the voice signals of the target path are determined, and the problem of poor anti-interference effect in the prior art is solved, and more accurate voice recognition is achieved.

CN120164483APending Publication Date: 2025-06-17TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311737397.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-15
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

The existing voice recognition technology has poor anti-interference effect during noise reduction processing, resulting in inaccurate speech recognition results.

Method used

By using at least two voice acquisition devices in the voice wake-up mode to collect the to-process voice signals and separate them through the separation module, it is determined that the voice signal of the target path is the target voice signal.

Benefits of technology

It improves the anti-interference ability of speech processing, enhances the voice separation effect, and the target voice signal obtained is purer, has better noise reduction effect, and improves the accuracy of subsequent speech recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120164483A_ABST
    Figure CN120164483A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a voice processing method and device, electronic equipment and a storage medium, and relates to the fields of data processing, artificial intelligence, cloud technology and the like. The method comprises the following steps: in a voice wake-up mode, a first terminal acquires to-be-processed voice signals of a target object acquired by each voice acquisition device, separates each to-be-processed voice to obtain third voice signals of two paths, and determines the third voice signal of the target path as a target voice signal. Wherein the target channel is determined based on whether the voice signals of the two channels obtained by separating the wake-up voice signals contain the target wake-up word, and the wake-up voice signals comprise the voice signals of the target object containing the target wake-up word and collected by the voice collection devices; the separation method of the wake-up voice signal is the same as the separation method of the to-be-processed voice signal. Based on the method, the anti-interference capability of voice separation is improved, and a better denoising effect is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of computer technology and may involve fields such as data processing, artificial intelligence, and cloud technology. Specifically, this application relates to a voice processing method, apparatus, electronic device, and storage medium. Background Art

[0002] With the rapid development of technology, speech recognition technology is also widely used in various application fields, such as smartphones, smart homes, and car navigation. Users can interact with devices through voice, bringing a convenient experience to users.

[0003] Taking speech recognition on a smartphone as an example, a smartphone is usually equipped with two microphones, including a main microphone at the bottom and a secondary microphone at the top or on the back. When performing voice processing, two-channel voice signals can be received through the two microphones, and noise reduction processing can be performed on the two-channel voice signals based on an adaptive beamforming algorithm. After that, speech recognition is performed on the noise-reduced voice signals to obtain a speech recognition result.

[0004] However, the anti-interference effect of the above noise reduction method is poor, resulting in inaccurate subsequent speech recognition results. Summary of the Invention

[0005] The purpose of the embodiments of this application is to provide a voice processing method, apparatus, electronic device, and storage medium that can effectively improve the anti-interference ability. To achieve this purpose, the technical solutions provided by the embodiments of this application are as follows:

[0006] On the one hand, the embodiments of this application provide a voice processing method, which is executed by a first terminal. The first terminal is configured with at least two voice collection devices. The method includes:

[0007] In the voice wake-up mode, obtain the to-be-processed voice signals of the target object collected by each of the voice collection devices respectively;

[0008] Separate each of the to-be-processed voice signals to obtain third voice signals of two channels;

[0009] Determine the third voice signal of the target channel in the two channels as the target voice signal of the target object;

[0010] Among them, the target channel is determined based on whether the voice signals of the two channels separated from the wake-up voice signal contain a target wake-up word. The wake-up voice signal includes the voice signals of the target object containing the target wake-up word collected by each of the voice collection devices. The separation method of the wake-up voice signal is the same as the separation method of the to-be-processed voice signal.

[0011] On the other hand, an embodiment of the present application further provides a voice processing device, which is deployed on a first terminal, and the first terminal is configured with at least two voice acquisition devices. The device includes:

[0012] An acquisition module, configured to acquire, in a voice wake-up mode, the to-be-processed voice signals of a target object respectively acquired by each of the voice acquisition devices;

[0013] A separation module, configured to separate each of the to-be-processed voice signals to obtain third voice signals of two channels;

[0014] A target voice determination module, configured to determine the third voice signal of the target channel in the two channels as the target voice signal of the target object;

[0015] Wherein, the target channel is determined based on whether the voice signals of the two channels separated from the wake-up voice signal contain a target wake-up word. The wake-up voice signal includes the voice signals of the target object containing the target wake-up word acquired by each of the voice acquisition devices, and the separation method of the wake-up voice signal is the same as that of the to-be-processed voice signal.

[0016] Optionally, the voice wake-up mode is enabled in the following manner:

[0017] In response to a first target operation on the first terminal, the voice wake-up mode is enabled;

[0018] The acquisition module may be configured to:

[0019] In the voice wake-up mode, when the voice signals of the target object containing the target wake-up word are acquired by each of the voice acquisition devices, it is determined that the voice wake-up is successful, and the to-be-processed voice signals of the target object are respectively acquired by each of the voice acquisition devices.

[0020] Optionally, the target channel is determined in the following manner:

[0021] In the voice wake-up mode, the voice signals of the target object containing the target wake-up word acquired by each of the voice acquisition devices are determined as the wake-up voice signals;

[0022] The wake-up voice signals are separated to obtain fifth voice signals of two channels;

[0023] The fifth voice signals of the two channels are input into a trained voice classification model to obtain the classification results of the fifth voice signals of each channel, and the classification results represent whether the fifth voice signals contain the target wake-up word;

[0024] The channel where the fifth voice signal containing the target wake-up word is located is determined as the target channel.

[0025] Optionally, the device further includes a filtering processing module, and the filtering processing module can be used for:

[0026] In a first processing mode, obtain the to-be-processed voice signals of the target object respectively collected by each of the voice collection devices; wherein, the first processing mode is enabled in response to a second target operation on the first terminal;

[0027] Through a trained voice separation model, separate the interfering voice signals and the first voice signal of the target object in each of the to-be-processed voice signals, where the interfering voice signals are other voice signals in the to-be-processed voice signals except the first voice signal of the target object;

[0028] Determine the filtering weight coefficients according to the interfering voice signals and the first voice signal of the target object in each of the to-be-processed voice signals;

[0029] Perform filtering processing on each of the to-be-processed voice signals based on the filtering weight coefficients to obtain the target voice signals.

[0030] Optionally, the filtering processing module can be used for:

[0031] Perform voice separation on each of the to-be-processed voice signals respectively through a trained voice separation model to obtain the background noise signals and the human voice signals in each of the to-be-processed voice signals;

[0032] For each voice frame in the human voice signal of each to-be-processed voice signal, determine the signal energy value of the voice frame; and determine the first voice signal of the target object in the to-be-processed signal based on the signal energy values of each voice frame and a preset energy threshold;

[0033] Wherein, the interfering voice signals include the background noise signals.

[0034] Optionally, the voice separation model is trained in the following manner:

[0035] Obtain a plurality of first training samples; each of the first training samples includes a first sample voice signal, the sample voice signal of the sample object corresponding to the first sample voice signal, and an interfering voice signal;

[0036] Perform a training operation on a first neural network model based on the plurality of first training samples until a first training end condition is satisfied to obtain a trained voice separation model, and the training operation includes:

[0037] Input the first sample voice signals of each first training sample into the first neural network model, and obtain the predicted sample voice signals and predicted interference voice signals corresponding to each first sample voice signal through the first neural network model;

[0038] Determine the first training loss based on the predicted sample voice signals, predicted interference voice signals, sample voice signals, and interference voice signals corresponding to each first sample voice signal;

[0039] Adjust the model parameters in the first neural network model based on the first training loss.

[0040] Optionally, the voice classification model is trained in the following manner:

[0041] Obtain multiple second sample voice signals with sample labels; the sample labels of the second sample voice signals indicate whether the second sample voice signals contain the target wake-up word;

[0042] Perform a training operation on the second neural network model based on multiple second sample voice signals until the second training end condition is met, and obtain the trained voice classification model. The training operation includes:

[0043] Input each second sample voice signal into the second neural network model, and obtain the predicted classification results of each second sample voice signal through the second neural network model;

[0044] Determine the second training loss based on the predicted classification results of each second sample voice signal and the sample labels of each second sample voice signal;

[0045] Adjust the model parameters in the second neural network model based on the second training loss.

[0046] Optionally, the filtering processing module can be used for:

[0047] The determining of the filtering weight coefficients according to the interference voice signals in each of the to-be-processed voice signals and the first voice signal of the target object includes:

[0048] Determine the noise covariance matrix according to the interference voice signals in each of the to-be-processed voice signals;

[0049] Determine the target human voice covariance matrix according to the first voice signal of the target object in each of the to-be-processed voice signals; based on the target human voice covariance matrix, determine the direction-of-arrival vector of the target human voice;

[0050] Determine the filtering weight coefficients according to the noise covariance matrix and the direction-of-arrival vector of the target human voice.

[0051] On the other hand, an embodiment of the present application provides a voice processing method, which is executed by a first terminal. The first terminal is configured with at least two voice collection devices, and the method includes:

[0052] Obtain the to-be-processed voice signals of the target object respectively collected by each of the voice collection devices;

[0053] Separate the interfering voice signals and the first voice signal of the target object in each of the to-be-processed voice signals through a trained voice separation model. The interfering voice signals are other voice signals in the to-be-processed voice signals except the first voice signal of the target object;

[0054] Determine the filtering weight coefficients according to the interfering voice signals and the first voice signal of the target object in each of the to-be-processed voice signals;

[0055] Perform filtering processing on each of the to-be-processed voice signals based on the filtering weight coefficients to obtain the target voice signal.

[0056] Optionally, the separating the interfering voice signals and the first voice signal of the target object in each of the to-be-processed voice signals through a trained voice separation model includes:

[0057] Perform voice separation on each of the to-be-processed voice signals respectively through a trained voice separation model to obtain the background noise signals and the human voice signals in each of the to-be-processed voice signals;

[0058] For each voice frame in the human voice signal of each to-be-processed voice signal, determine the signal energy value of the voice frame; and determine the first voice signal of the target object in the to-be-processed signal based on the signal energy values of the voice frames and a preset energy threshold;

[0059] Wherein, the interfering voice signals include the background noise signals.

[0060] Optionally, the determining the filtering weight coefficients according to the interfering voice signals and the first voice signal of the target object in each of the to-be-processed voice signals includes:

[0061] Determine the noise covariance matrix according to the interfering voice signals in each of the to-be-processed voice signals;

[0062] Determine the target human voice covariance matrix according to the first voice signal of the target object in each of the to-be-processed voice signals; based on the target human voice covariance matrix, determine the direction-of-arrival vector of the target human voice;

[0063] Determine the filtering weight coefficient according to the noise covariance matrix and the direction-of-arrival vector of the target human voice.

[0064] Optionally, the voice separation model is trained in the following manner:

[0065] Obtain a plurality of first training samples; each of the first training samples includes a first sample voice signal, a sample voice signal of a sample object corresponding to the first sample voice signal, and an interfering voice signal;

[0066] Perform a training operation on the first neural network model based on the plurality of first training samples until a first training end condition is satisfied, to obtain a trained voice separation model, where the training operation includes:

[0067] Input the first sample voice signals of the respective first training samples into the first neural network model, and obtain, through the first neural network model, predicted sample voice signals and predicted interfering voice signals corresponding to the first sample voice signals;

[0068] Determine a first training loss based on the predicted sample voice signals, predicted interfering voice signals, sample voice signals, and interfering voice signals corresponding to the respective first sample voice signals;

[0069] Adjust the model parameters in the first neural network model based on the first training loss.

[0070] Optionally, the method further includes:

[0071] Separate each of the to-be-processed voice signals to obtain third voice signals in two channels;

[0072] Determine the third voice signal of the target channel in the two channels as the target voice signal of the target object;

[0073] Wherein, the target channel is determined based on whether the voice signals in the two channels separated from the wake-up voice signal contain a target wake-up word, the wake-up voice signal includes the voice signals of the target object containing the target wake-up word collected by each of the voice acquisition devices, and the separation method of the wake-up voice signal is the same as the separation method of the to-be-processed voice signal.

[0074] Optionally, the target channel is determined in the following manner:

[0075] In the voice wake-up mode, determine the voice signals of the target object containing the target wake-up word collected by each of the voice acquisition devices as the wake-up voice signal;

[0076] Separate the wake-up voice signal to obtain fifth voice signals in two channels;

[0077] Input the fifth voice signals of the two channels into the trained voice classification model to obtain the classification results of the fifth voice signals of each channel, where the classification results indicate whether the fifth voice signals contain the target wake-up word;

[0078] Determine the channel where the fifth voice signal containing the target wake-up word is located as the target channel.

[0079] Optionally, the voice classification model is trained in the following manner:

[0080] Obtain multiple second sample voice signals with sample labels; the sample labels of the second sample voice signals indicate whether the second sample voice signals contain the target wake-up word;

[0081] Perform a training operation on the second neural network model based on multiple second sample voice signals until the second training end condition is met, to obtain the trained voice classification model. The training operation includes:

[0082] Input each second sample voice signal into the second neural network model, and obtain the predicted classification results of each second sample voice signal through the second neural network model;

[0083] Determine the second training loss based on the predicted classification results of each second sample voice signal and the sample labels of each second sample voice signal;

[0084] Adjust the model parameters in the second neural network model based on the second training loss.

[0085] On the other hand, an embodiment of the present application further provides a voice processing device, which is deployed on a first terminal, and the first terminal is configured with at least two voice collection devices. The device includes:

[0086] An acquisition module, configured to acquire the to-be-processed voice signals of the target object collected by each of the voice collection devices respectively;

[0087] A separation module, configured to separate the interference voice signals and the first voice signals of the target object in each of the to-be-processed voice signals through a trained voice separation model, where the interference voice signals are other voice signals in the to-be-processed voice signals except for the first voice signals of the target object;

[0088] A filtering weight determination module, configured to determine the filtering weight coefficients according to the interference voice signals and the first voice signals of the target object in each of the to-be-processed voice signals;

[0089] A target voice determination module, configured to perform filtering processing on each of the to-be-processed voice signals based on the filtering weight coefficients to obtain the target voice signal.

[0090] An embodiment of the present application further provides an electronic device, which includes a memory and a processor. A computer program is stored in the memory, and the processor executes the computer program to implement the method provided in any optional embodiment of the present application.

[0091] On the other hand, an embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the method provided in any optional embodiment of the present application is implemented.

[0092] On the other hand, an embodiment of the present application further provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, the method provided in any optional embodiment of the present application is implemented.

[0093] The beneficial effects brought by the technical solution provided by the embodiment of the present application are as follows:

[0094] The voice processing method provided by the embodiment of the present application, in the voice wake-up mode, separates the collected to-be-processed voice signals to obtain the voice signal of the target path. The voice separation effect is better, the cleaner target voice signal of the target object can be extracted, the anti-interference ability is stronger, the noise reduction effect is better, and the result of voice recognition based on the voice separation result is more accurate, better meeting the actual application requirements.

[0095] The voice processing method provided by the embodiment of the present application, in the first processing mode, separates each of the to-be-processed voice signals by using a voice separation model trained based on an artificial intelligence method, and the obtained interference voice signal and the first voice signal of the target object are more accurate. The filtering weight coefficients determined based on the separation result are also more accurate. Therefore, the filtering effect based on the filtering weight coefficients is better, and the cleaner target voice signal of the target object can be extracted. Description of the Drawings

[0096] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments of the present application.

[0097] Figure 1 It is a schematic flowchart of a voice processing method provided by an embodiment of the present application;

[0098] Figure 2 It is a schematic flowchart of wake-up detection provided by an embodiment of the present application;

[0099] Figure 3Schematic flowchart of the training method for the voice classification model provided by the embodiments of the present application;

[0100] Figure 4 Schematic structural diagram of a voice separation based on blind source separation provided by the embodiments of the present application;

[0101] Figure 5 Schematic flowchart of a voice processing method provided by the embodiments of the present application;

[0102] Figure 6 Schematic structural diagram of the linear array of the voice acquisition device provided by the embodiments of the present application;

[0103] Figure 7 Schematic flowchart of the training method for the voice separation model provided by the embodiments of the present application;

[0104] Figure 8 Schematic structural diagram of a voice separation based on the voice separation model provided by the embodiments of the present application;

[0105] Figure 9 Schematic structural diagram of a voice processing device provided by the embodiments of the present application;

[0106] Figure 10 Schematic structural diagram of a voice processing device provided by the embodiments of the present application;

[0107] Figure 11 Schematic structural diagram of an electronic device provided by the embodiments of the present application. Detailed implementation manners

[0108] The embodiments of the present application will be described below with reference to the accompanying drawings in the present application. It should be understood that the embodiments described below in conjunction with the accompanying drawings are exemplary descriptions for explaining the technical solutions of the embodiments of the present application, and do not constitute limitations on the technical solutions of the embodiments of the present application.

[0109] Those skilled in the art can understand that, unless specifically stated, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the terms "comprising" and "including" used in the embodiments of the present application mean that the corresponding features can be implemented as the presented features, information, data, steps, operations, elements and / or components, but do not exclude the implementation of other features, information, data, steps, operations, elements, components and / or their combinations supported by the technical field of the present application, etc. It should be understood that when we say an element is "connected" or "coupled" to another element, the one element can be directly connected or coupled to the other element, or it can mean that the one element and the other element establish a connection relationship through an intermediate element. In addition, the "connection" or "coupling" used herein can include wireless connection or wireless coupling. The term "and / or" used herein indicates at least one of the items defined by the term. For example, "A and / or B" can be implemented as "A", or implemented as "B", or implemented as "A and B". When describing multiple (two or more) items, if the relationship between the multiple items is not clearly defined, the multiple items can refer to one, multiple or all of the multiple items. For example, for the description of "parameter A includes A1, A2, A3", it can be implemented that parameter A includes A1 or A2 or A3, and it can also be implemented that parameter A includes at least two of the three items of parameter A1, A2, A3.

[0110] In one way, for the voice processing method, device, electronic device and storage medium provided by the embodiments of the present application, the voice separation model trained based on the artificial intelligence method is used to separate each voice signal to be processed, and the obtained interference voice signal and the first voice signal of the target object are more accurate. The filtering weight coefficient determined based on the separation result is also more accurate. Therefore, the filtering effect based on the filtering weight coefficient is better. In another way, in the voice wake-up mode, the voice signal of the target path is obtained by separating each voice signal to be processed collected, and the voice separation effect is better, and the obtained target voice signal is more accurate. Subsequently, the result of voice recognition based on the voice separation result is more accurate.

[0111] Among them, the method provided by the embodiments of the present application may involve artificial intelligence (AI) technology. For example, the background noise signal and the human voice signal in the voice signal are separated through the voice separation model, and the third voice signal of the two paths is classified through the voice classification model. Among them, the trained voice separation model and voice classification model can be trained based on the training sample set in the way of machine learning (ML).

[0112] Artificial intelligence uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, including theories, methods, technologies, and application systems for perceiving the environment, acquiring knowledge, and using knowledge to achieve the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable them to have the functions of perception, reasoning, and decision-making.

[0113] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning, autonomous driving, and intelligent transportation.

[0114] Among them, the key technologies of speech technology include automatic speech recognition technology, speech synthesis technology, and voiceprint recognition technology. Enabling computers to listen, see, speak, and feel is the future development direction of human-computer interaction, and speech has become one of the most promising human-computer interaction methods in the future.

[0115] Machine learning is an interdisciplinary subject involving multiple fields such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.

[0116] Optionally, the solution provided in the embodiments of this application may involve cloud technology. For example, the training methods of the above-mentioned voice separation model and voice classification model can be executed by a server, and the solution of the embodiments of this application can also be executed after the first terminal collects a voice signal and sends it to the server. Among them, the server can be a cloud server, and the data processing involved in the implementation process can be based on cloud technology, and the data storage involved in the implementation process can use cloud storage. For example, blind source separation of each voice signal to be processed can be implemented using cloud technology, voice recognition processing of the target voice signal can be implemented using cloud technology, and the training sample set of the voice separation model can be stored in the cloud server.

[0117] Among them, cloud technology is the general term for network technology, information technology, integration technology, management platform technology, application technology, etc. based on the cloud computing business model. It can form a resource pool, be used on demand, and is flexible and convenient. Cloud computing technology will become an important support. Cloud storage is a new concept extended and developed from the cloud computing concept. A distributed cloud storage system (hereinafter referred to as the storage system) refers to a storage system that combines a large number of different types of storage devices (storage devices are also called storage nodes) in the network through cluster applications, grid technology, and distributed file systems, and collaborates through application software or application interfaces to jointly provide data storage and business access functions to the outside world.

[0118] It should be noted that in the optional embodiments of the present application, for data related to object information (such as the voice signal of the collected target object), when the embodiments in the present application are applied to specific products or technologies, object permission or consent is required, and the collection, use, and processing of the relevant data need to comply with the relevant laws, regulations, and standards of the relevant countries and regions. That is to say, if the embodiments in the present application involve data related to the object, it needs to be obtained under the authorization and consent of the object, the authorization and consent of the relevant department, and compliance with the relevant laws, regulations, and standards of the country and region. If personal information is involved in the embodiments, the acquisition of all personal information requires the consent of the individual. If sensitive information is involved, the separate consent of the information subject is required, and the embodiments also need to be implemented under the authorization and consent of the object.

[0119] To better understand and illustrate the solutions of the embodiments of the present application, some technical terms involved in the embodiments of the present application are briefly described below.

[0120] VAD: Voice Activity Detection algorithm, used to detect whether there is a voice signal in the current voice signal, that is, to judge the input signal and distinguish the voice signal from various background noise signals.

[0121] MVDR: Minimum Variance Distortionless Response algorithm. By processing the input signal, the output signal has the minimum variance and a distortionless response. Since the variance of the output signal is the smallest, it can effectively remove noise and interference and improve the signal quality.

[0122] Blind source separation algorithm: It can separate the original signal components from the observed mixed signals. For example, the Information Maximization algorithm, the FastICA (Fixed Point Algorithm), and the JADE (Joint Approximate Diagonalization of Eigenmatrices Cardoso) algorithm.

[0123] MFCC: Mel-Frequency Cepstral Coefficient. The Mel frequency is extracted based on the auditory characteristics of the human ear and has a non-linear correspondence with the Hertz (Hz) frequency. The Mel-Frequency Cepstral Coefficient is a Hertz spectral feature determined based on this non-linear correspondence.

[0124] The technical solutions of the embodiments of the present application and the technical effects produced by the technical solutions of the present application will be described below through the description of several embodiments. It should be noted that the following embodiments can be referred to, learned from, or combined with each other. For the same terms, similar features, and similar implementation steps in different embodiments, they will not be described repeatedly.

[0125] The solution of the embodiment of the present application can be executed by a single electronic device or can be implemented by multiple electronic devices in cooperation. For example, it can be executed by the first terminal, or can be sent to the server for execution after the first terminal collects the voice signal.

[0126] Among them, the above server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server (which can be called the cloud) that provides cloud computing services. The terminal (which can also be called a user terminal or user device) can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart voice interaction device (such as a smart speaker), a wearable electronic device (such as a smart watch), a vehicle terminal, a smart home appliance (such as a smart TV), an AR / VR device, etc., but is not limited thereto. The terminal and the server can be directly or indirectly connected through wired or wireless communication methods, and the present application does not limit this.

[0127] Figure 1 The flowchart of a voice processing method provided by an embodiment of the present application is shown. This method can be executed by the first terminal, and the first terminal is configured with at least two voice collection devices.

[0128] As Figure 1 shown, the voice processing method provided by the embodiment of the present application may include the following steps S110 to step S130.

[0129] Step S110: In the voice wake-up mode, obtain the to-be-processed voice signals of the target object collected by each voice collection device respectively.

[0130] Among them, the voice collection device may be a sound sensor configured on the first terminal for collecting voices, such as a microphone, and the target object is the target user who gives voice instructions through the first terminal. The to-be-processed voice signals of the target object are voice signals containing the voices of the target object collected from the surrounding environment.

[0131] In the embodiment of the present application, this voice processing method may be executed in the voice wake-up mode based on a wake-up word. Specifically, for the voice wake-up mode, after recognizing that the target object utters a wake-up voice instruction containing the target wake-up word, this voice processing method is executed. Figure 2 FIG. is a schematic flow chart of wake-up detection. After the voice wake-up mode of the first terminal is turned on, the ambient voice signals in the surrounding environment are collected in real time, that is, the ambient voice signals in the surrounding environment are collected through each voice collection device in real time. Among them, the voice wake-up mode may be turned on in response to a first target operation on the first terminal. The first target operation may be a trigger operation such as a click or a slide on the voice wake-up mode opening entry. The present application does not limit this. For the collected ambient voice signals, the voice features of the ambient voice signals are extracted, and the voice features of the ambient voice signals are matched with a predefined wake-up word. If the match is successful, that is, a voice signal containing the target wake-up word is collected, then the voice wake-up is successful, and the to-be-processed voice signals of the target object are collected through the voice collection device on the first terminal. If the match fails, the wake-up fails, and the ambient voice signals in the surrounding environment continue to be collected.

[0132] Among them, the predefined target wake-up word may be one or more. When there are multiple predefined target wake-up words, the voice features of the ambient voice signal match any one of the target wake-up words, which is considered a successful wake-up. When the voice features of the ambient voice signals collected by any voice collection device match successfully, it can be considered that the wake-up is successful. The extracted voice features may be MFCC, Zero Crossing Rate (ZCR), etc. of the voice signal.

[0133] Step S120: Separate each to-be-processed voice signal to obtain third voice signals on two channels.

[0134] Step S130: Determine the third voice signal on the target channel in the two channels as the target voice signal of the target object.

[0135] In the embodiments of the present application, after the to-be-processed voice signals of the target object are collected, since the to-be-processed voice signals contain the human voice of the target object and various noises, the to-be-processed voice signals can be separated to separate the target voice signals of the relatively pure target object for a series of subsequent processes such as speech recognition. Optionally, a blind source separation algorithm can be used to separate the to-be-processed voice signals.

[0136] Among them, the target path is determined based on whether the voice signals of the two paths separated from the wake-up voice signal contain the target wake-up word. The wake-up voice signal includes the voice signals of the target object containing the target wake-up word collected by each voice collection device. Optionally, the wake-up voice signal can be the voice signal containing the target wake-up word collected in the current wake-up round, or the voice signal containing the target wake-up word collected in other wake-up rounds in history.

[0137] In the voice wake-up mode, when it is detected that the voice signals of the target object containing the target wake-up word are collected by each voice collection device, the voice signals containing the target wake-up word are determined as the wake-up voice signals. And perform blind source separation on the wake-up voice signals to obtain the fifth voice signals of two paths. Input the fifth voice signals of the two paths into the trained voice classification model to obtain the classification results of the fifth voice signals of each path. Among them, the classification result represents whether the fifth voice signal contains the target wake-up word. Based on the classification results of the fifth voice signals of the two paths, determine the path corresponding to the fifth voice signal with a relatively high probability of containing the target wake-up word as the target path. The same voice separation algorithm is used for separating the wake-up voice signals and separating the to-be-processed voice signals.

[0138] Among them, the voice classification model can be trained in the following way, and the training process is as Figure 3 shown.

[0139] Obtain multiple second sample voice signals with sample labels. Among them, the second sample voice signal is obtained by mixing the sample voice signal of the sample object corresponding to the second sample voice signal and the interference voice signal. The sample label of the second sample voice signal represents whether the second sample voice signal contains the target wake-up word. Optionally, the sample label of the second sample voice signal can be manually labeled, or labeled based on the matching result between the voice feature of the second sample voice signal and the wake-up word voice feature of the target wake-up word. If the voice feature of the second sample voice signal matches the wake-up word voice feature, it is considered that the second sample voice signal contains the target wake-up word, and the sample label can be labeled as 1; if not, it is considered that the second sample voice signal does not contain the target wake-up word, and the sample label can be labeled as 0.

[0140] Input each second sample voice signal into the second neural network model to be trained, and obtain the predicted classification result of each second sample voice signal through the second neural network model. Among them, the second neural network model can adopt a convolutional neural network model (Convolutional Neural Networks, CNN), a recurrent neural network model (Recurrent Neural Networks, RNN), a long short-term memory network model (Long Short Term Memory, LSTM), etc.

[0141] Determine the second training loss based on the predicted classification result of each second sample voice signal and the sample label of each second sample voice signal. Among them, the second training loss represents the difference between the predicted classification result obtained through the second neural network model and the true result of the sample label.

[0142] Adjust the model parameters in the second neural network model based on the second training loss.

[0143] Repeat the above training operations until the second training end condition is met. Take the second neural network model that meets the second training end condition as the trained voice classification model. Among them, the second training end condition can be that the second training loss is less than the third threshold, or the number of times of performing the training operation is greater than the fourth threshold. The third threshold and the fourth threshold can be set based on needs, and this application does not limit this.

[0144] Figure 4 This is a schematic structural diagram of voice separation based on blind source separation provided by an embodiment of this application, including two parts: path selection and target separation. Assume that there are two microphones configured in the first terminal: the first microphone and the second microphone, which are used to collect voice signals in the surrounding environment. When performing this voice processing method, first collect the voice signal Mic11 through the first microphone and the voice signal Mic21 through the second microphone, and input the collected voice signals Mic11 and Mic21 into the wake-up word detection module for wake-up word detection. If the target wake-up word is included in Mic11 and Mic21, it is determined that the wake-up is successful, and input Mic11 and Mic21 into the blind source separation module for blind source separation to obtain the voice signals of two paths. Input the voice signals of the two paths into the pre-trained voice classification model for classification, and based on the classification result, determine the path where the voice signal containing the target wake-up word is located as the target path.

[0145] After successful wake-up, continue to collect the voice signal Mic12 through the first microphone and the voice signal Mic22 through the second microphone, input the voice signals Mic12 and Mic22 into the blind source separation module for blind source separation, and perform blind source separation on Mic12 and Mic22 to obtain the voice signals of two channels. Determine the voice signal of the target channel from the voice signals of the two channels as the target voice signal of the target object.

[0146] It can be understood that when the target object is in different environments, there are also differences in the results of blind source separation of the voice signals to be processed collected in different environments. In the voice wake-up mode based on wake-up words, in order to more accurately separate and denoise the voice data to be processed collected in different environments, the target path corresponding to each wake-up round is respectively used to separate the voice signal to be processed in this wake-up round to obtain the target voice signal.

[0147] Specifically, for each wake-up round, perform blind source separation on the environmental voice signals that wake up this wake-up round to obtain the voice signals of two channels, and classify the voice signals of the two channels to determine the path where the voice signal containing the target wake-up word is located in the classification result as the target path corresponding to this wake-up round. Based on the target path corresponding to this wake-up round, determine the target voice signal of the voice signal to be processed in this wake-up round. Among them, the voice features of the environmental voice signals that wake up this wake-up round match the voice features of the wake-up word, that is, the environmental voice signals that wake up this wake-up round are voice signals containing the target wake-up word.

[0148] For example, assume that the target wake-up word is "Xiaowei". When the voice wake-up mode is turned on, continuously collect the environmental voice signals of the surrounding environment through each voice collection device, extract the features of the collected environmental voice signals, and match the voice features of the extracted environmental voice signals with the predefined target wake-up word. When the matching is successful, the wake-up is successful.

[0149] Perform blind source separation on the environmental voice signals containing the target wake-up word to obtain the voice signals of two channels, input the voice signals of the two channels into the trained voice classification model to obtain the classification result, and determine the path 1 where the voice signal containing "Xiaowei" is located in the classification result as the target path of the current wake-up round.

[0150] After the voice wake-up is successful, the target object's voice signals to be processed continue to be collected through various voice collection devices. Assuming that the collected voice signals to be processed are voice signals containing "How is the weather today", the voice signals to be processed are separated by a blind source separation algorithm, and the voice signal belonging to channel 1 of the two separated channels is determined as the target voice signal. The target voice signal can be subjected to voice recognition to obtain a voice recognition result of "How is the weather today?". The voice recognition result can be subsequently input into the natural language processing system for a corresponding reply.

[0151] If the target object's to-be-processed voice signal is not received within the preset time, the current wake-up round is deemed to be over. In the next wake-up round, the target path corresponding to the next wake-up round needs to be re-determined to separate the to-be-processed voice signal of the next wake-up round to obtain the target voice signal.

[0152] based on Figure 1 The voice processing method shown in the figure, in the voice wake-up mode, separates the collected voice signals to be processed to obtain the voice signal of the target path, and the voice separation effect is better, and the obtained target voice signal is more accurate. It can extract a purer target voice signal after denoising the target object, has stronger anti-interference ability, better noise reduction effect, and better meets the actual application needs.

[0153] In the embodiment of the present application, when in the first processing mode, speech separation can be performed by the following method to obtain the target speech signal.

[0154] Specifically, in response to the second target operation on the first terminal, the first processing mode is turned on, and each voice collection device configured on the first terminal is called to collect the voice signal to be processed of the target object in the surrounding environment. Among them, the second target operation can be a trigger operation of the target object on the screen display interface of the first terminal for the voice collection device, such as clicking, long pressing, etc., or it can be a specified gesture operation of the target object, such as recognizing the gesture of the target object in the air through the image sensor (such as a camera) configured on the first terminal. When the gesture of the target object is a predefined gesture instruction, the voice collection device is triggered to collect the voice signal to be processed.

[0155] After collecting each speech signal to be processed, the interfering speech signal and the first speech signal of the target object in each speech signal to be processed are separated by the trained speech separation model. The interfering speech signal is other speech signals in the speech signal to be processed except the first speech signal of the target object. According to the interfering speech signal in each speech signal to be processed and the first speech signal of the target object, the filtering weight coefficient is determined; based on the filtering weight coefficient, each speech signal to be processed is filtered to obtain the target speech signal.

[0156] Optionally, for the to-be-processed speech signal of the target object collected by each speech acquisition device, the to-be-processed speech signal can be input into a trained first speech separation model to separate the interfering speech signal in the to-be-processed speech signal and the first speech signal of the target object in the to-be-processed speech signal.

[0157] Optionally, for each to-be-processed speech signal, the to-be-processed speech signal can be input into a trained second speech separation model to separate the background noise signal and the human voice signal in the to-be-processed speech signal. Among them, the background noise signal refers to a non-human voice noise signal. Since the collected human voice signal is not pure, and the signal energy value of the first speech signal of the target object is usually higher than that of the interfering signal therein. Therefore, the first speech signal of the target object can be further screened based on the signal energy difference of the speech signal. The human voice signal in the to-be-processed speech signal is frame-divided to obtain multiple speech frames of the human voice signal. For each speech frame in the human voice signal, the signal energy value of the speech frame is determined. If the signal energy value of the speech frame is greater than a preset energy threshold, the speech frame is determined as a target speech frame, otherwise the target speech is a non-target speech frame. According to each target speech frame in the human voice signal, the first speech signal of the target object in the to-be-processed signal is determined. Among them, the first speech signal of the target object includes multiple target speech frames.

[0158] Among them, the detailed content of the method for performing speech separation using a speech separation model trained by an artificial intelligence method can be seen in steps S210 to S240 of the speech processing method shown later Figure 5 as shown.

[0159] The embodiment of the present application also provides a speech processing method, as Figure 5 shown, this method can be executed by a first terminal, and the first terminal is configured with at least two speech acquisition devices. This speech processing method may include the following steps S210 to S240.

[0160] Step S210: Obtain the to-be-processed speech signals of the target object respectively collected by each speech acquisition device.

[0161] Among them, the speech acquisition device may be a sound sensor configured on the first terminal for collecting speech, such as a microphone, and the target object is the target user who gives a speech instruction through the first terminal. The to-be-processed speech signal of the target object is a speech signal containing the human voice of the target object collected from the surrounding environment.

[0162] In the embodiment of the present application, this speech processing method can be implemented in the first processing mode or the speech wake-up mode based on a wake-up word. Among them, the first processing mode and the speech wake-up mode have been introduced in detail above, and the present application will not repeat them here. Please refer to the above content.

[0163] Step S220: Separate the interfering speech signals and the first speech signal of the target object from each speech signal to be processed through the trained speech separation model.

[0164] In the embodiment of the present application, the VAD algorithm can be used to preliminarily separate each speech signal to be processed, and the filtering weight coefficient is calculated based on the preliminary separation result. Specifically, for the speech signal to be processed of the target object collected by each voice acquisition device, the speech signal to be processed is input into the trained first speech separation model to separate the interfering speech signal in the speech signal to be processed and the first speech signal of the target object in the speech signal to be processed. Among them, the interfering speech signal is other speech signals in the speech signal to be processed except the first speech signal of the target object.

[0165] Optionally, for each speech signal to be processed, the speech signal to be processed is input into the trained second speech separation model to separate the background noise signal and the human voice signal in the speech signal to be processed. Among them, the background noise signal refers to a non-human voice noise signal. Since the collected human voice signal is not pure, and the signal energy value of the first speech signal of the target object is usually higher than that of the interfering signal therein. Therefore, the first speech signal of the target object can be further screened based on the signal energy difference of the speech signal. The human voice signal in the speech signal to be processed is frame-divided to obtain multiple speech frames of the human voice signal. Among them, the frame length is usually taken between 10 and 30 ms. The frame division can adopt the method of continuous segmentation or the method of overlapping segmentation. The frame shift of the overlapping part can be set as needed, such as set to 1 / 2 of the frame length. For each speech frame in the human voice signal, the signal energy value of the speech frame is determined. If the signal energy value of the speech frame is greater than the preset energy threshold, the speech frame is determined as the target speech frame, otherwise the target speech is the non-target speech frame. According to each target speech frame in the human voice signal, the first speech signal of the target object in the speech signal to be processed is determined. Among them, the first speech signal of the target object includes multiple target speech frames.

[0166] Step S230: Determine the filtering weight coefficient according to the interfering speech signal and the first speech signal of the target object in each speech signal to be processed.

[0167] After the preliminary speech separation is performed through the speech separation model, the filtering weight coefficient of the beamformer (filter) can be calculated based on the preliminary speech separation result, so as to perform filtering based on the beamforming algorithm to obtain a purer target speech signal of the target object.

[0168] Among them, the beamforming algorithm is used to merge the voice signals collected by the multi-channel voice acquisition device, suppress the voice signals in the non-target voice direction, enhance the voice signals in the target voice direction, achieve focused sound pickup in a specific direction, effectively improve the signal-to-noise ratio of the received signal, and achieve the effect of noise reduction. The beamforming algorithm in the embodiments of the present application can adopt the MVDR algorithm, the generalized sidelobe cancellation algorithm (Generalized Sidelobe Cancellation, GSC), etc.

[0169] According to the interfering voice signals in each voice signal to be processed, the noise covariance matrix is determined. According to the first voice signal of the target object in each voice signal to be processed, the target voice covariance matrix is determined, and based on the target voice covariance matrix, the direction-of-arrival vector of the target voice is determined. According to the noise covariance matrix and the direction-of-arrival (DOA) vector of the target voice, the filtering weight coefficient is determined. Among them, the noise covariance matrix characterizes the correlation between the interfering voice signals in each voice signal to be processed, the target voice covariance matrix characterizes the correlation between the first voice signals of the target object in each voice signal to be processed, and the direction-of-arrival vector of the target voice characterizes the direction of the signal source of the target voice. The direction-of-arrival vector of the target voice is determined by performing eigenvalue decomposition on the target voice covariance matrix and based on the eigenvector corresponding to the maximum eigenvalue.

[0170] Optionally, for each voice signal to be processed, the interfering voice signals in the voice signal to be processed may also include background noise signals and seventh voice signals. Among them, the seventh voice signal includes multiple non-target voice frames in the voice signal of the voice signal to be processed. The noise covariance matrix can be determined based on the interfering voice signals including the background noise signals and the seventh voice signals, and the filtering weight coefficient can be determined according to the noise covariance matrix and the direction-of-arrival vector of the target voice.

[0171] Optionally, when the MVDR algorithm is used for filtering, the filtering weight coefficient of the MVDR filter can be determined by the following formula:

[0172]

[0173] Among them, w represents the filtering weight coefficient of the MVDR filter, R n represents the noise covariance matrix of the interfering voice signals, R n -1 represents the inverse matrix of the noise covariance matrix, a(θ) represents the direction-of-arrival vector, that is, the DOA vector, a(θ) H represents the conjugate transpose vector of the direction-of-arrival vector.

[0174] The derivation process of the filtering weight coefficients of MVDR is as follows:

[0175] Assume that there is a uniform linear array of voice acquisition devices in space, which is used to receive voice signals in space. As Figure 6 shown, this linear array includes N array elements (voice acquisition devices x1, x2... xn), and the distance between adjacent array elements is d. When the voice signal S enters the array from space at an angle θ deviating from the normal, taking the leftmost first array element as a reference, the voice signal S needs to travel an extra distance of dsin(θ) to reach the next array element, and the corresponding delay generated is (c is the sound wave transmission speed in air).

[0176] Then the delay duration of each array element relative to the leftmost first array element is:

[0177]

[0178] The phase difference of each array element relative to the leftmost first array element can be expressed as:

[0179]

[0180] The voice signals received by the above array elements can be expressed as:

[0181]

[0182] It can be recorded as: x(n) = s(t)a(θ);

[0183] Among them,

[0184] Assume that there are complex voice signals corresponding to M sound sources in space, which enter the array at different angles respectively. Then the voice signals received by the array can be expressed as:

[0185]

[0186] Assume that the desired voice signal is S z in the direction of θ Z (t), and the interfering voice signals in other directions θ j are S j (t). Then the received voice signal can be expressed as:

[0187]

[0188] To enhance the desired voice signal S z in the direction of θ Z (t) and suppress the interfering signals S j in other directions θ j (t), that is, the output target is that x(t) output after filtering by the filter tends to SZ (t).

[0189] After filtering by the adaptive beamforming algorithm, the average power of the output signal can be expressed as:

[0190] P = E(|x(t)| 2 ) = w H E(x(t)*x(t) T )w = w H Rw

[0191] To remove the interfering speech signal and retain the desired speech signal, the output average power can be minimized to ensure that the interfering speech signal is minimized.

[0192] Convert it into a convex optimization problem with constraints:

[0193]

[0194] Solving this convex optimization problem based on the Lagrangian function gives the optimal weight coefficient

[0195]

[0196] Since the constraint condition is that the clean speech signal is distortionless, that is, the clean speech signal remains unchanged. To minimize the variance of the output, it is only necessary to minimize the interfering signal. Therefore, the output signal covariance matrix R in the above formula can be replaced by the noise covariance matrix Rn.

[0197] Thus, after replacement, the filtering weight coefficient of MVDR can be obtained

[0198]

[0199] Among them, the speech separation model is trained in the following way, and the training process is as Figure 7 shown.

[0200] Obtain multiple first training samples. Among them, each first training sample includes a first sample speech signal, the sample speech signal of the sample object corresponding to the first sample speech signal, and an interfering speech signal. The sample objects corresponding to each first sample speech signal can be different. The first sample speech signal is obtained by mixing the sample speech signal of the corresponding sample object and the interfering speech signal.

[0201] Input the first sample voice signals in each first training sample into the first neural network model to be trained, and obtain the predicted sample voice signals and predicted interference voice signals corresponding to each first sample voice signal through the first neural network model. Among them, the first neural network model can adopt a convolutional neural network model (Convolutional Neural Networks, CNN), a recurrent neural network model (Recurrent Neural Networks, RNN), a long short-term memory network model (Long Short Term Memory, LSTM), etc.

[0202] Determine the first training loss based on the predicted sample voice signals, predicted interference voice signals, sample voice signals, and interference voice signals corresponding to each first sample voice signal. Among them, the first training loss represents the minimum difference between the two groups of the predicted sample voice signal and the sample voice signal, and the predicted interference voice signal and the interference voice signal.

[0203] Adjust the model parameters in the first neural network model based on the first training loss.

[0204] Repeat the above training operations until the first training end condition is satisfied. Take the first neural network model that meets the first training end condition as the trained voice separation model. Among them, the first training end condition can be that the first training loss is less than the first threshold, or the number of times of performing the training operation is greater than the second threshold. The first threshold and the second threshold can be set based on needs, and this application does not limit this.

[0205] It should be noted that the first training samples used by the above first voice separation model and the second voice separation model are not exactly the same. The interference voice signals included in the first training samples used to train the first voice separation model include background noise signals and interference human voice signals of other objects. While the interference voice signals included in the first training samples used to train the first voice separation model only include background noise signals.

[0206] Optionally, the sample voice signals of the sample objects corresponding to the above first training samples are pure voice signals recorded by the sample objects in a quiet environment, and the interference voice signals can be noises corresponding to different signal-to-noise ratios and different noise types (noise environments). For example, record 100 pure voice signals, and mix them with noises of different signal-to-noise ratios of 10dB, 20dB, and 30dB and different noise environments (streets, office scenes, etc.) respectively to obtain the first sample voice signals.

[0207] Step S240: Filter each voice signal to be processed based on the filter weight coefficient to obtain the target voice signal.

[0208] In the embodiments of the present application, after obtaining the filtering weight coefficients of the MVDR filter through the above method, the target speech signal can be obtained by filtering each speech signal to be processed based on the following formula.

[0209] V t = w H * x

[0210] Wherein, V t represents a frame of target speech signal output after filtering, w H represents the conjugate transpose of the filtering weight coefficient, and x represents a frame of speech signal in the speech signal to be processed. For example, assume that x includes two speech signals x1 and x2 respectively corresponding to two speech collection devices. The conjugate transpose of the filtering weight coefficient is then w = [b 11 , b 12 . Thus, the target speech signal

[0211] Figure 8 is a schematic structural diagram of speech separation based on a speech separation model provided by the embodiments of the present application. Assume that two microphones are configured in the first terminal: a first microphone and a second microphone, which are used to collect speech signals in the surrounding environment. When executing this speech processing method, the first microphone signal Mic1 collected by the first microphone and the second microphone signal Mic2 collected by the second microphone can be input into a pre-trained speech separation model to separate the interfering speech signal of Mic1, the interfering speech signal of Mic2, the first speech signal of the target object in Mic1, and the first speech signal of the target object in Mic2. Then, the first microphone signal Mic1, the second microphone signal Mic2, and the separated interfering speech signal of Mic1, the interfering speech signal of Mic2, the first speech signal of the target object in Mic1, and the first speech signal of the target object in Mic2 are input into the MVDR filter. Based on the above method, the filtering weight coefficients of the MVDR filter are determined, and the first microphone signal Mic1 and the second microphone signal Mic2 are filtered based on the determined filtering weight coefficients to obtain the target speech signal of the target object.

[0212] Based on Figure 5 the speech processing method shown, the speech separation model trained based on the artificial intelligence method can be used to separate each speech signal to be processed, and the obtained interfering speech signal and the first speech signal of the target object are more accurate. The filtering weight coefficients determined based on the separation result are also more accurate. Therefore, the filtering effect based on the filtering weight coefficients is better, achieving a better denoising effect. The subsequent speech recognition result based on this speech separation result is more accurate, better meeting the actual application requirements.

[0213] The current voice separation solution determines the interfering voice signal and the first voice signal of the target object based on the energy value difference. If the signal energy value of the voice signal is greater than the fifth threshold, it is considered as the first voice signal of the target object; otherwise, it is considered as the interfering voice signal. Therefore, when performing DOA estimation, it is easily affected by point noise sources, resulting in the wrong sound source direction of the desired signal being located. However, the solution provided by the embodiments of the present application, based on the VAD detection algorithm, separates the interfering voice signal and the first voice signal of the target object through a voice separation model trained by artificial intelligence technology, and the separation result is more accurate. Moreover, subsequently, based on the separated first voice signal, DOA estimation is performed, shielding non-human voice signals and avoiding the sound source direction of non-human voice signals being located, achieving a better filtering effect.

[0214] It should be noted that the above two voice acquisition trigger modes and the above two voice separation methods can be combined arbitrarily. Among them, the two voice acquisition trigger modes include the first processing mode and the voice wake-up mode, and the two voice separation methods include: Figure 5 The voice separation method based on the voice separation model as shown; Figure 1 The voice separation method based on path selection as shown.

[0215] When combining the first processing mode and the voice separation method based on the voice separation model, in response to the second target operation, each voice acquisition device configured on the first terminal is called to acquire the to-be-processed voice signal of the target object in the surrounding environment. For the to-be-processed voice signal of the target object acquired by each voice acquisition device, the to-be-processed voice signal is input into the trained first voice separation model to separate the interfering voice signal in the to-be-processed voice signal and the first voice signal of the target object in the to-be-processed voice signal. According to the interfering voice signals in each to-be-processed voice signal, the noise covariance matrix Rn is determined. According to the first voice signals of the target object in each to-be-processed voice signal, the target human voice covariance matrix Rs is determined, and based on the target human voice covariance matrix Rs, the direction-of-arrival vector DOA of the target human voice is determined. According to the noise covariance matrix Rn and the direction-of-arrival vector DOA of the target human voice, the filtering weight coefficient w is determined. Based on the determined filtering weight coefficient w, each to-be-processed voice signal is filtered to obtain the target voice signal of the target object.

[0216] When combining the first processing mode with the voice separation method based on path selection, in response to the second target operation, each voice acquisition device configured on the first terminal is called to collect the voice signal to be processed of the target object in the surrounding environment. The target path can be determined based on the initial voice signal in each collected voice signal to be processed. For example, the voice signal in the initial 5s of each collected voice signal to be processed is used as the initial voice signal, and the voice signals of two paths obtained by performing blind source separation on each initial voice signal, and the voice signals of the two paths are classified, and the target path is determined based on the classification result.

[0217] Among them, when classifying the voice signals of the two paths, classification can be performed based on whether the target text is included, or classification can be performed based on whether the human voice of the target object is included, as long as the path containing the target voice signal of the target object can be distinguished. When classifying based on whether the target text is included, the target text can be the text included in the pre-agreed initial voice signal, such as "Please answer". When classifying based on whether the human voice of the target object is included, classification can be performed based on the voice classification model trained with the voice sample data of the target object in history for identifying the human voice of the target object.

[0218] After determining the target path based on the initial voice signal, continue to perform blind source separation on each voice signal to be processed to obtain the third voice signals of two paths, and determine the third voice signal of the target path in the two paths as the target voice signal.

[0219] When combining the voice wake-up mode with the voice separation method based on the voice separation model, after determining that the wake-up is successful, the voice signal to be processed collected after the wake-up is successful can be separated by the voice separation model to obtain the interference voice signal and the first voice signal of the target object, and based on the interference voice signal and the first voice signal of the target object of each voice signal to be processed, the filtering weight coefficient is determined, so as to filter each voice signal to be processed based on the filtering weight coefficient to obtain the target voice signal of the target object.

[0220] When combining the voice wake-up mode with the voice separation method based on path selection, after determining that the wake-up is successful, perform blind source separation on each voice signal containing the target wake-up word to obtain the fifth voice signals of two paths. The fifth voice signals of the two paths are input into the trained voice classification model to obtain the classification results of the fifth voice signals of each path. Based on the classification results of the fifth voice signals of the two paths, the path corresponding to the fifth voice signal with a higher probability of containing the target wake-up word is determined as the target path.

[0221] Continue to collect the voice signal to be processed after the wake-up is successful, perform blind source separation on each voice signal to be processed to obtain the third voice signals of two paths; determine the third voice signal of the target path in the two paths as the target voice signal.

[0222] The method provided in this application can be applied to the application scenario of voice interaction between a user and a first terminal. For example, the user asks, "What's the weather like today?" After the user triggers, the first terminal can collect the current voice signal through the voice collection device configured thereon, and perform noise reduction processing on the collected current voice signal through the voice processing method provided in the embodiments of this application to obtain the target voice signal of the user after noise reduction. Subsequently, voice recognition is performed based on the target voice signal, and a corresponding reply is made through the natural language processing system based on the voice recognition result.

[0223] Based on Figure 1 the same principle as the voice processing method shown, the embodiments of this application provide a voice processing device, as Figure 9 shown. The voice processing device 300 is deployed on the first terminal. The first terminal is configured with at least two voice collection devices. The voice processing device 300 may include an acquisition module 310, a separation module 320, and a target voice determination module 330.

[0224] The acquisition module 310 is configured to obtain the to-be-processed voice signals of the target object collected by each of the voice collection devices in the voice wake-up mode;

[0225] The separation module 320 is configured to separate each of the to-be-processed voice signals to obtain third voice signals of two channels;

[0226] The target voice determination module 330 is configured to determine the third voice signal of the target channel in the two channels as the target voice signal of the target object;

[0227] Wherein, the target channel is determined based on whether the voice signals of the two channels separated from the wake-up voice signal contain the target wake-up word. The wake-up voice signal includes the voice signals of the target object collected by each of the voice collection devices and containing the target wake-up word. The separation method of the wake-up voice signal is the same as the separation method of the to-be-processed voice signal.

[0228] Optionally, the voice wake-up mode is enabled in the following manner:

[0229] In response to a first target operation on the first terminal, the voice wake-up mode is enabled;

[0230] The acquisition module may be configured to:

[0231] In the voice wake-up mode, when the voice signals of the target object containing the target wake-up word are collected by each of the voice collection devices, it is determined that the voice wake-up is successful, and the to-be-processed voice signals of the target object are collected by each of the voice collection devices.

[0232] Optionally, the target path is determined as follows:

[0233] In the voice wake-up mode, the voice signal of the target object containing the target wake-up word collected by each of the voice collection devices is determined as the wake-up voice signal;

[0234] Separate the wake-up voice signal to obtain fifth voice signals of two paths;

[0235] Input the fifth voice signals of the two paths into the trained voice classification model to obtain the classification results of the fifth voice signals of each path, where the classification results indicate whether the fifth voice signal contains the target wake-up word;

[0236] Determine the path where the fifth voice signal containing the target wake-up word is located as the target path.

[0237] Optionally, the device further includes a filtering processing module, and the filtering processing module can be used for:

[0238] In the first processing mode, obtain the to-be-processed voice signals of the target object collected by each of the voice collection devices; wherein, the first processing mode is enabled in response to a second target operation on the first terminal;

[0239] Separate the interference voice signals and the first voice signals of the target object in each of the to-be-processed voice signals through the trained voice separation model, where the interference voice signals are other voice signals in the to-be-processed voice signals except for the first voice signals of the target object;

[0240] Determine the filtering weight coefficients according to the interference voice signals and the first voice signals of the target object in each of the to-be-processed voice signals;

[0241] Perform filtering processing on each of the to-be-processed voice signals based on the filtering weight coefficients to obtain the target voice signals.

[0242] Optionally, the filtering processing module can be used for:

[0243] Perform voice separation on each of the to-be-processed voice signals through the trained voice separation model to obtain the background noise signals and the human voice signals in each of the to-be-processed voice signals;

[0244] For each voice frame in the human voice signal of each to-be-processed voice signal, determine the signal energy value of the voice frame; and based on the signal energy values of the voice frames and a preset energy threshold, determine the first voice signal of the target object in the to-be-processed signal;

[0245] Among them, the interfering speech signal includes the background noise signal.

[0246] Optionally, the speech separation model is trained in the following manner:

[0247] Obtain a plurality of first training samples; each of the first training samples includes a first sample speech signal, a sample speech signal of a sample object corresponding to the first sample speech signal, and an interfering speech signal;

[0248] Perform a training operation on the first neural network model based on the plurality of first training samples until a first training end condition is satisfied, to obtain a trained speech separation model, and the training operation includes:

[0249] Input the first sample speech signals of the respective first training samples into the first neural network model, and obtain, through the first neural network model, predicted sample speech signals and predicted interfering speech signals corresponding to the respective first sample speech signals;

[0250] Determine a first training loss based on the predicted sample speech signals, predicted interfering speech signals, sample speech signals, and interfering speech signals corresponding to the respective first sample speech signals;

[0251] Adjust the model parameters in the first neural network model based on the first training loss.

[0252] Optionally, the speech classification model is trained in the following manner:

[0253] Obtain a plurality of second sample speech signals with sample labels; the sample labels of the second sample speech signals indicate whether the second sample speech signals contain a target wake-up word;

[0254] Perform a training operation on the second neural network model based on the plurality of second sample speech signals until a second training end condition is satisfied, to obtain a trained speech classification model, and the training operation includes:

[0255] Input the respective second sample speech signals into the second neural network model, and obtain, through the second neural network model, predicted classification results of the respective second sample speech signals;

[0256] Determine a second training loss based on the predicted classification results of the respective second sample speech signals and the sample labels of the respective second sample speech signals;

[0257] Adjust the model parameters in the second neural network model based on the second training loss.

[0258] Optionally, the filtering processing module can be used for:

[0259] Determining a filtering weight coefficient according to the interfering speech signals in each of the to-be-processed speech signals and the first speech signal of the target object includes:

[0260] Determining a noise covariance matrix according to the interfering speech signals in each of the to-be-processed speech signals;

[0261] Determining a target human voice covariance matrix according to the first speech signal of the target object in each of the to-be-processed speech signals; determining a direction-of-arrival vector of the target human voice based on the target human voice covariance matrix;

[0262] Determining a filtering weight coefficient according to the noise covariance matrix and the direction-of-arrival vector of the target human voice.

[0263] Based on Figure 9 In the voice wake-up mode, the voice processing device shown can separate the speech signals of the target path from the collected to-be-processed speech signals, with better voice separation effect, capable of extracting a purer target speech signal of the target object, stronger anti-interference ability, better noise reduction effect, and better meeting the actual application requirements.

[0264] Based on Figure 5 Based on the same principle as the voice processing method shown, an embodiment of the present application provides a voice processing device, as Figure 10 shown. The voice processing device 400 is deployed on a first terminal, and the first terminal is configured with at least two voice collection devices. The voice processing device 400 may include an acquisition module 410, a separation module 420, a filtering weight determination module 430, and a target voice determination module 440.

[0265] The acquisition module 410 is configured to acquire the to-be-processed speech signals of the target object collected by each of the voice collection devices;

[0266] The separation module 420 is configured to separate the interfering speech signals and the first speech signal of the target object in each of the to-be-processed speech signals through a trained voice separation model, where the interfering speech signals are other speech signals in the to-be-processed speech signals except the first speech signal of the target object;

[0267] The filtering weight determination module 430 is configured to determine a filtering weight coefficient according to the interfering speech signals and the first speech signal of the target object in each of the to-be-processed speech signals;

[0268] The target voice determination module 440 is configured to perform filtering processing on each of the to-be-processed speech signals based on the filtering weight coefficient to obtain the target voice signal.

[0269] Based on Figure 10The voice processing device shown can separate each voice signal to be processed based on a voice separation model trained by an artificial intelligence method, and the obtained interference voice signal and the first voice signal of the target object are more accurate. The filtering weight coefficients determined based on the separation result are also more accurate. Therefore, the filtering effect based on the filtering weight coefficients is better, achieving a better noise reduction effect. Subsequently, the result of voice recognition based on the voice separation result is more accurate, better meeting the actual application requirements.

[0270] The device according to the embodiment of the present application can execute the method provided by the embodiment of the present application, and its implementation principle is similar. The actions performed by each module in the device according to the embodiments of the present application correspond to the steps in the method according to the embodiments of the present application. For the detailed function description of each module of the device, reference can be specifically made to the description in the corresponding method shown above, and details are not described herein again.

[0271] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit including the function of the module or unit.

[0272] An electronic device is provided in the embodiments of the present application, including a memory, a processor, and a computer program stored on the memory. When the processor executes the computer program stored in the memory, the method in any optional embodiment of the present application can be implemented.

[0273] Figure 11 The structural schematic diagram of an electronic device applicable to the embodiments of the present invention is shown. As Figure 11 shown, the electronic device can be a server or a user terminal, and the electronic device can be used to implement the method provided in any embodiment of the present invention.

[0274] As Figure 11 shown in, the electronic device 2000 mainly includes at least one processor 2001 ( Figure 11 one is shown in), a memory 2002, a communication module 2003, and an input / output interface 2004 and other components. Optionally, the components can be connected and communicate through a bus 2005. It should be noted that Figure 11 the structure of the electronic device 2000 shown in is only schematic and does not constitute a limitation on the electronic device applicable to the method provided in the embodiments of the present application.

[0275] Among them, the memory 2002 can be used to store the operating system, application programs, etc. The application programs can include computer programs that implement the methods shown in the embodiments of the present invention when called by the processor 2001, and can also include programs for implementing other functions or services. The memory 2002 can be a ROM (Read Only Memory), or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory), or other types of dynamic storage devices that can store information and computer programs. It can also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.

[0276] The processor 2001 is connected to the memory 2002 through the bus 2005 and realizes corresponding functions by calling the application programs stored in the memory 2002. Among them, the processor 2001 can be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof, which can implement or execute various exemplary logic blocks, modules, and circuits described in connection with the disclosure of the present invention. The processor 2001 can also be a combination that realizes computing functions, such as a combination including one or more microprocessors, a combination of a DSP and a microprocessor, etc.

[0277] The electronic device 2000 can be connected to a network through the communication module 2003 (which can include but is not limited to components such as a network interface) to communicate with other devices (such as user terminals or servers, etc.) through the network, so as to realize data interaction, such as sending data to other devices or receiving data from other devices. Among them, the communication module 2003 can include a wired network interface and / or a wireless network interface, etc., that is, the communication module can include at least one of a wired communication module or a wireless communication module.

[0278] The electronic device 2000 can be connected to the required input / output devices through the input / output interface 2004, such as a keyboard, a display device, etc. The electronic device 200 itself can have a display device, and can also externally connect other display devices through the interface 2004. Optionally, a storage device, such as a hard disk, etc., can also be connected through the interface 2004, so as to store the data in the electronic device 2000 into the storage device, or read the data in the storage device, and can also store the data in the storage device into the memory 2002. It can be understood that the input / output interface 2004 can be a wired interface or a wireless interface. According to different actual application scenarios, the devices connected to the input / output interface 2004 can be components of the electronic device 2000 or external devices connected to the electronic device 2000 when needed.

[0279] The bus 2005 for connecting each component can include a path to transmit information between the above components. The bus 2005 can be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. According to different functions, the bus 2005 can be divided into an address bus, a data bus, a control bus, etc.

[0280] Optionally, for the solution provided by the embodiments of the present invention, the memory 2002 can be used to store a computer program for executing the solution of the present invention, and is run by the processor 2001. When the processor 2001 runs the computer program, it realizes the actions of the method or device provided by the embodiments of the present invention.

[0281] Based on the same principle as the method provided by the embodiments of the present application, the embodiments of the present application provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it can realize the corresponding content of the foregoing method embodiments.

[0282] The embodiment of the present application also provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the corresponding content of the foregoing method embodiment can be implemented.

[0283] It should be noted that the terms "first", "second", "third", "fourth", "1", "2", etc. (if any) in the description, claims and the above drawings of the present application are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than the illustrated or described order.

[0284] It should be understood that although the flowchart of the embodiment of the present application indicates each operation step by an arrow, the execution order of these steps is not limited to the order indicated by the arrow. Unless there is a clear description in this article, in some implementation scenarios of the embodiment of the present application, the implementation steps in each flowchart can be executed in other orders according to requirements. In addition, some or all of the steps in each flowchart may include multiple sub-steps or multiple stages based on the actual implementation scenario. Some or all of these sub-steps or stages can be executed at the same time, and each sub-step or stage among these sub-steps or stages can also be executed at different times. In the scenario where the execution times are different, the execution order of these sub-steps or stages can be flexibly configured according to requirements, and the embodiment of the present application does not limit this.

[0285] The above are only optional implementation manners of some implementation scenarios of the present application. It should be noted that for those of ordinary skill in the art, without departing from the technical concept of the solution of the present application, using other similar implementation means based on the technical idea of the present application also belongs to the protection scope of the embodiments of the present application.

Claims

1. A voice processing method, characterized in that, The method is executed by a first terminal, which is configured with at least two voice acquisition devices. The method includes: In the voice wake-up mode, obtain the to-be-processed voice signals of the target object respectively collected by each of the voice acquisition devices; Separate each of the to-be-processed voice signals to obtain third voice signals of two channels; Determine the third voice signal of the target channel in the two channels as the target voice signal of the target object; Wherein, the target channel is determined based on whether the voice signals of the two channels separated from the wake-up voice signal contain a target wake-up word. The wake-up voice signal includes the voice signals of the target object containing the target wake-up word collected by each of the voice acquisition devices, and the separation method of the wake-up voice signal is the same as that of the to-be-processed voice signal.

2. The method according to claim 1, characterized in that, The voice wake-up mode is enabled in the following manner: In response to a first target operation on the first terminal, enable the voice wake-up mode; The step of, in the voice wake-up mode, obtaining the to-be-processed voice signals of the target object respectively collected by each voice acquisition device includes: In the voice wake-up mode, when the voice signals of the target object containing the target wake-up word are collected by each of the voice acquisition devices, determine that the voice wake-up is successful, and respectively collect the to-be-processed voice signals of the target object by each of the voice acquisition devices.

3. The method according to claim 2, characterized in that, The target channel is determined in the following manner: In the voice wake-up mode, determine the voice signals of the target object containing the target wake-up word collected by each of the voice acquisition devices as the wake-up voice signal; Separate the wake-up voice signal to obtain fifth voice signals of two channels; Input the fifth voice signals of the two channels into a trained voice classification model to obtain the classification results of the fifth voice signals of each channel, where the classification results represent whether the fifth voice signals contain the target wake-up word; Determine the channel where the fifth voice signal containing the target wake-up word is located as the target channel.

4. The method according to claim 1, characterized in that, The method further includes: In a first processing mode, obtain the to-be-processed voice signals of the target object respectively collected by each of the voice acquisition devices; wherein, the first processing mode is enabled in response to a second target operation on the first terminal; Separate the interference voice signals and the first voice signals of the target object in each of the to-be-processed voice signals through a trained voice separation model, where the interference voice signals are other voice signals in the to-be-processed voice signals except for the first voice signals of the target object; Determine the filtering weight coefficients according to the interference voice signals and the first voice signals of the target object in each of the to-be-processed voice signals; Perform filtering processing on each of the to-be-processed voice signals based on the filtering weight coefficients to obtain the target voice signal.

5. The method according to claim 4, characterized in that, The step of separating the interference voice signals and the first voice signals of the target object in each of the to-be-processed voice signals through a trained voice separation model includes: Perform voice separation on each of the to-be-processed voice signals through the trained voice separation model to obtain the background noise signal and the human voice signal in each of the to-be-processed voice signals; For each voice frame in the human voice signal of each to-be-processed voice signal, determine the signal energy value of this voice frame; and based on the signal energy values of each voice frame and a preset energy threshold, determine the first voice signal of the target object in this to-be-processed signal; Among them, the interfering voice signal includes the background noise signal.

6. The method according to claim 4, characterized in that, The voice separation model is trained in the following manner: Obtain a plurality of first training samples; each of the first training samples includes a first sample voice signal, the sample voice signal of the sample object corresponding to the first sample voice signal, and an interfering voice signal; Perform a training operation on the first neural network model based on a plurality of first training samples until the first training end condition is satisfied to obtain a trained voice separation model. The training operation includes: Input the first sample voice signals of each first training sample into the first neural network model, and through the first neural network model, obtain the predicted sample voice signal and the predicted interfering voice signal corresponding to each first sample voice signal; Based on the predicted sample voice signal, the predicted interfering voice signal, the sample voice signal, and the interfering voice signal corresponding to each first sample voice signal, determine the first training loss; Adjust the model parameters in the first neural network model based on the first training loss.

7. The method according to claim 3, characterized in that, The voice classification model is trained in the following manner: Obtain a plurality of second sample voice signals with sample labels; the sample labels of the second sample voice signals indicate whether the second sample voice signals contain the target wake-up word; Perform a training operation on the second neural network model based on a plurality of second sample voice signals until the second training end condition is satisfied to obtain a trained voice classification model. The training operation includes: Input each second sample voice signal into the second neural network model, and through the second neural network model, obtain the predicted classification result of each second sample voice signal; Based on the predicted classification results of each second sample voice signal and the sample labels of each second sample voice signal, determine the second training loss; Adjust the model parameters in the second neural network model based on the second training loss.

8. The method according to claim 4, wherein The determining of the filtering weight coefficient according to the interfering voice signal in each of the to-be-processed voice signals and the first voice signal of the target object includes: Determine the noise covariance matrix according to the interfering voice signal in each of the to-be-processed voice signals; Determine the target human voice covariance matrix according to the first voice signal of the target object in each of the to-be-processed voice signals; based on the target human voice covariance matrix, determine the direction-of-arrival vector of the target human voice; Determine the filtering weight coefficient according to the noise covariance matrix and the direction-of-arrival vector of the target human voice.

9. A voice processing method, wherein The method is executed by a first terminal, and the first terminal is configured with at least two voice collection devices. The method includes: Obtain the to-be-processed voice signals of the target object collected by each of the voice collection devices respectively; Using the trained voice separation model, separate the interfering voice signals and the first voice signal of the target object from each of the to-be-processed voice signals, where the interfering voice signals are the other voice signals in the to-be-processed voice signals except for the first voice signal of the target object; Determine the filtering weight coefficients according to the interfering voice signals and the first voice signal of the target object in each of the to-be-processed voice signals; Perform filtering processing on each of the to-be-processed voice signals based on the filtering weight coefficients to obtain the target voice signals.

10. A voice processing device, wherein The device is deployed in a first terminal, and the first terminal is configured with at least two voice collection devices. The device includes: An acquisition module, configured to acquire the to-be-processed voice signals of the target object respectively collected by each of the voice collection devices in the voice wake-up mode; A separation module, configured to separate each of the to-be-processed voice signals to obtain third voice signals in two channels; A target voice determination module, configured to determine the third voice signal in the target channel among the two channels as the target voice signal of the target object; Wherein, the target channel is determined based on whether the voice signals in the two channels separated from the wake-up voice signal contain the target wake-up word. The wake-up voice signal includes the voice signals of the target object containing the target wake-up word collected by each of the voice collection devices, and the separation method of the wake-up voice signal is the same as that of the to-be-processed voice signal.

11. A voice processing device, wherein The device is deployed in a first terminal, and the first terminal is configured with at least two voice collection devices. The device includes: An acquisition module, configured to acquire the to-be-processed voice signals of the target object respectively collected by each of the voice collection devices; A separation module, configured to use the trained voice separation model to separate the interfering voice signals and the first voice signal of the target object from each of the to-be-processed voice signals, where the interfering voice signals are the other voice signals in the to-be-processed voice signals except for the first voice signal of the target object; A filtering weight determination module, configured to determine the filtering weight coefficients according to the interfering voice signals and the first voice signal of the target object in each of the to-be-processed voice signals; A target voice determination module, configured to perform filtering processing on each of the to-be-processed voice signals based on the filtering weight coefficients to obtain the target voice signals.

12. An electronic device, wherein The electronic device includes a memory and a processor. A computer program is stored in the memory, and the processor executes the computer program to implement the method according to any one of claims 1 to 8 or claim 9.

13. A computer-readable storage medium, wherein A computer program is stored in the storage medium, and when the computer program is executed by a processor, the method according to any one of claims 1 to 8 or claim 9 is implemented.