Data processing method, voice data processing method, and electronic device
By preprocessing pure sample data with different types of interference sample data, a second type of interference processing model is trained, which solves the problem of interference processing changing the characteristics of speech data in existing technologies and improves the quality of data transmission.
Patent Information
- Application Number
- CN202210293597.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-23
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2042-03-23
AI Technical Summary
In existing technologies, during voice data transmission, processing different interference data sequentially can alter the characteristics of the voice data or other interference data, affecting the extraction of clean voice data and reducing transmission quality.
By preprocessing the clean sample data with different types of interference sample data, a second type of interference processing model is trained to identify and eliminate the second type of interference data, thereby improving the data processing quality.
It enables effective identification and elimination of data after processing for the first type of interference, thereby improving the overall quality of data processing.
Smart Images

Figure CN114694669B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present application relate to the field of communication, and in particular, to a data processing method, a voice data processing method and an electronic device. BACKGROUND
[0002] In a communication scenario, when performing transmission of voice data and the like, a plurality of interference data such as echo, noise and the like will usually be accompanied, thereby affecting the transmission quality of the voice data. Therefore, in order to improve the transmission quality of the voice data, the interference data will be processed, such as echo cancellation, voice enhancement and the like, so as to extract relatively pure voice data.
[0003] In a conventional manner, when processing a plurality of interference data, different interference data is usually processed in sequence. However, when processing a certain type of interference data, the voice data or other interference data may be affected, thereby changing the characteristics of the voice data or other interference data, ultimately affecting the extraction of pure voice data and the transmission quality of the voice data. SUMMARY
[0004] Embodiments of the present application provide a data processing method, a voice data processing method and an electronic device, to solve the problem of poor data processing quality in the prior art.
[0005] In a first aspect, a data processing method is provided in embodiments of the present application, comprising:
[0006] mixing the pure sample data with first-class first interference sample data according to a first processing probability, and then performing first-class interference processing, to obtain first processing data;
[0007] mixing second-class interference sample data with first-class second interference sample data according to a second processing probability, and then performing first-class interference processing, to obtain second processing data;
[0008] mixing the pure sample data or the first processing data with the second-class interference sample data, or mixing the pure sample data or the first processing data with the second processing data, to obtain target sample data;
[0009] training a second-class interference processing model using the target sample data and the pure sample data; the second-class interference processing model is used to perform second-class interference processing on the to-be-processed data after first-class interference processing, to eliminate second-class interference data in the to-be-processed data.
[0010] In a second aspect, a voice data processing method is provided in embodiments of the present application, comprising:
[0011] The clean speech sample data is mixed with the first echo sample data according to a first processing probability, and then is subjected to echo cancellation processing to obtain first processing data.
[0012] The noise sample data is mixed with the second echo sample data according to a second processing probability, and then is subjected to echo cancellation processing to obtain second processing data.
[0013] The clean speech sample data or the first processing data is mixed with the noise sample data, or the clean speech sample data or the first processing data is mixed with the second processing data to obtain target sample data.
[0014] The noise processing model is trained by using the target sample data and the clean speech sample data, and the second type of interference processing model is used to perform speech enhancement processing on the to-be-processed speech data after echo cancellation processing to eliminate noise data in the to-be-processed speech data.
[0015] In a third aspect, a data processing method is provided in the embodiments of the present application, which comprises:
[0016] The clean sample data is mixed with the first type of first interference sample data, and then is subjected to first type of interference processing to obtain first processing data.
[0017] The second type of interference sample data is mixed with the first type of second interference sample data, and then is subjected to first type of interference processing to obtain second processing data.
[0018] A corresponding number of clean sample data and first processing data are selected and mixed with a corresponding number of second type of interference sample data and second processing data respectively to obtain a plurality of target sample data.
[0019] The second type of interference processing model is trained by using the plurality of target sample data and the respective corresponding clean sample data, and the second type of interference processing model is used to perform second type of interference processing on the to-be-processed data after first type of interference processing to eliminate second type of interference data in the to-be-processed data.
[0020] In a fourth aspect, a speech data processing method is provided in the embodiments of the present application, which comprises:
[0021] The clean speech sample data is mixed with the first echo sample data, and then is subjected to echo cancellation processing to obtain first processing data.
[0022] The noise sample data is mixed with the second echo sample data, and then is subjected to echo cancellation processing to obtain second processing data.
[0023] selecting a corresponding number of pure speech sample data and first processing data, respectively mixing a corresponding number of noise sample data and second processing data to obtain a plurality of target sample data;
[0024] training a noise processing model using the plurality of target sample data and the respective corresponding pure speech sample data; the noise processing model is used to perform speech enhancement processing on the to-be-processed speech data after echo cancellation processing to eliminate noise data in the to-be-processed speech data.
[0025] In a fifth aspect, an embodiment of the present application provides a data processing method, comprising:
[0026] mixing the pure sample data with the first type of interference sample data according to a third processing probability and performing first type of interference processing to obtain third processing data;
[0027] mixing the pure sample data or the third processing data with the second type of interference sample data to obtain target sample data;
[0028] training a second type of interference processing model using the target sample data and the pure sample data; the second type of interference processing model is used to perform second type of interference processing on the to-be-processed data after first type of interference processing to eliminate second type of interference data in the to-be-processed data.
[0029] In a sixth aspect, an embodiment of the present application provides a data processing method, comprising:
[0030] mixing the second type of interference sample data with the first type of interference sample data according to a fourth processing probability and performing first type of interference processing to obtain fourth processing data;
[0031] mixing the pure sample data with the second type of interference sample data or the fourth processing data to obtain target sample data;
[0032] training a second type of interference processing model using the target sample data and the pure sample data; the second type of interference processing model is used to perform second type of interference processing on the to-be-processed data after first type of interference processing to eliminate second type of interference data in the to-be-processed data.
[0033] In a seventh aspect, an embodiment of the present application provides a data processing method, comprising:
[0034] performing first type of interference processing on the to-be-processed data to obtain preprocessed data;
[0035] performing second type of interference processing on the preprocessed data using the second type of interference processing model to eliminate second type of interference data in the preprocessed data to obtain target data.
[0036] The second type of interference processing model is trained by using target sample data and pure sample data; the target sample data is obtained by mixing the pure sample data or first processing data with second type of interference sample data, or by mixing the pure sample data or the first processing data with second processing data; the first processing data is obtained by mixing the pure sample data with first type of first interference sample data and then performing first type of interference processing; and the second processing data is obtained by mixing the second type of interference sample data with first type of second interference sample data and then performing first type of interference processing.
[0037] In an eighth aspect, a speech data processing method is provided in the embodiments of the present application, and the method comprises:
[0038] Obtaining user speech data;
[0039] Performing echo cancellation processing on the user speech data to obtain preprocessed speech data;
[0040] Performing speech enhancement processing on the preprocessed speech data by using a noise processing model to eliminate noise data in the preprocessed speech data and obtain target speech data; the noise processing model is trained by using target sample data and pure speech sample data; the target sample data is obtained by mixing the pure speech sample data or first processing data with noise sample data, or by mixing the pure speech sample data or the first processing data with second processing data; the first processing data is obtained by mixing the pure speech sample data with first echo sample data and then performing echo cancellation processing; and the second processing data is obtained by mixing the noise sample data with second echo sample data and then performing echo cancellation processing.
[0041] Sending the target speech data.
[0042] In a ninth aspect, an electronic device is provided in the embodiments of the present application, which comprises a storage component and a processing component. The storage component stores one or more computer program instructions. The one or more computer program instructions are called and executed by the processing component to implement the data processing method in any one of the first aspect, the third aspect, the fifth aspect, the sixth aspect, or the seventh aspect, or the speech data processing method in any one of the second aspect, the fourth aspect, or the eighth aspect.
[0043] In the embodiments of the present application, the first type of interference processing is performed on the pure sample data and the second type of interference sample data according to the processing probability, the processed data is also used as sample data to participate in training, target sample data is obtained, and the second type of interference processing model is trained by using the target sample data, so that the second type of interference processing model can identify the data processed by the first type of interference processing, and the second type of interference processing is performed on the data to eliminate the second type of interference data. The problem that the first type of interference processing affects the pure data and the second type of interference data in the prior art, and affects the effect of the second type of interference processing is solved, and the data processing quality is improved.
[0044] These aspects or other aspects of the present application will be more apparent in the following description of the embodiments. BRIEF DESCRIPTION OF DRAWINGS
[0045] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings required to be used in the description of the embodiments or the prior art will be briefly introduced. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.
[0046] Figure 1a A system architecture schematic diagram in which the technical solutions of the present application are applied is shown;
[0047] Figure 1b A flowchart of one embodiment of a data processing method provided by the present application is shown;
[0048] Figure 2 A flowchart of another embodiment of a data processing method provided by the present application is shown;
[0049] Figure 3 A flowchart of another embodiment of a data processing method provided by the present application is shown;
[0050] Figure 4 A flowchart of another embodiment of a data processing method provided by the present application is shown;
[0051] Figure 5 A flowchart of one embodiment of a voice data processing method provided by the present application is shown;
[0052] Figure 6 A structural schematic diagram of one embodiment of a noise processing model provided by the present application is shown;
[0053] Figure 7 A flowchart of another embodiment of a voice data processing method provided by the present application is shown;
[0054] Figure 8 a flow chart illustrating another embodiment of a data processing method provided by the present application is shown;
[0055] Figure 9 a flow chart illustrating another embodiment of a voice data processing method provided by the present application is shown;
[0056] Figure 10 a structural schematic diagram illustrating one embodiment of a system architecture diagram provided by the present application is shown;
[0057] Figure 11 a structural schematic diagram illustrating one embodiment of a data processing apparatus provided by the present application is shown;
[0058] Figure 12 a structural schematic diagram illustrating one embodiment of a voice data processing apparatus provided by the present application is shown;
[0059] Figure 13 a structural schematic diagram illustrating one embodiment of an electronic device provided by the present application is shown. DETAILED DESCRIPTION
[0060] In order to enable persons skilled in the art to better understand the scheme of the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application.
[0061] In some of the descriptions in the specification and claims of the present application and the above-mentioned drawings, a plurality of operations appearing in a specific order are included, but it should be clearly understood that these operations can be executed or in parallel without the order in which they appear in this text, and the serial numbers of the operations, such as 101, 102, etc., are only used to distinguish different operations, and the serial numbers themselves do not represent any execution order. In addition, these processes can include more or fewer operations, and the operations can be executed in sequence or in parallel. It should be noted that the "first", "second", etc. in this text are used to distinguish different messages, devices, modules, etc., and do not represent the order of precedence. Also, "first" and "second" are not of different types.
[0062] In actual application, when data transmission is performed in a communication scenario, such as voice data transmission, it is usually accompanied by various interference signals, such as echo, noise, etc., thereby affecting the transmission quality of voice data. Therefore, in order to improve the transmission quality of voice data, interference data is processed, such as echo cancellation, speech enhancement, etc., to extract relatively pure voice data for transmission, so as to improve the communication quality.
[0063] In the conventional manner, when processing multiple interference data, different interference data is usually processed in sequence. However, when processing a certain type of interference data, it may affect the voice data or other interference data, thereby changing the characteristics of the voice data or other interference data, ultimately affecting the extraction of pure voice data and the transmission quality of voice data. For example, when performing echo cancellation and noise cancellation on voice mixed with echo and noise, the echo cancellation process may distort the characteristics of the voice data, damaging the voice data, resulting in the damaged voice data being identified as noise data and being eliminated during subsequent noise cancellation, causing voice data loss, etc. Or, the echo cancellation process may distort the form of the noise, damaging the noise, making it difficult to identify the damaged noise during subsequent noise cancellation, resulting in incomplete noise cancellation, etc.
[0064] To solve the above technical problems, the inventors found that during subsequent interference data processing, it is usually impossible to identify the damaged data after the previous interference processing, resulting in poor data processing quality. Therefore, after a series of thinking and experiments, the technical solution of the present application is proposed, which provides a data processing method, including mixing pure sample data with first type first interference sample data according to a first processing probability and performing first type interference processing to obtain first processing data; mixing second type interference sample data with first type second interference sample data according to a second processing probability and performing first type interference processing to obtain second processing data; mixing the pure sample data or the first processing data with the second type interference sample data, or mixing the pure sample data or the first processing data with the second processing data to obtain target sample data; training a second type interference processing model using the target sample data and the pure sample data; and using the second type interference processing model to perform second type interference processing on the data to be processed after first type interference processing to eliminate second type interference data in the data to be processed.
[0065] In the embodiments of the present application, the pure sample data and the second type interference sample data are pre-processed according to the processing probability, and the processed data is also used as sample data for training to obtain target sample data. The second type interference processing model is trained using the target sample data, which can identify the data after first type interference processing and perform second type interference processing to eliminate second type interference data. This solves the problem in the prior art that first type interference processing affects pure data and second type interference data, affecting the effectiveness of second type interference processing, and improves the data processing quality.
[0066] With reference to the drawings of the embodiments of the present application, the technical solutions in the embodiments of the present application will be described clearly and completely. Obviously, the described embodiments are only a part of the embodiments of the present application, but not all the embodiments of the present application. Based on the embodiments of the present application, all the other embodiments obtained by those skilled in the art without creative work fall within the scope of the present application.
[0067] The technical solutions of the present application are applied to a communication scenario, and in one actual application, are particularly applicable to a communication scenario implemented based on RTC (Real-Time Communication) technology.
[0068] RTC technology refers to a communication technology capable of transmitting and receiving text, audio and video in real time, and is applicable to scenarios such as live broadcast, on-demand, video conference, online classroom, online chat room and game interaction, to realize real-time transmission of pure audio data, video data and the like. The technical solutions of the present application can be specifically applied to communication scenarios such as live broadcast, on-demand, video conference, online classroom, online chat room and game interaction implemented based on RTC.
[0069] Referring to Figure 1a , a system architecture diagram in which the technical solutions of the embodiments of the present application can be applied is shown. The system can include a server 100 and a plurality of clients 200. The plurality of clients 200 can establish a communication connection through the server 100, and in an RTC scenario, the server 100 is used to provide RTC services between the plurality of clients 200. The plurality of clients 200 can respectively act as a sending end or a receiving end, and realize real-time communication through the server 100. It should be noted that the number of clients shown in the figure is only illustrative, and the present application is not limited thereto.
[0070] A user can interact with the server 100 through the client 200 to receive data sent by other clients 200, or send data to other clients 200, and the like. In an RTC scenario, a user can publish a data stream to the server 100 through the client 200, and the client 200 pushes the data stream to a client subscribing to the data stream. The data stream can be, for example, audio stream, video stream and the like. For example, in a live broadcast scenario, a host user can collect media data in real time through a client and send it to a server. The media data of different host users is distinguished by a live broadcast room, and the server can push the media data of the host user to the client of a watching user entering the live broadcast room corresponding to the host user. For another example, in a conference scenario, a participant user can collect media data in real time through a client and send it to a server, and the server can push the media data sent by each client to the client of another participant user, and the like.
[0071] It can be understood that the data transmitted by the client 200 can need to be encoded, transcoded, compressed, and the like before being published to the server 100, and can also be subjected to interference processing and the like according to the technical solutions provided in the embodiments of the present application. Details will be described below.
[0072] Among them, the client 200 and the server 100 establish a connection through a network. The network provides a medium for a communication link between the client and the server. The network can include various connection types, such as wired, wireless communication links, or fiber optic cables, and the like.
[0073] Among them, the client 200 can be a browser, an APP (Application, application), or a web application such as an H5 (HyperText Markup Language 5, HyperText Markup Language version 5) application, or a light application (also known as a small program, a lightweight application), or a cloud application, etc. The client 200 can be based on the SDK (Software Development Kit, software development kit) provided by the server for the corresponding service, such as based on the RTC SDK development obtained, etc. The client 200 can be deployed in an electronic device, and needs to rely on the device to run or run in some app in the device, etc. The electronic device may, for example, have a display screen and support information browsing, etc., and can be a personal mobile terminal such as a mobile phone, a tablet computer, a personal computer, a smart wearable device, etc. Various other types of applications can also be configured in the electronic device, such as human-computer dialogue type applications, model training type applications, text processing type applications, web browser applications, shopping type applications, search type applications, instant messaging tools, email clients, social platform software, etc.
[0074] The server 100 can include servers that provide various services, such as servers that provide communication services for multiple clients, servers that provide support for models used on clients for background training, servers that process data sent by clients, and the like.
[0075] It should be noted that the server 100 can be implemented as a distributed server cluster composed of multiple servers, or as a single server. The server can be a server of a distributed system, or a server combined with a blockchain. The server can also be a cloud server or an intelligent cloud host that provides cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content distribution networks (Content Delivery Network, CDN), and big data and artificial intelligence platforms, and the like. Basic cloud computing services.
[0076] It should be noted that the data processing method and the voice data processing method provided in the embodiments of the present application are generally executed by a server, and the corresponding data processing apparatus and voice data processing apparatus are generally arranged in the server. However, in other embodiments of the present application, a client can also have similar functions as the server, so as to execute the data processing method and the voice data processing method provided in the embodiments of the present application. In other embodiments, the data processing method and the voice data processing method provided in the embodiments of the present application can also be executed by the client and the server together.
[0077] Figure 1b A flowchart of one embodiment of the data processing method provided in the embodiments of the present application can include the following steps:
[0078] 101: After mixing the pure sample data according to the processing probability and the first type of interference sample data, the first type of interference processing is performed on the mixed data to obtain processed data.
[0079] 102: The pure sample data or the processed data is mixed with the second type of interference sample data to obtain target sample data.
[0080] 103: The second type of interference processing model is trained by using the target sample data and the pure sample data.
[0081] The second type of interference processing model is used to perform the second type of interference processing on the data to be processed after the first type of interference processing, so as to eliminate the second type of interference data in the data to be processed.
[0082] The pure data can refer to single data without interference data, such as voice data, image data, etc. The interference data can refer to data that interferes with the pure data. For example, if the pure data is voice data, the interference data can be noise data, echo data, etc. In actual application, the pure data is usually mixed with multiple types of interference data, and the pure data is obtained by sequentially processing different interference data.
[0083] In the embodiments, two types of interference data are taken as examples for description. For ease of description, the first type of interference data and the second type of interference data can be referred to respectively. For the mixed data in which the pure data is mixed with the first type of interference data and the second type of interference data, the first type of interference processing can be performed on the first type of interference data first, and then the second type of interference processing can be performed on the second type of interference data, so as to obtain the pure data. For example, the first type of interference data is echo data, and the second type of interference data is noise data. Correspondingly, the first type of interference data can be echo cancellation processing, and the second type of interference processing can be noise cancellation processing, such as voice enhancement processing, without limitation.
[0084] The second type of interference processing can be performed by using a second type of interference processing model. The second type of interference processing model can be a neural network model, such as a convolutional neural network (CNN) model, a fully convolutional network (FCN) model, and the like, which can be set according to an actual application scenario. The second type of interference processing model can perform second type of interference processing on the processed data after the first type of interference processing to eliminate the second type of interference data in the processed data.
[0085] When the first type of interference processing is performed, the first type of interference processing can affect the pure data to some extent, thereby affecting the subsequent second type of interference processing. For example, when the pure data is speech data, the first type of interference data is echo data, and the second type of interference data is noise data, the echo cancellation processing can distort the characteristics of the speech data, damage the speech data, and cause the damaged speech data to be recognized as noise data and eliminated during subsequent noise cancellation, resulting in missing speech data. Therefore, when training the second type of interference processing model using sample data, the sample data damaged by the first type of interference processing can also be used as training samples to participate in the training of the second type of interference processing model.
[0086] Specifically, for the pure sample data, the pure sample data can be mixed with the first type of interference sample data according to a processing probability and then subjected to the first type of interference processing to obtain processed data, which can be damaged pure sample data. The processing probability can be set in advance, such as 5% to 10%, and can be set according to an actual application scenario. The pure sample data or the processed data can be mixed with the second type of interference sample data to obtain target sample data. When the pure sample data is mixed with the first type of interference sample data and subjected to the first type of interference processing, the processed data can be mixed with the second type of interference sample data to obtain the target sample data. When the pure sample data is not mixed with the first type of interference sample data and subjected to the first type of interference processing, the pure sample data can be mixed with the second type of interference sample data to obtain the target sample data. Then, the target sample data can be used as training samples, and the pure sample data can be used as training labels to train the second type of interference processing model. The specific model training process can refer to the implementation manner in a conventional scheme, and will not be described herein.
[0087] Optionally, the pure sample data and the second type of interference sample data can be obtained from the training data set, and the first type of interference sample data can be pre-set. Taking the first type of interference sample data as echo sample data as an example, the echo sample data can be pre-set echo simulation sample, and the application does not limit the setting of the echo simulation sample.
[0088] In the embodiment, the pure sample data is pre-processed by the first type of interference according to the processing probability, the processed data is also used as sample data for training to obtain target sample data, and the second type of interference processing model is trained by the target sample data, so that the second type of interference processing model can recognize the data processed by the first type of interference and perform the second type of interference processing to eliminate the second type of interference data. The problem that the first type of interference processing affects the pure data and affects the effect of the second type of interference processing in the prior art is solved, and the data processing quality is improved.
[0089] In actual application, when the first type of interference processing is performed, the first type of interference processing can also affect the second type of interference data to some extent, thereby affecting the subsequent second type of interference processing. Taking the pure data as voice data, the first type of interference data as echo data, and the second type of interference data as noise data as an example, the echo cancellation processing can distort the form of the noise data, damage the noise data, and cause the noise cancellation to be incomplete when the subsequent noise cancellation is performed. Figure 2 A flowchart of another embodiment of a data processing method provided by the application can include the following steps:
[0090] 201: The second type of interference sample data is mixed with the first type of interference sample data according to the processing probability, and then the first type of interference processing is performed to obtain processed data.
[0091] 202: The pure sample data is mixed with the second type of interference sample data or the processed data to obtain target sample data.
[0092] 203: The second type of interference processing model is trained by using the target sample data and the pure sample data.
[0093] The second type of interference processing model is used to perform the second type of interference processing on the to-be-processed data processed by the first type of interference to eliminate the second type of interference data in the to-be-processed data.
[0094] In this embodiment, for the second type of interference sample data, the second type of interference sample data can be mixed with the first type of interference sample data according to the processing probability, and then the first type of interference processing is performed to obtain processed data, which can be the second type of interference sample data after damage. For ease of description, the processing probability in the embodiment shown in FIG. 1 can be referred to as the first processing probability, the obtained processed data can be referred to as the first processed data, and the processing probability in this embodiment can be referred to as the second processing probability, and the obtained processed data can be referred to as the second processed data. The second processing probability and the first processing probability are independent of each other, and can be set according to actual application scenarios, such as setting the second processing probability to 2% to 4%.
[0095] Then, the pure sample data can be mixed with the second type of interference sample data or the second processed data to obtain target sample data. When the second type of interference sample data is mixed with the first type of interference sample data and the first type of interference processing is performed, the pure sample data can be mixed with the second processed data to obtain the target sample data. When the second type of interference sample data is not mixed with the first type of interference sample data and the first type of interference processing is not performed, the pure sample data can be mixed with the second type of interference sample data to obtain the target sample data. The target sample data is used as a training sample, and the pure sample data is used as a training label to train the second type of interference processing model, which will not be described again.
[0096] In this embodiment, by pre-processing the second type of interference sample data according to the processing probability, the processed data is also used as sample data for training to obtain target sample data. Training the second type of interference processing model with the target sample data can enable the second type of interference processing model to recognize the data after the first type of interference processing and perform the second type of interference processing to eliminate the second type of interference data. This solves the problem in the prior art that the first type of interference processing affects the second type of interference data, resulting in an impact on the second type of interference processing effect, and improves the data processing quality.
[0097] In actual applications, when the first type of interference processing is performed, the first type of interference processing can simultaneously affect the pure sample data and the second type of interference data, thereby affecting the subsequent second type of interference processing. Figure 3 A flowchart of another embodiment of a data processing method provided by the embodiments of the present application can include the following steps:
[0098] 301: The pure sample data is mixed with the first type of first interference sample data according to the first processing probability, and then the first type of interference processing is performed to obtain the first processed data.
[0099] 302: The second type of interference sample data is mixed with the first type of second interference sample data according to a second processing probability, and then first type of interference processing is performed to obtain second processing data.
[0100] In this embodiment, the pure sample data and the second type of interference sample data can be mixed with the first type of interference sample data and subjected to the first type of interference processing. For ease of description, the first type of interference sample data mixed with the pure sample data can be referred to as first type of first interference sample data, and the first type of interference sample data mixed with the second type of interference sample data can be referred to as first type of second interference sample data. Optionally, the first type of first interference sample data and the first type of second interference sample data can be the same or different.
[0101] 303: The pure sample data or the first processing data is mixed with the second type of interference sample data, or the pure sample data or the first processing data is mixed with the second processing data to obtain target sample data.
[0102] The pure sample data can be mixed with the second type of interference sample data to obtain the target sample data, the first processing data can be mixed with the second type of interference sample data to obtain the target sample data, the pure sample data can be mixed with the second processing data to obtain the target sample data, and the first processing data can be mixed with the second processing data to obtain the target sample data. Four different target sample data can be obtained.
[0103] 304: The second type of interference processing model is trained by using the target sample data and the pure sample data.
[0104] The second type of interference processing model can be used to perform second type of interference processing on the to-be-processed data after the first type of interference processing to eliminate the second type of interference data in the to-be-processed data.
[0105] Training the second type of interference processing model by using the four different target sample data can enable the second type of interference processing model to extract pure data from mixed data containing pure sample data and second type of interference sample data, from mixed data containing first processing data and second type of interference sample data, from mixed data containing pure sample data and second processing data, and from mixed data containing first processing data and second processing data.
[0106] In this embodiment, the pure sample data and the second type of interference sample data are pre-processed according to the processing probability, the processed data is also used as sample data for training, and the target sample data is obtained. The second type of interference processing model is trained by using the target sample data, so that the second type of interference processing model can recognize the data processed by the first type of interference processing, and the second type of interference processing is performed to eliminate the second type of interference data. The problem that the first type of interference processing affects the pure data and the second type of interference data in the prior art, resulting in the problem of affecting the second type of interference processing effect is solved, and the data processing quality is improved.
[0107] Optionally, when the second type of interference sample data is mixed with the first type of second interference sample data according to the second processing probability and then subjected to the first type of interference processing, the second type of interference sample data can be mixed with the first type of second interference sample data and the pure sample data according to the second processing probability and then subjected to the first type of interference processing to obtain second processing data. For example, when the pure sample data is pure speech sample data, the first type of interference sample data is echo sample data, and the second type of interference sample data is noise sample data, the noise sample data can be mixed with the pure speech sample data and the echo sample data according to the second processing probability and then subjected to echo cancellation processing to obtain second processing data. The second processing data can include speech data with damaged noise data.
[0108] In practical applications, the data processing method can further include:
[0109] The pure sample data and the second type of interference sample data are randomly selected from the training data set.
[0110] The training data set can include multiple types of pure sample data and second type of interference sample data. For example, when the pure sample data is pure speech sample data and the second type of interference sample data is noise sample data, the training data set can include different types of pure speech sample data, such as male speech sample data, female speech sample data, old speech sample data, and child speech sample data. Different types of noise sample data, such as air conditioner noise sample data and office noise sample data, can be included. The sample data in the training data set can be pre-collected. A pure sample data and a second type of interference sample data are randomly selected from the training data set for the above processing.
[0111] In some embodiments, the method of mixing the pure sample data with the first type of first interference sample data according to the first processing probability and then performing the first type of interference processing to obtain the first processing data can include:
[0112] Determining whether to perform the first type of interference processing on the pure sample data according to the first processing probability;
[0113] If yes, the pure sample data is mixed with the first type of first interference sample data and then is subjected to the first type of interference processing to obtain first processing data.
[0114] For a certain pure sample data, a judgment can be made as to whether it is mixed with the first type of interference sample data and subjected to the first type of interference processing according to the first processing probability. If the judgment result is yes, the pure sample data is mixed with the first type of first interference sample data and then is subjected to the first type of interference processing to obtain first processing data.
[0115] Optionally, if the judgment result is no, the above processing is not performed and the pure sample data remains unchanged.
[0116] Correspondingly, the method for mixing the second type of interference sample data with the first type of second interference sample data according to the second processing probability and then subjecting the mixed sample data to the first type of interference processing to obtain second processing data can include:
[0117] determining whether the second type of interference sample data is subjected to the first type of interference processing according to the second processing probability;
[0118] If yes, the second type of interference sample data is mixed with the first type of second interference sample data and then is subjected to the first type of interference processing to obtain second processing data.
[0119] For a certain second type of interference sample data, a judgment can be made as to whether it is mixed with the first type of interference sample data and subjected to the first type of interference processing according to the second processing probability. If the judgment result is yes, the second type of interference sample data is mixed with the first type of first interference sample data and then is subjected to the first type of interference processing to obtain second processing data.
[0120] Optionally, if the judgment result is no, the above processing is not performed and the second type of interference sample data remains unchanged.
[0121] Further, the method for mixing the pure sample data or the first processing data with the second type of interference sample data or mixing the pure sample data or the first processing data with the second processing data to obtain target sample data can include:
[0122] in a case where it is determined that the pure sample data is not subjected to the first type of interference processing and the second type of interference sample data is not subjected to the first type of interference processing, mixing the pure sample data with the second type of interference sample data to obtain target sample data;
[0123] in a case where it is determined that the pure sample data is subjected to the first type of interference processing and the second type of interference sample data is not subjected to the first type of interference processing, mixing the first processing data with the second type of interference sample data to obtain target sample data;
[0124] If it is determined that the pure sample data has not undergone the first type of interference processing and the second type of interference sample data has undergone the first type of interference processing, the pure sample data and the second-processed data are mixed to obtain the target sample data.
[0125] Given that the clean sample data undergoes Type I interference processing and the Type II interference sample data undergoes Type I interference processing, the first processed data and the second processed data are mixed to obtain the target sample data.
[0126] like Figure 4 The diagram shown is a flowchart of another embodiment of a data processing method provided in this application. The method may include the following steps:
[0127] 401: After mixing the clean sample data with the first type of interference sample data, perform first type of interference processing to obtain the first processed data.
[0128] In this embodiment, for any clean sample data in the training dataset, it can be mixed with first type of first interference sample data and subjected to first type of interference processing to obtain first processed data.
[0129] 402: Mix the second type of interference sample data with the first type of second interference sample data and then perform first type of interference processing to obtain second processed data.
[0130] For any second-type interference sample data in the training dataset, it can be mixed with the first-type second-type interference sample data and subjected to first-type interference processing to obtain second-processed data.
[0131] 403: Select an appropriate number of clean sample data and first-processed data, and mix them with an appropriate number of second-type interference sample data and second-processed data respectively to obtain multiple target sample data.
[0132] Select a corresponding number of clean sample data and first-processed data from the training dataset, and mix them with a corresponding number of second-type interference sample data and second-processed data to obtain multiple target sample data. There are multiple ways to achieve this, which will be described in subsequent embodiments and will not be elaborated here.
[0133] 404: Train a second type of interference processing model using multiple target sample data and their corresponding clean sample data.
[0134] The second type of interference processing model is used to perform second type of interference processing on the data to be processed after the first type of interference processing in order to eliminate the second type of interference data in the data to be processed.
[0135] The plurality of target sample data are taken as training data, and the corresponding pure sample data are taken as training labels to train the second interference processing model.
[0136] In this embodiment, the pure sample data and the second interference sample data are pre-processed by the first interference processing, and a corresponding number of processed data are selected as sample data to participate in training to obtain the target sample data. The second interference processing model is trained by the target sample data, so that the second interference processing model can recognize the data processed by the first interference processing and perform the second interference processing to eliminate the second interference data. The problem that the pure data and the second interference data are affected by the first interference processing in the prior art, resulting in the problem of affecting the second interference processing effect, is solved, and the data processing quality is improved.
[0137] The target sample data can be obtained in various ways, which are described below.
[0138] As an optional implementation, the first number of pure sample data and the first number of second interference sample data can be one-to-one mixed, the second number of pure sample data and the second number of second processing data can be one-to-one mixed, the third number of first processing data and the third number of second interference sample data can be one-to-one mixed, and the fourth number of first processing data and the fourth number of second processing data can be one-to-one mixed; wherein the second number, the third number, and the fourth number are less than the first number.
[0139] For example, the first number can be 100, 100 pure sample data and 100 second interference sample data can be selected from the training data set to be one-to-one mixed to obtain 100 target sample data; the second number can be 50, 50 pure sample data and 50 second processing data can be selected from the training data set to be one-to-one mixed to obtain 50 target sample data; the third number can be 30, 30 first processing data and 30 second interference sample data can be selected from the training data set to be one-to-one mixed to obtain 30 target sample data; the fourth number can be 20, 20 first processing data and 20 second processing data can be selected from the training data set to be one-to-one mixed to obtain 20 target sample data. The second interference processing model is trained by using the above 200 target sample data and the corresponding pure sample data.
[0140] As another optional implementation manner, the fifth quantity of pure sample data can be mixed with data randomly selected from the sixth quantity of second type interference sample data or the seventh quantity of second processing data respectively to obtain a plurality of target sample data; and the eighth quantity of first processing data can be mixed with data randomly selected from the sixth quantity of second type interference sample data or the seventh quantity of second processing data respectively to obtain a plurality of target sample data; wherein the fifth quantity is greater than the eighth quantity, and the sixth quantity is greater than the seventh quantity.
[0141] As another optional implementation manner, the ninth quantity of pure sample data can be mixed with data randomly selected from a set composed of the tenth quantity of second type interference sample data and the eleventh quantity of second processing data respectively to obtain a plurality of target sample data; and the twelfth quantity of first processing data can be mixed with data randomly selected from the set composed of the tenth quantity of second type interference sample data and the eleventh quantity of second processing data respectively to obtain a plurality of target sample data; wherein the ninth quantity is greater than the twelfth quantity, and the tenth quantity is greater than the eleventh quantity.
[0142] The target sample data can also be obtained in other implementation manners, which can be set according to actual application scenarios and are not limited.
[0143] In actual application, the pure sample data can include pure speech sample data, the first type interference sample data can include echo sample data, reverberation sample data or howling sample data; and the second type interference sample data can include noise sample data.
[0144] Hereinafter, the technical solution of the present application is described by taking pure sample data as pure speech sample data, first type interference sample data as echo sample data, and second type interference sample data as noise sample data as an example. As shown in FIG. 1, a flowchart of one embodiment of a speech data processing method provided by an embodiment of the present application can include the following steps: Figure 5
[0145] 501: The pure speech sample data is mixed with the first echo sample data according to a first processing probability, and then echo cancellation processing is performed to obtain first processing data.
[0146] 502: The noise sample data is mixed with the second echo sample data according to a second processing probability, and then echo cancellation processing is performed to obtain second processing data.
[0147] 503: The pure speech sample data or the first processing data is mixed with the noise sample data, or the pure speech sample data or the first processing data is mixed with the second processing data to obtain target sample data.
[0148] 504: training the noise processing model by using the target sample data and the clean speech sample data.
[0149] The noise processing model is used to perform speech enhancement processing on the to-be-processed speech data after echo cancellation processing to eliminate noise data in the to-be-processed speech data.
[0150] In the embodiment, the pre-echo cancellation processing is performed on the clean speech sample data and the noise disturbance sample data according to the processing probability, the processed data is also used as sample data to participate in training to obtain target sample data, and the noise processing model is trained by using the target sample data. The noise processing model can identify the data after echo cancellation processing and perform speech enhancement processing on the data to eliminate noise data, thereby solving the problem that the echo cancellation processing affects the clean speech data and noise data in the prior art, affecting the speech enhancement effect, and improving the speech data processing quality.
[0151] Optionally, the noise sample data can be mixed with the second echo sample data and the clean speech sample data according to the second processing probability, and then the echo cancellation processing is performed to obtain second processing data.
[0152] Optionally, the clean speech sample data and the noise sample data can be randomly selected from the training data set, and then the clean speech sample data is processed according to the first processing probability and the noise sample data is processed according to the second processing probability.
[0153] Specifically, it is determined whether to perform echo cancellation processing on the clean speech sample data according to the first processing probability. If yes, the clean speech sample data is mixed with the first echo sample data and then the echo cancellation processing is performed to obtain first processing data. If no, the echo cancellation processing is not performed, and the clean speech sample data remains unchanged. It is also determined whether to perform echo cancellation processing on the noise sample data according to the second processing probability. If yes, the noise sample data is mixed with the second echo sample data and then the echo cancellation processing is performed to obtain second processing data. If no, the echo cancellation processing is not performed, and the noise sample data remains unchanged.
[0154] Furthermore, if it is determined that neither the clean speech sample data nor the noise sample data has undergone echo cancellation processing, the clean speech sample data and the noise sample data can be mixed to obtain the target sample data; if it is determined that the clean speech sample data has undergone first-type interference processing and the noise sample data has not undergone echo cancellation processing, the first-processed data can be mixed with the noise sample data to obtain the target sample data; if it is determined that neither the clean speech sample data nor the noise sample data has undergone echo cancellation processing, the clean speech sample data can be mixed with the second-processed data to obtain the target sample data; if it is determined that both the clean speech sample data and the noise sample data have undergone echo cancellation processing, the first-processed data and the second-processed data can be mixed to obtain the target sample data.
[0155] For ease of understanding, Figure 6 A schematic diagram of one embodiment of a noise processing model is shown. Figure 6 As shown, clean speech sample data is mixed with first echo sample data according to a first processing probability 'a' and then subjected to echo cancellation processing to obtain first processed data. Noise sample data is mixed with second echo sample data according to a second processing probability 'b' and then subjected to echo cancellation processing to obtain second processed data. If it is determined that neither the clean speech sample data nor the noise sample data has undergone echo cancellation processing, the clean speech sample data and the noise sample data are mixed to obtain target sample data. If it is determined that the clean speech sample data undergoes Type I interference processing and the noise sample data has not undergone echo cancellation processing, the first processed data and the noise sample data are mixed to obtain target sample data. If it is determined that neither the clean speech sample data nor the noise sample data has undergone echo cancellation processing, the clean speech sample data and the second processed data are mixed to obtain target sample data. If it is determined that both the clean speech sample data and the noise sample data have undergone echo cancellation processing, the first processed data and the second processed data can be mixed to obtain target sample data. The target sample data is used as the training sample and input into the noise processing model. The clean speech sample data is used as the training label to train the noise processing model.
[0156] In practical applications, Figure 5 When the technical solution of the illustrated embodiment is executed by the server, in some embodiments, the above method may further include: sending the noise processing model to the client so that the client can perform speech enhancement processing on the speech data to be processed after echo cancellation processing. The specific implementation method will be described in subsequent embodiments.
[0157] like Figure 7As shown, it is a flow chart of one embodiment of a voice data processing method provided by the embodiment of the present application. The method can include the following steps:
[0158] 701: Perform echo cancellation processing on the mixed pure voice sample data and the first echo sample data to obtain first processing data.
[0159] 702: Perform echo cancellation processing on the mixed noise sample data and the second echo sample data to obtain second processing data.
[0160] 703: Select a corresponding number of pure voice sample data and first processing data, respectively, and mix them with a corresponding number of noise sample data and second processing data to obtain a plurality of target sample data.
[0161] Optionally, the first number of pure voice sample data and the first number of noise interference sample data can be one-to-one mixed, the second number of pure voice sample data and the second number of second processing data can be one-to-one mixed, the third number of first processing data and the third number of noise sample data can be one-to-one mixed, and the fourth number of first processing data and the fourth number of second processing data can be one-to-one mixed; wherein the second number, the third number, and the fourth number are less than the first number.
[0162] Optionally, the fifth number of pure voice sample data can be mixed with data randomly selected from the sixth number of noise sample data or the seventh number of second processing data, respectively, to obtain a plurality of target sample data; and the eighth number of first processing data can be mixed with data randomly selected from the sixth number of noise sample data or the seventh number of second processing data, respectively, to obtain a plurality of target sample data; wherein the fifth number is greater than the eighth number, and the sixth number is greater than the seventh number.
[0163] Optionally, the ninth number of pure voice sample data can be mixed with data randomly selected from a set composed of the tenth number of noise sample data and the eleventh number of second processing data, respectively, to obtain a plurality of target sample data; and the twelfth number of first processing data can be mixed with data randomly selected from the set composed of the tenth number of noise sample data and the eleventh number of second processing data, respectively, to obtain a plurality of target sample data; wherein the ninth number is greater than the twelfth number, and the tenth number is greater than the eleventh number.
[0164] 704: Use the plurality of target sample data and the respective corresponding pure voice sample data to train a noise processing model.
[0165] The noise processing model is used to perform speech enhancement processing on the to-be-processed speech data after echo cancellation processing to eliminate noise data in the to-be-processed speech data.
[0166] In this embodiment, the pure speech sample data and the noise sample data are preprocessed by the first type of interference, and a corresponding number of processed data is selected as sample data for training. The target sample data is obtained, and the noise processing model is trained by the target sample data. The noise processing model can identify the data after echo cancellation processing, and perform speech enhancement processing to eliminate noise data. The problem of affecting the pure speech data and noise data by echo cancellation processing in the prior art, which affects the speech enhancement effect, is solved, and the speech data processing quality is improved.
[0167] In practical applications, Figure 7 When the technical solution of the embodiment shown is executed by the server, in some embodiments, the method can further include: distributing the noise processing model to the client to perform speech enhancement processing on the to-be-processed speech data after echo cancellation processing. The specific implementation will be described in subsequent embodiments.
[0168] The process of using the second type of interference processing model to process data is described below. The method can be executed by the client. Figure 8 is a flowchart of another embodiment of the data processing method provided by the present application. The method can include the following steps:
[0169] 801: Perform first type of interference processing on the to-be-processed data to obtain preprocessed data.
[0170] In this embodiment, the to-be-processed data can refer to mixed data containing pure data, first type of interference data, and second type of interference data, which can be obtained by the client.
[0171] 802: Use the second type of interference processing model to perform second type of interference processing on the preprocessed data to eliminate the second type of interference data in the preprocessed data and obtain target data.
[0172] The second type of interference processing model is trained using target sample data and pure sample data. The target sample data is obtained by mixing pure sample data or first processed data with second type of interference sample data, or by mixing pure sample data or first processed data with second processed data. The first processed data is obtained by mixing pure sample data with first type of first interference sample data and then performing first type of interference processing. The second processed data is obtained by mixing second type of interference sample data with first type of second interference sample data and then performing first type of interference processing.
[0173] The generation process of the second type of interference processing model has been described in detail in the foregoing embodiments, and will not be described again here.
[0174] In this embodiment, during training of the second type of interference processing model, the pure sample data and the second type of interference sample data are pre-processed by the first type of interference processing according to the processing probability. The processed data is also used as sample data for training to obtain target sample data. The second type of interference processing model is trained based on the target sample data, so that the second type of interference processing model can recognize the data processed by the first type of interference processing and perform the second type of interference processing to eliminate the second type of interference data. The problem that the pure data and the second type of interference data are affected by the first type of interference processing in the prior art, resulting in an impact on the effect of the second type of interference processing, is solved, and the data processing quality is improved.
[0175] In the following, taking the pure data as pure speech data, the first type of interference data as echo data, and the second type of interference data as noise data as an example, the process of processing speech data by using a noise processing model is described. The method can be executed by a client. Figure 9 This is a flowchart of another embodiment of a speech data processing method provided by the present application. The method can include the following steps:
[0176] 901: Obtain user speech data.
[0177] In this embodiment, the user speech data can refer to mixed speech data containing pure speech data, echo data, and noise data, which can be obtained by a client.
[0178] 902: Perform echo cancellation processing on the user speech data to obtain preprocessed speech data.
[0179] 903: Perform speech enhancement processing on the preprocessed speech data by using a noise processing model to eliminate noise data in the preprocessed speech data and obtain target speech data.
[0180] The noise processing model is trained by using target sample data and pure speech sample data. The target sample data is obtained by mixing pure speech sample data or first processing data with noise sample data, or by mixing pure speech sample data or first processing data with second processing data. The first processing data is obtained by mixing pure speech sample data with first echo sample data and then performing echo cancellation processing. The second processing data is obtained by mixing noise sample data with second echo sample data and then performing echo cancellation processing.
[0181] The specific training method of the noise processing model can be found in the foregoing detailed embodiments, which will not be described again here. The noise processing model can be trained by a server and then distributed to a client.
[0182] The generation process of the noise processing model has been described in detail in the foregoing embodiments, and will not be described again here.
[0183] 904: Send the target voice data.
[0184] In this embodiment, when the noise processing model is trained, the pure voice sample data and the noise interference sample data are pre-echo-canceled according to the processing probability, the processed data is also used as sample data to participate in the training, and the target sample data is obtained. The noise processing model is trained with the target sample data, which can realize that the noise processing model can identify the data after echo cancellation, and perform voice enhancement processing to eliminate noise data, thereby solving the problem that the echo cancellation processing in the prior art affects the pure voice data and noise data, and affects the voice enhancement effect, and improving the voice data processing quality.
[0185] The target voice data can be sent to the server and then sent to the receiving end via the server. Of course, it can be understood that, in order to facilitate and realize the transmission of voice data, the target voice data can also need to be encoded, compressed, and the like before being transmitted, which is not limited in the present application.
[0186] In actual application, taking a live broadcast scene as an example, the method for obtaining user voice data can include:
[0187] Collecting user voice data in a live broadcast scene. The user voice data may, for example, refer to voice data of an anchor.
[0188] The method for sending the target voice data can include:
[0189] Sending the target voice data to the server, so as to be sent to the receiving end by the server.
[0190] The receiving end may, for example, refer to a client that watches a live broadcast in a live broadcast scene.
[0191] Of course, the technical solution of the present application embodiment can be applied not only to a live broadcast scene, but also to other communication scenes, such as a conference, an online classroom, and the like real-time communication scenes. The user voice data can be collected by a sending end, processed by the technical solution of the present application embodiment to obtain target voice data, and then sent to a receiving end via a server.
[0192] For ease of understanding, in one actual application, the technical solution of the present application embodiment can be applied to a communication system architecture as shown in Figure 10 The communication system can be a real-time communication system, such as a live broadcast system, a conference system, an online chat room system, or an online classroom system, and the like. The specific implementation of the system architecture can refer to Figure 1aAs shown in the middle, in order to facilitate understanding, Figure 10 As shown in the middle, in order to facilitate understanding, Figure 10 As shown in the middle, in order to facilitate understanding,
[0193] In addition, the noise processing model can be pre-trained by the server B and distributed to the client A.
[0194] Figure 11 An embodiment of a data processing apparatus provided by the present application is shown in the structure diagram. The apparatus can include the following modules:
[0195] The first processing module 1101 is configured to mix the pure sample data with the first type of first interference sample data according to a first processing probability, and then perform first type of interference processing, to obtain first processing data.
[0196] The second processing module 1102 is configured to mix the second type of interference sample data with the first type of second interference sample data according to a second processing probability, and then perform first type of interference processing, to obtain second processing data.
[0197] The first mixing module 1103 is configured to mix the pure sample data or the first processing data with the second type of interference sample data, or mix the pure sample data or the first processing data with the second processing data, to obtain target sample data.
[0198] The first training module 1104 is configured to train a second type of interference processing model using the target sample data and the pure sample data. The second type of interference processing model is used to perform second type of interference processing on the to-be-processed data after the first type of interference processing, to eliminate the second type of interference data in the to-be-processed data.
[0199] In some embodiments, the second processing module 1102 can also be configured to mix the second type of interference sample data with the first type of second interference sample data and the pure sample data according to the second processing probability, and then perform first type of interference processing, to obtain second processing data.
[0200] In some embodiments, the apparatus can further include:
[0201] The selection module is configured to randomly select pure sample data and second-class interference sample data from the training data set;
[0202] The first processing module 1101 can be specifically configured to determine whether to perform first-class interference processing on the pure sample data according to a first processing probability; if yes, the pure sample data is mixed with first-class first interference sample data and then subjected to first-class interference processing to obtain first processing data;
[0203] The second processing module 1102 can be specifically configured to determine whether to perform first-class interference processing on the second-class interference sample data according to a second processing probability; if yes, the second-class interference sample data is mixed with first-class second interference sample data and then subjected to first-class interference processing to obtain second processing data;
[0204] The first mixing module 1103 can be specifically configured to mix the pure sample data with the second-class interference sample data to obtain target sample data in a case where it is determined that the pure sample data is not subjected to first-class interference processing and the second-class interference sample data is not subjected to first-class interference processing; mix the first processing data with the second-class interference sample data to obtain target sample data in a case where it is determined that the pure sample data is subjected to first-class interference processing and the second-class interference sample data is not subjected to first-class interference processing; mix the pure sample data with the second processing data to obtain target sample data in a case where it is determined that the pure sample data is not subjected to first-class interference processing and the second-class interference sample data is subjected to first-class interference processing; and mix the first processing data with the second processing data to obtain target sample data in a case where it is determined that the pure sample data is subjected to first-class interference processing and the second-class interference sample data is subjected to first-class interference processing.
[0205] Figure 11 The data processing apparatus can perform Figure 3 The data processing method of the embodiments shown above has the same implementation principles and technical effects, and will not be described in detail. The specific operation modes of each module and unit of the data processing apparatus in the above embodiments have been described in detail in the embodiments related to the method, and will not be described in detail here.
[0206] Figure 12 FIG. 1 is a structural schematic diagram of a voice data processing apparatus according to an embodiment of the present application. The apparatus can include the following modules:
[0207] The third processing module 1201 is configured to mix the pure voice sample data with the first echo sample data according to a first processing probability and then perform echo cancellation processing to obtain first processing data;
[0208] The fourth processing module 1202 is configured to perform echo cancellation processing on the noise sample data mixed with the second echo sample data according to the second processing probability, to obtain second processing data.
[0209] The second mixing module 1203 is configured to mix the clean speech sample data or the first processing data with the noise sample data, or mix the clean speech sample data or the first processing data with the second processing data, to obtain target sample data.
[0210] The second training module 1204 is configured to train the noise processing model by using the target sample data and the clean speech sample data.
[0211] Figure 12 The speech data processing apparatus can perform Figure 5 The speech data processing method described in the embodiments has the implementation principle and technical effects which will not be repeated here. The specific operation modes of each module and unit of the speech data processing apparatus in the above embodiments have been described in detail in the embodiments related to the method, and will not be described in detail here.
[0212] The embodiments of the present application further provide a data processing apparatus, which can include the following modules:
[0213] The fifth processing module is configured to perform first-type interference processing on the to-be-processed data, to obtain preprocessed data.
[0214] The sixth processing module is configured to perform second-type interference processing on the preprocessed data by using a second-type interference processing model, to eliminate second-type interference data in the preprocessed data, and obtain target data.
[0215] The data processing apparatus described in the embodiments can perform Figure 8 The data processing method described in the embodiments has the implementation principle and technical effects which will not be repeated here. The specific operation modes of each module and unit of the data processing apparatus in the above embodiments have been described in detail in the embodiments related to the method, and will not be described in detail here.
[0216] The embodiments of the present application further provide a speech data processing apparatus, which can include the following modules:
[0217] The acquisition module is configured to acquire user speech data.
[0218] The echo processing module is configured to perform echo cancellation processing on the user speech data, to obtain preprocessed speech data.
[0219] The voice enhancement processing module is configured to perform voice enhancement processing on the preprocessed voice data by using a noise processing model to eliminate noise data in the preprocessed voice data and obtain target voice data.
[0220] The sending module is configured to send the target voice data.
[0221] The voice data processing apparatus described in the embodiments can perform Figure 9 The voice data processing method described in the embodiments has the same implementation principles and technical effects as the voice data processing method described above. The specific manners in which the modules and units of the voice data processing apparatus perform operations have been described in detail in the embodiments of the method, and thus will not be described here.
[0222] The embodiments of the present application also provide a computing device, as shown in the figure, Figure 13 The device can include a storage component 1301 and a processing component 1302.
[0223] The storage component 1301 stores one or more computer program instructions, wherein the one or more computer program instructions are called and executed by the processing component 1302 to implement the data processing method shown in FIG. 1, Figure 2 、 Figure 3 、 Figure 4 、 Figure 8 the voice data processing method shown in the figure. Figure 5 、 Figure 7 、 Figure 9
[0224] Among them, Figure 10 The client in the above method can be configured in an electronic device.
[0225] The processing component can include one or more processors to execute computer instructions to complete all or part of the steps in the above method. Of course, the processing component can also be one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors or other electronic components, for executing the above method.
[0226] The storage component is configured to store various types of data to support the operation of the terminal. The storage component can be realized by any type of volatile or non-volatile storage device or their combination, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.
[0227] Of course, the electronic device described above can also include other components, such as an input / output interface, a communication component, and the like.
[0228] The input / output interface provides an interface between the processing component and a peripheral interface module, which can be an output device, an input device, and the like. The communication component is configured to facilitate wired or wireless communication between the computing device and other devices.
[0229] The embodiments of the present application also provide a computer readable storage medium storing a computer program, which, when executed by a computer, can implement the data processing method shown in FIG. 1, Figure 2 、 Figure 3 、 Figure 4 、 Figure 8 the voice data processing method shown in FIG. 1, Figure 5 、 Figure 7 、 Figure 9 . The computer readable medium can be included in the electronic device described in the above embodiments; or can exist separately and not be assembled into the electronic device.
[0230] The embodiments of the present application also provide a computer program product, which includes a computer program carried on a computer readable storage medium, and the computer program, when executed by a computer, can implement the data processing method shown in FIG. 1, Figure 2 、 Figure 3 、 Figure 4 、 Figure 8 the voice data processing method shown in FIG. 1, Figure 5 、 Figure 7 、 Figure 9 .
[0231] In such embodiments, the computer program can be downloaded and installed from a network, and / or installed from a removable medium. When the computer program is executed by a processor, various functions defined in the system of the present application are performed.
[0232] It should be noted that the electronic device described above can be a physical device or an elastic computing host provided by a cloud computing platform. It can be implemented as a distributed cluster composed of multiple servers or terminal devices, or as a single server or a single terminal device.
[0233] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the system, device and unit described above can refer to the corresponding process in the foregoing method embodiments, which will not be described here.
[0234] The device embodiments described above are merely illustrative, wherein the units described as separate components can or can not be physically separate, and the components displayed as units can or can not be physical units, i.e., can be located in one place, or can be distributed to multiple network units. Part or all of the modules can be selected to achieve the purposes of the embodiments according to actual needs. Those skilled in the art can understand and implement without creative labor.
[0235] Through the description of the above embodiments, those skilled in the art can clearly understand that the embodiments can be realized by means of software and necessary general hardware platforms, and of course can also be realized by hardware. Based on such understanding, the above technical solutions can be embodied in the form of software products, and the computer software products can be stored in a computer readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and include a plurality of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute the methods described in each embodiment or some parts of the embodiments.
[0236] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A voice data processing method characterized by comprising: The method comprises the following steps: mixing the pure speech sample data with the first type of first interference sample data according to a first processing probability, and then performing first type interference processing to obtain first processing data; mixing the second type of interference sample data with the first type of second interference sample data according to a second processing probability, and then performing first type interference processing to obtain second processing data; mixing the first processing data with the second type of interference sample data, or mixing the pure speech sample data with the second processing data, or mixing the first processing data with the second processing data, to obtain target sample data; training a second type interference processing model by using the target sample data and the pure speech sample data, wherein the second type interference processing model is used to perform second type interference processing on the to-be-processed data after first type interference processing to eliminate the second type interference data in the to-be-processed data.
2. The method of claim 1, wherein, The method further comprises the following steps: randomly selecting the pure speech sample data and the second type of interference sample data from a training data set; 3. The method of claim 1, wherein, The method further comprises the following steps: determining whether to perform first type interference processing on the pure speech sample data according to the first processing probability; if yes, mixing the pure speech sample data with the first type of first interference sample data, and then performing first type interference processing to obtain first processing data; The method further comprises the following steps: determining whether to perform first type interference processing on the second type of interference sample data according to the second processing probability; if yes, mixing the second type of interference sample data with the first type of second interference sample data, and then performing first type interference processing to obtain second processing data; The method further comprises the following steps: in a case where it is determined that the pure speech sample data is subjected to first type interference processing and the second type of interference sample data is not subjected to first type interference processing, mixing the first processing data with the second type of interference sample data to obtain target sample data; in a case where it is determined that the pure speech sample data is subjected to first type interference processing and the second type of interference sample data is not subjected to first type interference processing, mixing the first processing data with the second type of interference sample data to obtain target sample data; In a case where it is determined that the pure speech sample data is not subjected to the first type of interference processing and the second type of interference sample data is subjected to the first type of interference processing, mixing the pure speech sample data and the second processing data to obtain target sample data; In a case where it is determined that the pure speech sample data is subjected to the first type of interference processing and the second type of interference sample data is subjected to the first type of interference processing, mixing the first processing data and the second processing data to obtain target sample data.
4. The method of claim 1, wherein, The pure sample data includes pure speech sample data, the first type of interference sample data includes echo sample data, reverberation sample data or howling sample data, and the second type of interference sample data includes noise sample data.
5. A voice data processing method characterized by comprising: The method comprises: mixing the pure speech sample data and the first echo sample data according to a first processing probability, and then performing echo cancellation processing to obtain first processing data; mixing the noise sample data and the second echo sample data according to a second processing probability, and then performing echo cancellation processing to obtain second processing data; mixing the first processing data and the noise sample data, or mixing the pure speech sample data and the second processing data, or mixing the first processing data and the second processing data to obtain target sample data; training a noise processing model by using the target sample data and the pure speech sample data; The noise processing model is used to perform speech enhancement processing on the to-be-processed speech data after echo cancellation processing to eliminate noise data in the to-be-processed speech data.
6. The method of claim 5, wherein, The method further comprises: distributing the noise processing model to a client, so that the client performs speech enhancement processing on the to-be-processed speech data after echo cancellation processing.
7. A voice data processing method characterized by comprising: The method comprises: mixing the pure speech sample data and the first type of first interference sample data, and then performing first type of interference processing to obtain first processing data; mixing the second type of interference sample data and the first type of second interference sample data, and then performing first type of interference processing to obtain second processing data; selecting a corresponding number of pure speech sample data and first processing data, and mixing them with a corresponding number of second type of interference sample data and second processing data respectively to obtain a plurality of target sample data; training a second type of interference processing model by using the plurality of target sample data and the respective corresponding pure speech sample data; the second type of interference processing model is used to perform second type of interference processing on to-be-processed data after first type of interference processing to eliminate second type of interference data in the to-be-processed data.
8. The method of claim 7, wherein, The method of selecting a corresponding number of pure speech sample data and first processing data, and mixing them with a corresponding number of second type of interference sample data and second processing data respectively to obtain a plurality of target sample data comprises: The first number of pure speech sample data is mixed with the first number of second type interference sample data one by one, the second number of pure speech sample data is mixed with the second number of second processing data one by one, the third number of first processing data is mixed with the third number of second type interference sample data one by one, and the fourth number of first processing data is mixed with the fourth number of second processing data one by one; wherein, the second number, the third number and the fourth number are less than the first number.
9. A voice data processing method characterized by comprising: Comprise: After mixing the pure speech sample data with the first echo sample data, perform echo cancellation processing to obtain first processing data; After mixing the noise sample data with the second echo sample data, perform echo cancellation processing to obtain second processing data; Select a corresponding number of pure speech sample data and first processing data, respectively, and mix a corresponding number of noise sample data and second processing data to obtain a plurality of target sample data; Use the plurality of target sample data and the respective corresponding pure speech sample data to train a noise processing model; The noise processing model is used to perform speech enhancement processing on the to-be-processed speech data after echo cancellation processing to eliminate noise data in the to-be-processed speech data.
10. A voice data processing method characterized by comprising: Comprise: After mixing the pure speech sample data with the first type interference sample data according to a third processing probability, perform first type interference processing to obtain third processing data; Mix the third processing data with the second type interference sample data to obtain target sample data; Use the target sample data and the pure speech sample data to train a second type interference processing model; the second type interference processing model is used to perform second type interference processing on the to-be-processed data after first type interference processing to eliminate second type interference data in the to-be-processed data.
11. A voice data processing method characterized by comprising: Comprise: After mixing the second type interference sample data with the first type interference sample data according to a fourth processing probability, perform first type interference processing to obtain fourth processing data; Mix the pure speech sample data with the fourth processing data to obtain target sample data; Use the target sample data and the pure speech sample data to train a second type interference processing model; the second type interference processing model is used to perform second type interference processing on the to-be-processed data after first type interference processing to eliminate second type interference data in the to-be-processed data.
12. A voice data processing method characterized by comprising: Comprise: Perform first type interference processing on the to-be-processed data to obtain preprocessed data; Use the second type interference processing model to perform second type interference processing on the preprocessed data to eliminate second type interference data in the preprocessed data to obtain target data; The second type of interference processing model is trained by using target sample data and pure speech sample data; the target sample data is obtained by mixing first processing data and second type of interference sample data, or by mixing the pure speech sample data and second processing data, or by mixing the first processing data and the second processing data; the first processing data is obtained by mixing pure speech sample data and first type of first interference sample data and then performing first type of interference processing; and the second processing data is obtained by mixing second type of interference sample data and first type of second interference sample data and then performing first type of interference processing.
13. A voice data processing method characterized by comprising: The method comprises: obtaining user speech data; performing echo cancellation processing on the user speech data to obtain preprocessed speech data; performing speech enhancement processing on the preprocessed speech data by using a noise processing model to eliminate noise data in the preprocessed speech data and obtain target speech data; wherein the noise processing model is trained by using target sample data and pure speech sample data; the target sample data is obtained by mixing first processing data and noise sample data, or by mixing the pure speech sample data and second processing data, or by mixing the first processing data and the second processing data; the first processing data is obtained by mixing pure speech sample data and first echo sample data and then performing echo cancellation processing; and the second processing data is obtained by mixing noise sample data and second echo sample data and then performing echo cancellation processing; sending the target speech data.
14. An electronic device, comprising: The storage component stores one or more computer program instructions, and the processing component invokes and executes the one or more computer program instructions to implement the data processing method in any one of claims 1-4 or any one of claims 7-8 or any one of claims 10-12, or to implement the speech data processing method in any one of claims 5-6 or claim 9 or claim 13.
Citation Information
Patent Citations
Voice signal processing method and device and electronic equipment
CN113286047A
Speech enhancement model training method and device and speech enhancement method and device
CN113823301A