An audio information processing method, apparatus and electronic device
By jointly training the audio content extractor, voiceprint feature extractor and audio synthesizer, voiceprint desensitized audio information is generated, which solves the problem of insufficient protection of user voiceprint information and realizes efficient and safe protection of user information.
Patent Information
- Application Number
- CN202210329461.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-31
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-03-31
AI Technical Summary
The prior art has shortcomings in protecting user voiceprint information, especially in recognition technology, voiceprint information is easily leaked, resulting in a decrease in user information security.
By obtaining sample audio information, the pre-trained audio content extractor, voiceprint feature extractor and audio synthesizer are jointly trained to obtain the audio information to be desensitized and the target voiceprint audio information. The extractor extracts the audio content and voiceprint features and inputs them into the audio synthesizer to generate voiceprint desensitized audio information.
Effectively hide user's voiceprint information, improve user information security, avoid voiceprint information leakage, and retain other audio features in the audio information.
Smart Images

Figure CN114937457B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computers, and in particular to an audio information processing method, device and electronic device. Background Art
[0002] With the development of technology, voice, especially voiceprint information, has gradually been regarded as an important user information and widely used in identification technology, such as to verify the identity of the user. However, everything has two sides. The emergence of this technology will invisibly leak the user information contained in the voice, especially the user's privacy information, resulting in a decrease in the security of user information. At present, the main focus is on desensitizing the user information content, that is, eliminating the ID number, name and address information in the call recording, but the voiceprint information is often ignored and does not get the protection it deserves. Summary of the invention
[0003] The embodiments of this specification provide an audio information processing method, device and electronic device to hide the user's voiceprint information and ensure the security of user information.
[0004] The present invention provides an audio information processing method, including:
[0005] Acquire sample audio information, and jointly train a pre-trained audio content extractor, a voiceprint feature extractor, and an audio synthesizer according to the sample audio information;
[0006] Acquire the audio information to be desensitized and the target voiceprint audio information, extract the audio content of the audio information to be desensitized by using the audio content extractor, and extract the target voiceprint feature of the target voiceprint audio information by using the voiceprint feature extractor;
[0007] The audio content and the target voiceprint feature are input into the audio synthesizer, and after being processed by the audio synthesizer, voiceprint desensitized audio information is output.
[0008] Optionally, the jointly training a pre-trained audio content extractor, a voiceprint feature extractor, and an audio synthesizer according to the sample audio information includes:
[0009] Extracting audio content from the sample audio information using the audio content extractor;
[0010] Extracting voiceprint features from the sample audio information using the voiceprint feature extractor;
[0011] The audio synthesizer predicts an audio point according to the audio content and the voiceprint feature;
[0012] Calculate the deviation value based on the predicted audio points and the sample audio information, and adjust the parameters of the audio content extractor, the voiceprint feature extractor, and the audio synthesizer according to the deviation value, and perform iterative training until the calculated deviation value is less than the threshold.
[0013] Optionally, calculating the deviation value based on the predicted audio points and the sample audio information includes:
[0014] Sample from the waveform of the sample audio information, calculate the difference between each sampling point and the predicted audio points, and sum all the differences to obtain the deviation value.
[0015] Optionally, the method further includes: pre-training the audio content extractor and the voiceprint feature extractor using the sample audio information, including:
[0016] Set an audio label for the sample audio information according to the audio content in the sample audio information, and train the audio content extractor based on the audio label;
[0017] Set different voiceprint labels according to different users to whom the different sample audio information belongs, and train the voiceprint feature extractor based on the voiceprint labels.
[0018] Optionally, after being processed by the audio synthesizer, outputting voiceprint desensitized audio information includes:
[0019] The audio synthesizer predicts the audio points of the audio content according to the frequency domain distribution of the target voiceprint feature, and synthesizes the voiceprint desensitized audio information according to the predicted audio points.
[0020] Optionally, it further includes:
[0021] Deploy the jointly trained audio content extractor, the voiceprint feature extractor, and the audio synthesizer to a network access device;
[0022] The network access device obtains the audio information to be desensitized sent by the terminal, desensitizes it using the deployed model, synthesizes the voiceprint desensitized audio information, and uploads the voiceprint desensitized audio information to the network.
[0023] An embodiment of this specification further provides an audio information processing device, including:
[0024] A joint training module, which obtains sample audio information and jointly trains a pre-trained audio content extractor, a voiceprint feature extractor, and an audio synthesizer according to the sample audio information;
[0025] A feature extraction module that obtains the audio information to be desensitized and the target voiceprint audio information, and is used to extract the audio content of the audio information to be desensitized by using the audio content extractor, and extract the target voiceprint feature in the target voiceprint audio information by using the voiceprint feature extractor;
[0026] An audio synthesis module that inputs the audio content and the target voiceprint feature into the audio synthesizer, and after being processed by the audio synthesizer, outputs the voiceprint desensitized audio information.
[0027] Optionally, the joint training of the pre-trained audio content extractor, voiceprint feature extractor, and audio synthesizer according to the sample audio information includes:
[0028] Using the audio content extractor to extract the audio content from the sample audio information;
[0029] Using the voiceprint feature extractor to extract the voiceprint feature from the sample audio information;
[0030] The audio synthesizer predicts audio points according to the audio content and the voiceprint feature;
[0031] Calculating a deviation value according to the predicted audio points and the sample audio information, and adjusting the parameters of the audio content extractor, the voiceprint feature extractor, and the audio synthesizer according to the deviation value, and performing iterative training until the calculated deviation value is less than a threshold.
[0032] An embodiment of this specification also provides an electronic device, where the electronic device includes:
[0033] A processor; and,
[0034] A memory storing a computer-executable program, and when the executable program is executed, the processor executes any one of the above methods.
[0035] An embodiment of this specification also provides a computer-readable storage medium, where the computer-readable storage medium stores one or more programs, and when the one or more programs are executed by a processor, any one of the above methods is implemented.
[0036] The various technical solutions provided in the embodiments of this specification obtain sample audio information, jointly train a pre-trained audio content extractor, a voiceprint feature extractor, and an audio synthesizer, obtain the audio information to be desensitized and the target voiceprint audio information, extract the audio content therein using the audio content extractor, extract the target voiceprint features therein using the voiceprint feature extractor, input the audio content and the target voiceprint features into the audio synthesizer, and after being processed by the audio synthesizer, output the voiceprint desensitized audio information. Through joint training, the training audio content extractor and the voiceprint feature extractor can accurately extract the audio content and the voiceprint features, enabling the audio synthesizer to learn the logic of synthesizing audio by combining the audio content and the voiceprint features. When processing the audio information to be desensitized and the target voiceprint audio information, it synthesizes the desensitized audio information with the target voiceprint, hiding the user's voiceprint information and ensuring the security of user information. Brief Description of the Drawings
[0037] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The illustrative embodiments and descriptions thereof of the present application are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:
[0038] Figure 1 It is a schematic diagram of the principle of an audio information processing method provided by an embodiment of this specification;
[0039] Figure 2 It is a schematic diagram of the structure of an audio information processing method provided by an embodiment of this specification;
[0040] Figure 3 It is a schematic diagram of the principle of an audio information processing device provided by an embodiment of this specification;
[0041] Figure 4 It is a schematic diagram of the structure of an electronic device provided by an embodiment of this specification;
[0042] Figure 5 It is a schematic diagram of the principle of a computer-readable medium provided by an embodiment of this specification. Detailed Embodiments
[0043] Now, the exemplary embodiments of the present invention will be described more fully with reference to the accompanying drawings. However, the exemplary embodiments can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, providing these exemplary embodiments enables the present invention to be more comprehensive and complete, and more conveniently conveys the inventive concept to those skilled in the art. Identical reference numerals in the figures represent identical or similar elements, components, or parts, and thus their repeated description will be omitted.
[0044] Under the premise of being consistent with the technical concept of the present invention, the features, structures, characteristics or other details described in a specific embodiment do not exclude that they can be combined in one or more other embodiments in a suitable manner.
[0045] In the description of specific embodiments, the features, structures, characteristics or other details described in the present invention are intended to enable those skilled in the art to fully understand the embodiments. However, it does not exclude that those skilled in the art can practice the technical solutions of the present invention without one or more of the specific features, structures, characteristics or other details.
[0046] The flowcharts shown in the accompanying drawings are only exemplary and do not necessarily include all the contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps can be decomposed, and some operations / steps can be combined or partially combined, so the actual execution order may change according to actual conditions.
[0047] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities may be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0048] The term "and / or" or "and / or" includes all combinations of any one or more of the associated listed items.
[0049] Figure 1 A schematic diagram of a method for processing audio information provided in an embodiment of this specification, the method may include:
[0050] S101: Acquire sample audio information, and jointly train a pre-trained audio content extractor, voiceprint feature extractor, and audio synthesizer according to the sample audio information.
[0051] In the disclosed embodiment, the audio content includes not only text information, but also other audio information such as pauses, speaking volume, tone, and speaking breath; and the voiceprint features in the disclosed embodiment include not only timbre features, but also accent features, such as the erhua sound of Beijingers, and the indistinguishability of "f" and "h" of Fujianese. In the prior art, voice calls are usually translated directly into text, and the transmission of information loss is basically ignored, which requires extremely high accuracy of the intermediate results. If the voice environment is noisy or the speaker's Mandarin is not standard, information loss will occur and will be continuously amplified during the transmission process; in the disclosed embodiment, in order to avoid information loss transmission, the volume of the audio content, the accent of the human voice in the audio, and other factors are taken into consideration, effectively reducing information loss.
[0052] In order to desensitize the voice, hide the real voiceprint information, and achieve the change of voice timbre, a tool that can synthesize a new audio according to the voiceprint characteristics and audio content in the audio, namely an audio synthesizer, can be constructed.
[0053] When training the audio synthesizer, the voiceprint characteristics and audio content need to be utilized. However, the voiceprint characteristics and audio content in the sample audio information are fused together. Therefore, before training the audio synthesizer, a tool that can extract (or separate) the voiceprint characteristics and audio content respectively needs to be constructed.
[0054] Therefore, a tool for extracting audio content from sample audio, namely an audio content extractor, and a tool for extracting voiceprint characteristics from sample audio, namely a voiceprint feature extractor, can be pre-trained first. Then, the audio content extractor and the voiceprint feature extractor can be used as the input layer of the audio synthesizer for joint training. In this way, the final trained composite model includes an audio content extractor, a voiceprint feature extractor, and an audio synthesizer.
[0055] Before training, a large amount of audio information of users can be collected as sample audio information, and then labels can be set for the sample audio information for training.
[0056] Among them, the source of the audio information can have various ways. It can be the audio already existing in the database, the audio temporarily obtained by the service platform, or the audio sent by the user terminal to the access point.
[0057] Specifically, in the embodiments of this specification, the method further includes: pre-training an audio content extractor and a voiceprint feature extractor using the sample audio information, including:
[0058] Setting an audio label for the sample audio information according to the audio content in the sample audio information, and training the audio content extractor based on the audio label;
[0059] Setting different voiceprint labels according to different users to whom the different sample audio information belongs, and training the voiceprint feature extractor based on the voiceprint labels.
[0060] Among them, the audio content can be manually identified from the audio label, or semantic recognition can be performed using a program to obtain the audio content, and after obtaining the audio content, a label is set according to the audio content.
[0061] Generally, the timbres of different audios of the same user are the same, and the voiceprints (including timbres) of different users are different. Therefore, the same voiceprint label can be set for different audios of the same user, different voiceprint labels can be set for different audios of different users, and multiple sample audio information with different semantics, speech rates, and loudnesses emitted by the same user can also be collected to set the same voiceprint label to improve the sample diversity.
[0062] Among them, the audio content extractor and the voiceprint feature extractor can convert the audio information in the time domain form into the audio information in the frequency domain form. The audio information in the frequency domain form can be recorded by feature vectors, and multiple feature vectors form a feature matrix. The process of feature extraction is performed using the feature matrix, so that the audio content extractor and the voiceprint feature extractor can accurately learn the features of the audio and timbre attributes in the frequency domain representation information respectively.
[0063] Among them, the process of converting audio information into a feature matrix can be regarded as encoding. Correspondingly, the process of generating audio from the feature matrix can be regarded as decoding.
[0064] In order to enable the audio synthesizer to accurately learn the processing logic of the synthesized audio and further optimize the audio content extractor and the voiceprint feature extractor, therefore, in the embodiments of this specification, the audio content extractor, the voiceprint feature extractor, and the audio synthesizer are jointly trained.
[0065] In the embodiments of this specification, the audio synthesizer can specifically be a distribution model, which is used to predict audio points according to the required frequency domain distribution. The audio synthesizer can have a decoder for converting the data in matrix form into audio.
[0066] Among them, the inputs of the audio synthesizer are the audio content and the voiceprint feature. Therefore, the audio content extractor, the voiceprint feature extractor, and the audio synthesizer can be associated, and the outputs of the audio content extractor and the voiceprint feature extractor are used as the inputs of the audio synthesizer.
[0067] During the joint training, iteration can be performed according to the accuracy of the audio synthesized by the audio synthesizer. The accuracy can be characterized by calculating the deviation value according to the predicted audio points and the sample audio information.
[0068] Therefore, in the embodiments of this specification, the joint training of the pre-trained audio content extractor, the voiceprint feature extractor, and the audio synthesizer according to the sample audio information includes:
[0069] Using the audio content extractor to extract the audio content from the sample audio information;
[0070] Using the voiceprint feature extractor to extract the voiceprint feature from the sample audio information;
[0071] The audio synthesizer predicts audio points according to the audio content and the voiceprint feature;
[0072] Calculate the deviation value based on the predicted audio points and the sample audio information, and adjust the parameters of the audio content extractor, the voiceprint feature extractor, and the audio synthesizer according to the deviation value, and perform iterative training until the calculated deviation value is less than the threshold.
[0073] Through joint training, the parameters of the audio content extractor, the voiceprint feature extractor, and the audio synthesizer are continuously optimized. Eventually, the audio content extractor can accurately extract the audio content, the voiceprint feature extractor can accurately extract the voiceprint features, and the audio synthesizer can accurately synthesize according to the extracted audio content and voiceprint features.
[0074] Among them, calculating the deviation value is essentially an evaluation of the similarity of the audio and timbre between the synthesized audio and the sample audio.
[0075] There are various ways to calculate the deviation value. For example, take the difference between the waveform of the synthesized audio and the sample audio, or take the difference after converting the synthesized audio to the frequency domain, and then quantify the difference into a numerical value.
[0076] Specifically, in the embodiments of this specification, the calculating the deviation value based on the predicted audio points and the sample audio information may include:
[0077] Sample from the waveform of the sample audio information, calculate the difference between each sampling point and the predicted audio points, and sum all the differences to obtain the deviation value.
[0078] S102: Obtain the audio information to be desensitized and the target voiceprint audio information, use the audio content extractor to extract the audio content of the audio information to be desensitized, and use the voiceprint feature extractor to extract the target voiceprint features of the target voiceprint audio information.
[0079] The sample audio information is used in the joint training stage. However, in the usage stage, since the ultimate goal is to make the synthesized audio have specific voiceprint features, that is, the target voiceprint features, it is necessary to obtain the audio information to be desensitized and the target voiceprint audio information.
[0080] Among them, the target voiceprint audio information may specifically be the audio of the artificial customer service timbre or the audio of the robot timbre, which will not be specifically elaborated and limited here.
[0081] Input the audio content of the audio information to be desensitized into the audio content extractor, and the audio content extractor can extract the audio content therein.
[0082] Input the target voiceprint audio information into the voiceprint feature extractor, and the voiceprint feature extractor can extract the voiceprint features therein, that is, the target voiceprint features.
[0083] S103: Input the audio content and the target voiceprint feature into the audio synthesizer. After being processed by the audio synthesizer, the voiceprint desensitized audio information is output.
[0084] By obtaining sample audio information, jointly train the pre-trained audio content extractor, voiceprint feature extractor, and audio synthesizer to obtain the audio information to be desensitized and the target voiceprint audio information. Use the audio content extractor to extract the audio content therein, use the voiceprint feature extractor to extract the target voiceprint feature therein, input the audio content and the target voiceprint feature into the audio synthesizer, and after being processed by the audio synthesizer, output the voiceprint desensitized audio information. Through joint training, the trained audio content extractor and voiceprint feature extractor can accurately extract the audio content and voiceprint feature, enabling the audio synthesizer to learn the logic of synthesizing audio by combining the audio content and voiceprint feature. When processing the audio information to be desensitized and the target voiceprint audio information, desensitized audio information with the target voiceprint is synthesized, hiding the user's voiceprint information and ensuring user information security.
[0085] When the audio synthesizer synthesizes using the audio content and the target voiceprint feature in the audio information to be desensitized, the synthesized audio not only has the audio in the audio information to be desensitized but also has the target voiceprint. It not only well hides the original voiceprint of the audio information to be desensitized, realizes audio desensitization, hides the user's real voiceprint information in the audio to be desensitized, and ensures user information security; moreover, it retains to the greatest extent other audio information in the audio information to be desensitized except for the text information; at the same time, the technical solution disclosed in this application does not require pre-saving or establishing a voiceprint library / timbre library in advance, saving a large amount of costs.
[0086] Specifically, the synthesis process can be regarded as the reverse process of feature extraction. Essentially, it is to predict audio points according to the distribution model, so that the frequency domain distribution performance of the predicted numerous audio points is consistent with the frequency domain distribution performance of the target voiceprint.
[0087] Specifically, in the embodiments of this specification, after being processed by the audio synthesizer, outputting the voiceprint desensitized audio information includes:
[0088] The audio synthesizer predicts the audio points of the audio content according to the frequency domain distribution of the target voiceprint feature, and synthesizes the voiceprint desensitized audio information according to the predicted audio points.
[0089] After joint training, the trained model can be deployed in the business platform for desensitization; it can also be deployed in the network access device connected to the Internet for desensitization before uploading the audio of the client to the Internet.
[0090] Therefore, in the embodiments of this specification, this method may further include:
[0091] Deploy the jointly trained audio content extractor, the voiceprint feature extractor, and the audio synthesizer to a network access device;
[0092] The network access device obtains the audio information to be desensitized sent by a terminal, desensitizes it using the deployed model, synthesizes voiceprint desensitized audio information, and uploads the voiceprint desensitized audio information to the network.
[0093] Figure 2 It is a schematic diagram of the principle of an audio information processing method provided by an embodiment of this specification, showing the principle of training a voiceprint feature converter.
[0094] Before training, first construct the structures of an audio content extractor (Ec), a voiceprint feature extractor (Es), and the audio synthesizer (D). Among them, the audio content extractor and the voiceprint feature extractor can have encoders for converting audio into a matrix format, and the audio synthesizer has a decoder for finally converting the calculated matrix into an audio format.
[0095] Then, obtain sample audio information X1, Z1, U1..., set audio labels according to the sample audio information, set voiceprint labels according to the timbre characteristics (including but not limited to timbre and breath, etc.) of the sample audio information, and use the labeled sample audio information to train the audio content extractor (Ec) and the voiceprint feature extractor (Es) respectively.
[0096] After pre-training, connect the output ends of the audio content extractor (Ec) and the voiceprint feature extractor (Es) to the input end of the audio synthesizer (D) respectively for joint training.
[0097] Specifically, the audio content extractor (Ec) outputs audio content C1, the voiceprint feature extractor (Es) outputs voiceprint feature S1, and the audio synthesizer (D) synthesizes audio according to the audio content C1 and the voiceprint feature S1, denoted as The synthesized audio is evaluated and quantified with the sample audio information X1 (usually by taking the difference). The quantified deviation value can reflect whether the synthesis ability of the audio synthesizer meets the preset requirements. If it does not meet the requirements, adjust the parameters of the audio content extractor (Ec), the voiceprint feature extractor (Es), and the audio synthesizer (D).
[0098] Subsequently, continue to perform iterative training on the audio content extractor (Ec), the voiceprint feature extractor (Es), and the audio synthesizer (D) with adjusted parameters until the calculated deviation value is less than the threshold.
[0099] Different from the pre-training stage, in the joint training stage, the parameters in the audio content extractor (Ec) and the voiceprint feature extractor (Es) are adjusted according to the audio difference and timbre difference between the audio synthesized by the audio synthesizer (D) and the sample audio information, rather than directly adjusted according to the outputs of the audio content extractor (Ec) and the voiceprint feature extractor (Es) respectively. Therefore, the audio content extractor (Ec) and the voiceprint feature extractor (Es) can adapt to the audio synthesizer (D), making the jointly trained composite model most conducive to the accurate synthesis of audio. By using the joint training method, each model is trained as a whole, which externally appears as a de-sensitized model. The achieved effect is that when the voice is input into the de-sensitized model, the de-sensitized voice is directly output. Among them, the text content of the audio no longer undergoes speech recognition (translating the text content) and subsequent speech synthesis, ensuring a certain degree of fault tolerance for the audio content.
[0100] During the process of performing audio information processing, instead of inputting the sample audio to the audio content extractor (Ec) and the voiceprint feature extractor (Es), the audio content to be de-sensitized is input to the audio content extractor (Ec), and the target voiceprint audio information is input to the voiceprint feature extractor (Es). The audio content extractor (Ec) and the voiceprint feature extractor (Es) respectively extract the audio content C1 and the voiceprint feature S1, and their outputs are passed to the audio synthesizer (D), and the audio synthesizer (D) synthesizes the audio and outputs the synthesized audio. That is, the audio content is expressed using the target voiceprint, realizing the de-sensitization of the audio information, and well hiding the voiceprint information of the user corresponding to the audio content to be de-sensitized.
[0101] Furthermore, the de-sensitized audio information can be provided to various speech processing software for further processing and application, such as automatic speech recognition (ASR), automatic speaker verification (ASV), etc.
[0102] Figure 3 The following is a schematic structural diagram of an audio information processing device provided by an embodiment of this specification. The device may include:
[0103] A joint training module 301, which obtains sample audio information and jointly trains the pre-trained audio content extractor, voiceprint feature extractor, and audio synthesizer according to the sample audio information;
[0104] A feature extraction module 302, which obtains the audio content to be de-sensitized and the target voiceprint audio information, and is used to extract the audio content of the audio content to be de-sensitized by using the audio content extractor, and extract the target voiceprint feature in the target voiceprint audio information by using the voiceprint feature extractor;
[0105] The audio synthesis module 303 inputs the audio content and the target voiceprint feature into the audio synthesizer. After being processed by the audio synthesizer, voiceprint desensitized audio information is output.
[0106] Among them, the joint training of the pre-trained audio content extractor, voiceprint feature extractor, and audio synthesizer according to the sample audio information includes:
[0107] Using the audio content extractor to extract audio content from the sample audio information;
[0108] Using the voiceprint feature extractor to extract voiceprint features from the sample audio information;
[0109] The audio synthesizer predicts audio points according to the audio content and the voiceprint feature;
[0110] Calculate the deviation value according to the predicted audio points and the sample audio information, and adjust the parameters of the audio content extractor, the voiceprint feature extractor, and the audio synthesizer according to the deviation value, and perform iterative training until the calculated deviation value is less than the threshold.
[0111] The device obtains sample audio information, jointly trains the pre-trained audio content extractor, voiceprint feature extractor, and audio synthesizer, obtains the audio information to be desensitized and the target voiceprint audio information, extracts the audio content therein by using the audio content extractor, extracts the target voiceprint feature therein by using the voiceprint feature extractor, inputs the audio content and the target voiceprint feature into the audio synthesizer, and after being processed by the audio synthesizer, outputs voiceprint desensitized audio information. Through joint training, the trained audio content extractor and voiceprint feature extractor can accurately extract audio content and voiceprint features, enabling the audio synthesizer to learn the logic of synthesizing audio by combining audio content and voiceprint features. When processing the audio information to be desensitized and the target voiceprint audio information, desensitized audio information with the target voiceprint is synthesized, hiding the user's voiceprint information and ensuring user information security.
[0112] Figure 4 It is a schematic structural diagram of an electronic device provided by an embodiment of this specification. The following refers to Figure 4 to describe the electronic device 400 according to this embodiment of the present invention. Figure 4 The shown electronic device 400 is only an example and should not bring any limitation to the functions and usage scope of the embodiments of the present invention.
[0113] Such as Figure 4As shown, the electronic device 400 is presented in the form of a general-purpose computing device. The components of the electronic device 400 may include, but are not limited to: at least one processing unit 410, at least one storage unit 420, a bus 430 connecting different system components (including the storage unit 420 and the processing unit 410), a display unit 440, etc.
[0114] Among them, the storage unit stores program code, and the program code can be executed by the processing unit 410, so that the processing unit 410 executes the steps according to various exemplary embodiments of the present invention described in the above processing method part of this specification. For example, the processing unit 410 can execute as Figure 1 shown in the steps.
[0115] The storage unit 420 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 4201 and / or a cache storage unit 4202, and may further include a read-only storage unit (ROM) 4203.
[0116] The storage unit 420 may also include a program / utility 4204 having a set (at least one) of program modules 4205. Such program modules 4205 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment.
[0117] The bus 430 may represent one or more of several types of bus structures, including a storage unit bus or a storage unit controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any bus structure in a variety of bus structures.
[0118] The electronic device 400 can also communicate with one or more external devices 500 (such as a keyboard, a pointing device, a Bluetooth device, etc.), can also communicate with one or more devices that enable a user to interact with the electronic device 400, and / or communicate with any device that enables the electronic device 400 to communicate with one or more other computing devices (such as a router, a modem, etc.). Such communication can be carried out through an input / output (I / O) interface 450. And, the electronic device 400 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter 460. The network adapter 460 can communicate with other modules of the electronic device 400 through the bus 430. It should be understood that although Figure 4is not shown in the figure, and other hardware and / or software modules can be used in combination with the electronic device 400, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0119] Through the description of the above embodiments, those skilled in the art can easily understand that the exemplary embodiments described in the present invention can be implemented by software, or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present invention can be embodied in the form of a software product, which can be stored in a computer-readable storage medium (which can be a CD-ROM, USB flash drive, mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, server, or network device, etc.) to execute the above method according to the present invention. When the computer program is executed by a data processing device, the computer-readable medium can implement the above method of the present invention, that is: as Figure 1 the method shown.
[0120] Figure 5 is a schematic diagram of the principle of a computer-readable medium provided by an embodiment of this specification.
[0121] Implement Figure 1 The computer program for implementing the method shown can be stored on one or more computer-readable media. The computer-readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0122] The computer-readable storage medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries the readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable storage medium can also be any readable medium other than the readable storage medium, which can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the readable storage medium can be transmitted by any appropriate medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination of the above.
[0123] The program code for performing the operations of the present invention can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, executed as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., by using an Internet service provider to connect through the Internet).
[0124] In summary, the present invention can be implemented in hardware, or in software modules running on one or more processors, or in a combination thereof. Those skilled in the art should understand that general-purpose data processing devices such as microprocessors or digital signal processors (DSPs) can be used in practice to implement some or all of the functions of some or all of the components in accordance with some embodiments of the present invention. The present invention can also be implemented as a device or apparatus program (e.g., a computer program and a computer program product) for performing some or all of the methods described herein. Such a program implementing the present invention can be stored on a computer-readable medium, or can be in the form of one or more signals. Such signals can be downloaded from an Internet website, or provided on a carrier signal, or provided in any other form.
[0125] The specific embodiments described above have further elaborated on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the present invention is not inherently related to any specific computer, virtual device, or electronic device, and various general-purpose devices can also implement the present invention. The above are only specific embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
[0126] Each embodiment in this specification is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and the key points of each embodiment are to illustrate the differences from other embodiments.
[0127] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.
Claims
1. An audio information processing method, characterized in that, Including: Obtain sample audio information, and jointly train a pre-trained audio content extractor, a voiceprint feature extractor, and an audio synthesizer according to the sample audio information, where: The audio content extracted by the audio content extractor includes text information, pauses, volume, tone, and speaking breath; The voiceprint features extracted by the voiceprint feature extractor include timbre features and accent features; Obtain the audio information to be desensitized and the target voiceprint audio information, use the audio content extractor to extract the audio content of the audio information to be desensitized, and use the voiceprint feature extractor to extract the target voiceprint features of the target voiceprint audio information; Input the audio content and the target voiceprint features into the audio synthesizer, and the audio synthesizer predicts the audio points of the audio content according to the frequency domain distribution of the target voiceprint features, and synthesizes the voiceprint desensitized audio information according to the audio points.
2. The method according to claim 1, characterized in that The joint training of the pre-trained audio content extractor, voiceprint feature extractor, and audio synthesizer according to the sample audio information includes: Use the audio content extractor to extract the audio content from the sample audio information; Use the voiceprint feature extractor to extract the voiceprint features from the sample audio information; The audio synthesizer predicts audio points according to the audio content and the voiceprint features; Calculate the deviation value according to the predicted audio points and the sample audio information, and adjust the parameters of the audio content extractor, the voiceprint feature extractor, and the audio synthesizer according to the deviation value, and perform iterative training until the calculated deviation value is less than the threshold.
3. The method according to claim 2, wherein The calculating the deviation value according to the predicted audio points and the sample audio information includes: Sample from the waveform of the sample audio information, calculate the difference between each sampling point and the predicted audio point, and sum all the differences to obtain the deviation value.
4. The method according to claim 2, characterized in that, The method further includes: Set an audio label for the sample audio information according to the audio content in the sample audio information, and train the audio content extractor based on the audio label; Set different voiceprint labels according to different users to whom the different sample audio information belongs, and train the voiceprint feature extractor based on the voiceprint labels.
5. The method according to any one of claims 1-4, characterized in that, Also including: Deploy the pre-trained audio content extractor, voiceprint feature extractor, and audio synthesizer to a network access device; The network access device obtains the audio information to be desensitized sent by the terminal, desensitizes it using the deployed model, and uploads the voiceprint desensitized audio information to the network after synthesizing the voiceprint desensitized audio information.
6. An audio information processing device, characterized in that, Including: A joint training module, which obtains sample audio information and jointly trains a pre-trained audio content extractor, a voiceprint feature extractor, and an audio synthesizer according to the sample audio information, where: the audio content extracted by the audio content extractor includes text information, pauses, volume, tone, and speaking breath; the voiceprint features extracted by the voiceprint feature extractor include timbre features and accent features; A feature extraction module that obtains the audio information to be desensitized and the target voiceprint audio information, and is used to extract the audio content of the audio information to be desensitized by using the audio content extractor, and extract the target voiceprint features in the target voiceprint audio information by using the voiceprint feature extractor; An audio synthesis module that inputs the audio content and the target voiceprint features into the audio synthesizer. The audio synthesizer predicts the audio points of the audio content according to the frequency domain distribution of the target voiceprint features, and synthesizes the voiceprint desensitized audio information according to the audio points.
7. The device according to claim 6, characterized in that, The joint training module is specifically used for: Using the audio content extractor to extract the audio content from the sample audio information; Using the voiceprint feature extractor to extract the voiceprint features from the sample audio information; The audio synthesizer predicts the audio points according to the audio content and the voiceprint features; Calculating the deviation value according to the predicted audio points and the sample audio information, and adjusting the parameters of the audio content extractor, the voiceprint feature extractor and the audio synthesizer according to the deviation value, and performing iterative training until the calculated deviation value is less than the threshold.
8. An electronic device, wherein, The electronic device includes: A processor; and, A memory storing a computer-executable program, where the executable program, when executed, causes the processor to execute the method according to any one of claims 1-5.
9. A computer-readable storage medium, wherein, The computer-readable storage medium stores one or more programs, and when the one or more programs are executed by a processor, the method according to any one of claims 1-5 is implemented.
Citation Information
Patent Citations
Speech synthesis method and device, electronic equipment and storage medium
CN112786012A
Dialogue audio data processing method, electronic equipment and computer readable storage medium
CN112966090A
Construction method of voice synthesizer, and voice synthesis method and device
CN113823257A