Voice processing method and device, equipment, medium and product

The pronunciation and text identification are extracted through audio encoding and text encoding models and fused into speech element data, solving the problem of single speech synthesis in the prior art and improving user experience.

CN120260540APending Publication Date: 2025-07-04YIDIAN LINGXI INFORMATION TECHNOLOGY (GUANGZHOU) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510245027.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing voice synthesis technology makes the user synthesized voice more single, affecting the user experience.

Method used

By obtaining audio data and text data, the audio encoding model extracts pronunciation identifiers, and the text encoding model extracts text identifiers, and inputs them into the speech conversion model to fuse them into pronunciation speech element data.

Benefits of technology

It realizes the synthesizing of voice without fixed sound lines or replicated authorized sound lines, improving the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120260540A_ABST
    Figure CN120260540A_ABST
Patent Text Reader

Abstract

The invention relates to a voice processing method and device, equipment, a medium and a product, and belongs to the technical field of voice processing, and the method comprises the steps: obtaining audio data and text data; inputting the audio data into a preset audio coding model to obtain a rhythm identifier; inputting the text data into a preset text coding model to obtain a text identifier; and inputting the rhythm identifier and the text identifier into a preset voice conversion model to obtain voice element data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of speech processing, and more particularly, to a speech processing method, apparatus, device, medium, and product. Background Art

[0002] With the rapid development of speech synthesis technology, speech synthesis technology has gradually evolved from mechanical "mechanical voice" to highly anthropomorphic "intelligent voice generation". Most existing speech synthesis schemes convert the text input by the user into a fixed voice line or reproduce an authorized voice line, which makes the synthesized speech obtained by the user relatively single and affects the user experience. Summary of the Invention

[0003] An object of an embodiment of the present disclosure is to provide a new technical solution for speech processing.

[0004] According to a first aspect of the present disclosure, there is provided a speech processing method, the method comprising:

[0005] Obtain audio data and text data;

[0006] Input the audio data into a preset audio encoding model to obtain a prosody identifier;

[0007] Input the text data into a preset text encoding model to obtain a text identifier;

[0008] Input the prosody identifier and the text identifier into a preset speech conversion model to obtain speech element data.

[0009] In a possible implementation, the audio encoding model includes a first encoding layer and a quantization layer, and the quantization layer is a vector quantization layer;

[0010] The inputting the audio data into a preset audio encoding model to obtain a prosody identifier includes:

[0011] Input the audio data into the first encoding layer of the preset audio encoding model to obtain a first encoding sequence;

[0012] Input the first encoding sequence into the quantization layer of the audio encoding model to obtain a prosody identifier.

[0013] In a possible implementation, the method further includes:

[0014] Obtain training samples;

[0015] Train the audio encoding model with the training samples to obtain a trained audio encoding model.

[0016] In a possible implementation, training the audio encoding model with the training samples to obtain a trained audio encoding model includes:

[0017] Inputting the training samples into the audio encoding model to obtain discrete identifiers;

[0018] Inputting the discrete identifiers into a preset audio decoding model to obtain posterior probabilities corresponding to the audio encoding model;

[0019] Adjusting the audio encoding model according to the posterior probabilities and the training samples.

[0020] In a possible implementation, the audio decoding model includes a second encoding layer and a decoding layer;

[0021] The step of inputting the discrete identifiers into a preset audio decoding model to obtain posterior probabilities corresponding to the audio encoding model includes:

[0022] Inputting the discrete identifiers into the second encoding layer of the preset audio decoding model to obtain a second encoding sequence;

[0023] Inputting the second encoding sequence into the decoding layer of the audio decoding model to obtain posterior probabilities corresponding to the audio encoding model.

[0024] In a possible implementation, the speech conversion model is an autoregressive conversion model.

[0025] According to a second aspect of the present disclosure, there is also provided a speech processing apparatus, including:

[0026] An acquisition module, configured to acquire audio data and text data;

[0027] A first obtaining module, configured to input the audio data into a preset audio encoding model to obtain prosody identifiers; wherein, the audio encoding model is additionally configured with a quantization layer;

[0028] A second obtaining module, configured to input the text data into a preset text encoding model to obtain text identifiers;

[0029] A third obtaining module, configured to input the prosody identifiers and the text identifiers into a preset speech conversion model to obtain speech feature data.

[0030] According to a third aspect of the present disclosure, there is also provided a computer system. The computer system includes a processor, and when the processor executes program instructions or code, the computer system implements the speech processing method in the first aspect. Exemplarily, the computer system further includes a memory, and the memory is used to store the program instructions or code.

[0031] According to a fourth aspect of the present disclosure, there is also provided a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the above-mentioned voice processing method when running.

[0032] According to a fifth aspect of the present disclosure, there is also provided a computer program product including a game program, which, when executed, causes a computer to execute the steps of the above-mentioned voice processing method.

[0033] According to a sixth aspect of the present disclosure, there is also provided an electronic device including a memory and a processor. A computer program is stored in the memory, and the processor is configured to execute the above-mentioned voice processing method through the computer program.

[0034] One beneficial effect of the embodiments of the present disclosure is that the voice processing method provided by the embodiments of the present disclosure can obtain audio data and text data, input the audio data into an audio encoding model to obtain a prosody identifier, input the text data into a text encoding model to obtain a text identifier, and input the prosody identifier and the text identifier into a preset voice conversion model to obtain voice element data. Furthermore, by extracting the prosody in the audio input by the user and fusing it with the text input by the user, voice element data with prosody can be obtained, thereby realizing the synthesis of voice without a fixed voice line or a replicated authorized voice line, effectively improving the user experience.

[0035] Through the following detailed description of the exemplary embodiments of the present specification with reference to the accompanying drawings, the features and advantages of the embodiments of the present specification will become clear. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] The drawings incorporated in the specification and constituting a part of the specification illustrate the embodiments of the present specification and, together with the description, are used to explain the principles of the embodiments of the present specification.

[0037] Figure 1 A schematic diagram of the hardware structure of an electronic device that can be used to implement the voice processing method according to the embodiments of the present disclosure is shown;

[0038] Figure 2 A schematic flowchart of the voice processing method according to some embodiments is shown;

[0039] Figure 3 A schematic flowchart framework of the voice processing according to some embodiments is shown;

[0040] Figure 4 A schematic diagram of the structure of an audio encoding model according to some embodiments is shown;

[0041] Figure 5Shows a schematic structural diagram of an audio encoding model and an audio decoding model according to some embodiments;

[0042] Figure 6 Shows a schematic structural diagram of a voice processing device according to some embodiments;

[0043] Figure 7 Shows a schematic hardware structure diagram of an electronic device according to some embodiments. Detailed implementation manners

[0044] Now, various exemplary embodiments of this specification will be described in detail with reference to the accompanying drawings.

[0045] The following description of at least one exemplary embodiment is merely illustrative in nature and is in no way a limitation on the embodiments of this specification or their application or use.

[0046] It should be noted that similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further discussed in subsequent drawings.

[0047] It should be noted that all actions of obtaining signals, information, or data in the embodiments of the present disclosure are carried out on the premise of complying with the corresponding data protection regulations and policies of the country where the location is located and obtaining authorization from the corresponding device owner.

[0048] The embodiments of the present disclosure provide a new voice processing solution, which allows a user to provide an audio segment and a text segment. Through a preset audio encoding model, the prosody identifier of the audio is extracted, and through a preset text encoding model, the text identifier of the text is extracted. Then, through a preset voice conversion model, the prosody identifier and the text identifier are fused into a voice with prosody and capable of reflecting the text content, thereby realizing the synthesis of voice without a fixed voice line or a replicated authorized voice line, effectively improving the user experience.

[0049] Figure 1 Shows a schematic hardware structure diagram of an electronic device that can be used to implement the voice processing method according to the embodiments of the present disclosure.

[0050] The electronic device 1000 is a device capable of running graphic design software. The graphic design software can be a local application installed on the electronic device, or a web application, a lightweight application, or a mini-program, etc., which is not limited herein. The electronic device 1000 can be a mobile phone, a tablet computer, a PC, etc., which is not limited herein.

[0051] Such as Figure 1As shown, the electronic device 1000 may include a processor 1101, a memory 1102, an interface device 1103, a communication device 1104, an output device 1105, an input device 1106, and so on. Figure 1 The hardware configuration shown is merely illustrative and is in no way intended to limit the present disclosure, its application, or use.

[0052] The processor 1101 is used to execute a computer program, which can be written in an instruction set such as x86, Arm, RISC, MIPS, SSE, etc. The memory 1102 includes, for example, ROM (Read-Only Memory), RAM (Random Access Memory), non-volatile memory such as a hard disk, and the like. The interface device 1103 includes, for example, a USB interface, a network cable interface, a headphone interface, and the like. The communication device 1104 can perform wired or wireless communication, for example. The communication device 1104 may include at least one short-range communication module, for example, any module that performs short-range wireless communication based on short-range wireless communication protocols such as the Hilink protocol, WiFi (IEEE 802.11 protocol), Mesh, Bluetooth, ZigBee, Thread, Z-Wave, NFC, UWB, LiFi, etc. The communication device 1104 may also include a remote communication module, for example, any module that performs WLAN, GPRS, 2G / 3G / 4G / 5G remote communication. The output device 1105 may include, for example, a liquid crystal display screen or a touch display screen, a speaker, and the like. The input device 1106 may include, for example, a touch screen, a keyboard, a microphone, various sensors, and the like.

[0053] In this embodiment, the memory 1102 of the electronic device 1000 is used to store a computer program, which is used to control the processor 1101 to operate to execute the voice processing method according to any embodiment of the present disclosure.

[0054] Next, taking the electronic device 1000 as an example of the implementation subject, various embodiments of the voice processing method will be described. Figure 1 as the electronic device 1000, various embodiments of the voice processing method will be described.

[0055] <First Embodiment>

[0056] Figure 2 A voice processing method according to some embodiments is shown. The voice processing method may include the following steps S210 to step S240:

[0057] Step S210, obtain audio data and text data.

[0058] In this embodiment, the audio data may be any piece of audio, may be the audio in a video file, may be a given Mel spectrogram, or may be a string representing audio, and is not limited herein.

[0059] In this embodiment, the reference text includes text and a string reflecting the text. That is, the reference text can be any piece of text or a string representing a piece of text, and no limitation is imposed here.

[0060] Step S220: Input the audio data into a pre-set audio encoding model to obtain a prosody identifier.

[0061] In this embodiment, as Figure 3 shown, the audio encoding model can be a multilingual encoder (SpeechTokenizer), which is fine-tuned and optimized based on the Whisper Automatic Speech Recognition model (Whisper ASR model). The audio encoding model can encode the input audio data to obtain a prosody identifier, that is, prosody tokens. Here, the prosody tokens are the numbers corresponding to the basic units of the feature vectors of prosody expressions.

[0062] Step S230: Input the text data into a pre-set text encoding model to obtain a text identifier.

[0063] In this embodiment, as Figure 3 shown, the text encoding model can be a text encoder (TextTokenizer), which can encode the input text data to obtain a text identifier, that is, text tokens. Here, the text tokens are the numbers corresponding to the basic units of the feature vectors of text expressions.

[0064] Step S240: Input the prosody identifier and the text identifier into a pre-set speech conversion model to obtain speech element data.

[0065] In this embodiment, the speech element data is, for example, speech tokens, and the speech element data is data with prosody and capable of reflecting the text content of the text data.

[0066] In some embodiments, as Figure 3 shown, the speech conversion model can be an autoregressive conversion model (LLM-Autoregressive Transformer). When the speech conversion model is an autoregressive conversion model, through self-supervised learning of a large amount of text, high-quality speech generation, understanding, and reasoning capabilities are achieved, and thus the prosody identifier and the text identifier are converted into speech element data that meets the requirements.

[0067] The voice processing method according to the first embodiment of the present invention solves the problem in the prior art that the text input by the user is converted into a fixed voice line or a reproduced authorized voice line, which makes the synthesized voice obtained by the user relatively single. Based on this method, audio data and text data are acquired. The audio data is input into an audio encoding model to obtain a prosody identifier, and the text data is input into a text encoding model to obtain a text identifier. The prosody identifier and the text identifier are input into a preset voice conversion model, and voice element data can be obtained. Furthermore, by extracting the prosody in the audio input by the user and fusing it with the text input by the user, voice element data with prosody is obtained, thereby realizing the synthesis of voice without using a fixed voice line or a reproduced authorized voice line, effectively improving the user experience.

[0068] <Second Embodiment>

[0069] In this embodiment, in order to improve the accuracy of the prosody identifier output by the audio encoding model, the audio encoding model may include a first encoding layer and a quantization layer. The quantization layer is a vector quantization layer. Through the quantization layer, data can be compressed and feature discretization processing can be performed, thereby improving the accuracy of the prosody identifier output by the audio encoding model.

[0070] In these embodiments, compared with the above first embodiment, step S220 may include the following step S310 and step S320:

[0071] Step S310, input the audio data into the first encoding layer of a preset audio encoding model to obtain a first encoding sequence.

[0072] In this embodiment, as Figure 4 shown, the audio encoding model (Speech Tokenizer) includes a first encoding layer (Transformer Encoder1) and a quantization layer (Residual Vector Quantizer). After the audio data (Speech X) passes through rotary positional encoding (Rotary Positional Embedding), it is input into the first encoding layer (Transformer Encoder1), and a first encoding sequence can be obtained.

[0073] Step S320, input the first encoding sequence into the quantization layer of the audio encoding model to obtain a prosody identifier.

[0074] In this embodiment, as Figure 4 shown, the first encoding sequence is input into the quantization layer (Residual VectorQuantizer), and a prosody identifier can be obtained.

[0075] In this embodiment, a prosody emotion cloning technology through transfer learning of a multi-lingual speech encoder and a speech conversion model, combined with the Repetition Aware Sampling method, is used to achieve applications in prosody emotion expression and multiple languages.

[0076] In this embodiment, by accurately extracting prosody identifiers from audio data and migrating the prosody identifiers to speech element data, the emotion reflected by the audio data is maintained, effectively reducing the occurrence of repeated or distorted generated content. And through this audio encoding model, the subsequent speech conversion model can meet the requirements of high complexity and depth in understanding natural language and audio emotions.

[0077] <Third Embodiment>

[0078] In this embodiment, in order to improve the accuracy of the prosody identifiers output by the audio encoding model, the audio encoding model needs to be trained.

[0079] In these embodiments, compared with the above-mentioned first embodiment, the method further includes the following steps S410 and S420:

[0080] Step S410, obtain training samples.

[0081] In this embodiment, the training sample can be any piece of audio, which can be the audio in a video file, a given Mel spectrogram, or a string representing audio, and is not limited herein.

[0082] Step S420, train the audio encoding model with the training samples to obtain a trained audio encoding model.

[0083] In this embodiment, through training the audio encoding model, multi-lingual audio encoding quantization is achieved, shortening the cumbersome work such as text-to-phoneme conversion and forced alignment, effectively improving the convenience of expanding languages, and improving the accuracy of multi-lingual pronunciation.

[0084] <Fourth Embodiment>

[0085] In this embodiment, in order to improve the accuracy of the prosody identifiers output by the audio encoding model, an audio decoding model can be additionally configured. Through the audio encoding model and the audio decoding model, the training samples are converted into posterior probabilities.

[0086] In these embodiments, compared with the above-mentioned third embodiment, the step S420 may include the following steps S510 to S530:

[0087] Step S510, input the training samples into the audio encoding model to obtain discrete identifiers.

[0088] In this embodiment, if Figure 5 As shown, the training sample (Speech X') is input into the audio coding model (SpeechTokenizer) to obtain discrete tokens (Speech Tokens').

[0089] Step S520: input the discrete identifier into a preset audio decoding model to obtain a posterior probability corresponding to the audio coding model.

[0090] In this embodiment, if Figure 5 As shown, discrete tokens (Speech Tokens') are input into the audio decoding model to obtain the posterior probability (P(Y丨X)) corresponding to the audio encoding model.

[0091] Step S530: adjusting the audio coding model according to the posterior probability and the training samples.

[0092] In this embodiment, a first coding layer Transformer Encoder1 and a second coding layer Transformer Encoder2 are set, and a quantization layer (Residual Vector Quantizer) is inserted between the first coding layer Transformer Encoder1 and the second coding layer Transformer Encoder2. Given a training sample (Speech X') as input, it is then passed through rotation position encoding and the first coding layer Transformer Encoder1 to obtain a context-aware representation H, and then the quantization layer (Residual Vector Quantizer) is used to obtain a discrete tag Token. The discrete tag Token corresponds to the vector embedding of the quantization layer Codebook, and then passes through the second coding layer Transformer Encoder2 and the ASR decoder to predict the posterior probability of the text mark corresponding to the audio. The audio coding model can be adjusted through the posterior probability and training samples.

[0093] <Fifth Embodiment>

[0094] In this embodiment, in order to obtain the posterior probability corresponding to the audio encoding model through the audio decoding model, the audio decoding model may include a second encoding layer and a decoding layer.

[0095] In these embodiments, relative to the fourth embodiment, after step S230, the method further includes the following steps S610 and S620:

[0096] Step S610: input the discrete identifier into the second coding layer of the preset audio decoding model to obtain a second coding sequence.

[0097] In this embodiment, as Figure 5 shown, discrete identifiers (Speech Tokens’) are input into the second encoding layer (Transformer Encoder2) to obtain a second encoded sequence.

[0098] Step S620: Input the second encoded sequence into the decoding layer of the audio decoding model to obtain the posterior probability corresponding to the audio encoding model.

[0099] In this embodiment, the second encoded sequence is input into the decoding layer (Transformer Decoder) to obtain the posterior probability corresponding to the audio encoding model.

[0100] <Device Embodiment>

[0101] Figure 6 shows a schematic structural diagram of a speech processing device according to an embodiment of the present disclosure. As Figure 6 shown, the speech processing device 300 includes an acquisition module 310, a first obtaining module 320, a second obtaining module 330, and a third obtaining module 340.

[0102] The acquisition module 310 is configured to acquire audio data and text data;

[0103] The first obtaining module 320 is configured to input the audio data into a preset audio encoding model to obtain prosody identifiers; wherein, the audio encoding model is additionally configured with a quantization layer;

[0104] The second obtaining module 330 is configured to input the text data into a preset text encoding model to obtain text identifiers;

[0105] The third obtaining module 340 is configured to input the prosody identifiers and text identifiers into a preset speech conversion model to obtain speech element data.

[0106] In some embodiments, the first obtaining module 320 is further configured to input the audio data into the first encoding layer of a preset audio encoding model to obtain a first encoded sequence; and input the first encoded sequence into the quantization layer of the audio encoding model to obtain prosody identifiers.

[0107] In some embodiments, the speech processing device 300 further includes a training module, and the training module is configured to acquire training samples; and train the audio encoding model through the training samples to obtain a trained audio encoding model.

[0108] In some embodiments, the training module is further configured to input the training samples into the audio encoding model to obtain discrete identifiers; input the discrete identifiers into a preset audio decoding model to obtain the posterior probability corresponding to the audio encoding model; and adjust the audio encoding model according to the posterior probability and the training samples.

[0109] In some embodiments, the training module is further configured to input the discrete identifiers into the second encoding layer of the preset audio decoding model to obtain a second encoding sequence; and input the second encoding sequence into the decoding layer of the audio decoding model to obtain the posterior probability corresponding to the audio encoding model.

[0110] <Device Embodiments>

[0111] Figure 7 FIG. shows a schematic hardware structure diagram of an electronic device according to some other embodiments. As Figure 7 shown, the electronic device 600 includes a processor 610 and a memory 620. The memory 620 is used to store a computer program, and the computer program is used to control the processor 610 to operate so as to control the electronic device 600 to execute the speech processing method according to any embodiment of the present disclosure.

[0112] Embodiments of the present disclosure also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the speech processing method according to any embodiment of the present disclosure.

[0113] Embodiments of the present disclosure also provide a computer program product, which includes a computer program or instruction that, when executed by a processor, implements the speech processing method according to any embodiment of the present disclosure.

[0114] The various embodiments in this specification are all described in a progressive manner. The same or similar parts among the various embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device and equipment embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.

[0115] The above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than in the embodiments and still achieve the desired results. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired results. In certain implementations, multitasking and parallel processing are also possible or may be advantageous.

[0116] Embodiments of this specification may be devices, methods, and / or computer program products. A computer program product may include a computer-readable storage medium having computer-readable program instructions thereon for causing a processor to implement various aspects of the embodiments of this specification.

[0117] A computer-readable storage medium may be a tangible device that can retain and store instructions for use by an instruction execution device. A computer-readable storage medium may be, for example, but is not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium would include: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device such as a punch card or raised structures in grooves having instructions stored thereon, and any suitable combination of the foregoing. The computer-readable storage medium as used herein is not to be construed as a transitory signal per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission medium (e.g., optical pulses through an optical fiber cable), or electrical signals transmitted through a wire.

[0118] The computer-readable program instructions described herein may be downloaded from a computer-readable storage medium to respective computing / processing devices, or may be downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, optical fiber transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.

[0119] The computer program instructions for performing the operations of the embodiments of this specification may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine - related instructions, microcode, firmware instructions, state - setting data, or source code or object code written in any combination of one or more programming languages. The programming languages include object - oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer - readable program instructions may be executed entirely on the first user computer, partially on the first user computer, executed as a stand - alone software package, partially on the first user computer and partially on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the first user computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., by using an Internet service provider to connect through the Internet). In some embodiments, by using the state information of the computer - readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field - programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer - readable program instructions to implement various aspects of the embodiments of this specification.

[0120] Aspects of the embodiments of this specification are described herein with reference to the flowcharts and / or block diagrams of methods, apparatus (devices), and computer program products according to the embodiments of this specification. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer - readable program instructions.

[0121] These computer - readable program instructions can be provided to a processor of a general - purpose computer, a special - purpose computer, or other programmable data - processing apparatus to produce a machine such that, when the instructions are executed by the processor of the computer or other programmable data - processing apparatus, a device is produced that implements the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams. These computer - readable program instructions can also be stored in a computer - readable storage medium, which causes a computer, a programmable data - processing apparatus, and / or other devices to work in a particular manner. Thus, the computer - readable medium storing the instructions includes a manufacture that includes instructions for implementing various aspects of the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams.

[0122] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other devices, causing a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other devices to generate a computer-implemented process such that the instructions executed on the computer, other programmable data processing apparatus, or other devices implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.

[0123] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present specification. In this regard, each box in the flowchart or block diagram may represent a module, a segment of a program, or a part of an instruction, and the module, segment of a program, or part of an instruction contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the boxes may occur in a different order than noted in the figures. For example, two consecutive boxes may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each box in the block diagrams and / or flowcharts, and combinations of boxes in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or acts, or by a combination of dedicated hardware and computer instructions. It is well known to those skilled in the art that implementation by hardware, implementation by software, and implementation by a combination of software and hardware are equivalent.

[0124] The embodiments of the present specification have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, the practical application, or the improvement of technologies in the market, or to enable other ordinary skilled persons in the art to understand the embodiments disclosed herein.

Claims

1. A voice processing method, the method comprising: Obtaining audio data and text data; Inputting the audio data into a preset audio encoding model to obtain a prosody identifier; Inputting the text data into a preset text encoding model to obtain a text identifier; Inputting the prosody identifier and the text identifier into a preset voice conversion model to obtain voice element data.

2. The method according to claim 1, wherein The audio encoding model includes a first encoding layer and a quantization layer, and the quantization layer is a vector quantization layer; The inputting the audio data into a preset audio encoding model to obtain a prosody identifier includes: Inputting the audio data into the first encoding layer of the preset audio encoding model to obtain a first encoding sequence; Inputting the first encoding sequence into the quantization layer of the audio encoding model to obtain a prosody identifier.

3. The method according to claim 1, wherein The method further includes: Obtaining a training sample; Training the audio encoding model with the training sample to obtain a trained audio encoding model.

4. The method according to claim 3, wherein The training the audio encoding model with the training sample to obtain a trained audio encoding model includes: Inputting the training sample into the audio encoding model to obtain a discrete identifier; Inputting the discrete identifier into a preset audio decoding model to obtain a posterior probability corresponding to the audio encoding model; Adjusting the audio encoding model according to the posterior probability and the training sample.

5. The method according to claim 4, wherein, The audio decoding model includes a second encoding layer and a decoding layer; The inputting the discrete identifier into a preset audio decoding model to obtain a posterior probability corresponding to the audio encoding model includes: Inputting the discrete identifier into the second encoding layer of the preset audio decoding model to obtain a second encoding sequence; Inputting the second encoding sequence into the decoding layer of the audio decoding model to obtain a posterior probability corresponding to the audio encoding model.

6. The method according to claim 1, wherein The voice conversion model is an autoregressive conversion model.

7. A voice processing device, wherein, The apparatus includes: An obtaining module, configured to obtain audio data and text data; A first obtaining module, configured to input the audio data into a preset audio encoding model to obtain a prosody identifier; wherein, the audio encoding model is additionally configured with a quantization layer; A second obtaining module, configured to input the text data into a preset text encoding model to obtain a text identifier; A third obtaining module, configured to input the prosody identifier and the text identifier into a preset voice conversion model to obtain voice element data.

8. An electronic device, wherein, Including a memory and a processor, the memory is used to store a computer program, and the processor is used to execute the method steps according to any one of claims 1 to 6 under the control of the computer program.

9. A computer-readable storage medium, the computer-readable storage medium including a stored computer program, wherein, When the computer program is run, it executes the method steps according to any one of claims 1 to 6.

10. A computer program product, characterized in that, Including a computer program, when the computer program is executed by a processor, it implements the method according to any one of claims 1 to 6.