Voice data processing method, device and equipment and computer readable storage medium

By determining the target timbre from multiple candidate timbres and performing encoding and feature transformation processing, the problem of inconsistent timbres in speech synthesis is solved, personalized timbre speech data processing is realized, and the naturalness and diversity of speech synthesis are improved.

CN119851676BActive Publication Date: 2025-12-12UBTECH ROBOTICS CORP LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411993615.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-12-12
Estimated Expiration
2044-12-30

AI Technical Summary

Technical Problem

In existing technologies, due to insufficient data and language differences, the timbre lacks richness and consistency during speech synthesis, making it impossible to achieve personalized timbre uniformity. This is especially true in multilingual scenarios where the speech synthesis effect is unnatural.

Method used

The target timbre is determined from multiple candidate timbres, and audio data and text are obtained and encoded. Language conversion markers are added, and speech data of the target timbre is generated using feature conversion and speech conversion techniques.

Benefits of technology

It achieves personalized timbre unification for speech data from different languages, improves the naturalness and timbre diversity of speech synthesis, and adapts to the timbre preferences of different users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119851676B_ABST
    Figure CN119851676B_ABST
Patent Text Reader

Abstract

The application provides a voice data processing method and device, equipment and a computer readable storage medium. The method comprises: determining target audio data corresponding to a target voice tone from candidate audio data corresponding to N candidate voice tones; obtaining first audio data, first text corresponding to the first audio data, and performing encoding processing on the first audio data and the first text to obtain an initial mark sequence; when the first text comprises at least two languages, adding a language conversion mark to the initial mark sequence to obtain a first target mark sequence; performing feature conversion on the first target mark sequence based on the target audio data to obtain first audio features of the first target mark sequence; and performing voice conversion on the first audio features to obtain second audio data corresponding to the first text, wherein the voice tone of the second audio data is the target voice tone. Through the application, the voice tone diversity of voice data can be improved, and the voice tone of voice data in different languages can be unified.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a voice data processing method, apparatus, device, and computer-readable storage medium. Background Technology

[0002] In the field of speech synthesis, generating natural-sounding speech typically requires voice samples from multiple speakers. Voice unification technology for mixed speech can help synthesize more natural and consistent speech; voice unification is a key step in achieving this goal. Voice unification technology can be used to create personalized voice services, such as customized voice assistants and game character voice-overs. In the medical field, voice unification technology can be used to assist in diagnosis by analyzing changes in the timbre of a patient's voice to diagnose certain diseases. In the cultural and entertainment industries such as film, television series, and animation, voice unification technology can be used to create more realistic and consistent speech effects.

[0003] In the process of speech synthesis, related technologies suffer from insufficient data to satisfy timbre preferences, resulting in a lack of timbre richness. Furthermore, differences in training data across different languages ​​lead to variations in the timbre of synthesized speech and unnatural transitions, making it impossible to achieve personalized timbre unification for speech data from different languages. Summary of the Invention

[0004] This application provides a voice data processing method, apparatus, device, and computer-readable storage medium, which can improve the timbre diversity of voice data and perform personalized timbre unification for voice data of different languages.

[0005] The technical solution of this application embodiment is implemented as follows:

[0006] This application provides a voice data processing method, the method comprising:

[0007] From the candidate audio data corresponding to each of the N candidate timbres, determine the target audio data corresponding to the target timbre;

[0008] Obtain first audio data and the first text corresponding to the first audio data, and encode the first audio data and the first text to obtain an initial tag sequence;

[0009] When the first text contains at least two languages, language conversion markers are added to the initial marker sequence to obtain the first target marker sequence;

[0010] Based on the target audio data, feature transformation is performed on the first target marker sequence to obtain the first audio feature of the first target marker sequence;

[0011] The first audio feature is converted into speech to obtain the second audio data corresponding to the first text, and the timbre of the second audio data is the target timbre.

[0012] This application provides a voice data processing device, including:

[0013] The timbre determination module is used to determine the target audio data corresponding to the target timbre from the candidate audio data corresponding to each of the N candidate timbres.

[0014] The first tagging module is used to acquire first audio data and first text corresponding to the first audio data, and to encode the first audio data and the first text to obtain an initial tagging sequence;

[0015] The second tagging module is used to add language conversion tags to the initial tagging sequence when the first text includes at least two languages, to obtain a first target tagging sequence;

[0016] The feature conversion module is used to perform feature conversion on the first target marker sequence based on the target audio data to obtain the first audio feature of the first target marker sequence;

[0017] The speech conversion module is used to convert the first audio feature into speech to obtain the second audio data corresponding to the first text, wherein the timbre of the second audio data is the target timbre.

[0018] This application provides an electronic device, the electronic device comprising:

[0019] Memory is used to store executable instructions or computer programs.

[0020] The processor, when executing computer-executable instructions or computer programs stored in the memory, implements the voice data processing method provided in the embodiments of this application.

[0021] This application provides a computer-readable storage medium storing computer-executable instructions or computer programs, which, when executed by a processor, implement the voice data processing method provided in this application.

[0022] This application provides a computer program product, including computer-executable instructions or a computer program, which, when executed by a processor, implements the voice data processing method provided in this application.

[0023] The embodiments of this application have the following beneficial effects:

[0024] By applying the embodiments of this application, target audio data corresponding to a target timbre is determined from candidate audio data corresponding to N candidate timbres. This achieves the determination of personalized timbre audio data from audio data of multiple timbres. Then, first audio data and the first text corresponding to the first audio data are obtained, and the first audio data and the first text are encoded to obtain an initial marker sequence. When the first text includes at least two languages, language conversion markers are added to the initial marker sequence to obtain a first target marker sequence, which can convert different languages ​​in the text. Based on the target audio data, feature conversion is performed on the first target marker sequence to obtain the first audio feature of the first target marker sequence. Then, speech conversion is performed on the first audio feature to obtain the second audio data corresponding to the first text. The timbre of the second audio data is the target timbre. In this way, speech conversion is performed on other audio data based on the personalized timbre audio data, realizing personalized timbre unification of speech data of different languages. Attached Figure Description

[0025] Figure 1 This is a schematic diagram illustrating the application mode of the voice data processing method provided in the embodiments of this application;

[0026] Figure 2 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application;

[0027] Figure 3A This is a first flowchart illustrating the voice data processing method provided in the embodiments of this application;

[0028] Figure 3B This is a schematic diagram of the second process of the voice data processing method provided in the embodiments of this application;

[0029] Figure 3C This is a schematic diagram of the third process of the voice data processing method provided in the embodiments of this application;

[0030] Figure 3D This is a schematic diagram of the fourth process of the voice data processing method provided in the embodiments of this application;

[0031] Figure 4 This is a schematic diagram of the voice data processing flow provided in the embodiments of this application;

[0032] Figure 5 This is a schematic diagram of the Mel spectrum diffusion process provided in the embodiments of this application;

[0033] Figure 6 This is a schematic diagram of the speech tagging process provided in an embodiment of this application;

[0034] Figure 7 This is a schematic diagram of the structure of an input sequence of the same language provided in an embodiment of this application;

[0035] Figure 8 This is a schematic diagram of the structure of input sequences in different languages ​​provided in the embodiments of this application.

[0036] It should be noted that the terms "first" and "second" mentioned above are only used to distinguish between different options and do not represent the degree of superiority or inferiority of the options or their priority in the implementation process. Detailed Implementation

[0037] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0038] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0039] In the following description, the terms "first, second, third" are used merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first, second, third" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.

[0040] In the implementation of this application, the collection and processing of relevant data should strictly comply with the requirements of relevant laws and regulations, obtain the informed consent or separate consent of the personal information subject, and carry out subsequent data use and processing within the scope of laws and regulations and the authorization of the personal information subject.

[0041] In this application embodiment, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.

[0042] Unless otherwise defined, all technical and scientific terms used in the embodiments of this application have the same meaning as commonly understood by one of ordinary skill in the art. The terminology used in the embodiments of this application is for the purpose of describing the embodiments of this application only and is not intended to limit this application.

[0043] Before providing a further detailed description of the embodiments of this application, the nouns and terms involved in the embodiments of this application will be explained, and the nouns and terms involved in the embodiments of this application shall be interpreted as follows.

[0044] 1) Position encoding: is a technique that introduces position information of sequence elements into a sequence model.

[0045] 2) Mel spectrogram: A visualization tool used to represent the spectral characteristics of speech signals, which converts the time and frequency dimensions of the speech signal into a two-dimensional image.

[0046] 3) Timbre: refers to the quality or characteristics of a sound, which is determined by a variety of factors such as the composition of the sound spectrum, the shape of the waveform, and the harmonic structure.

[0047] This application provides a voice data processing method, apparatus, device, and computer-readable storage medium, which can improve the timbre diversity of voice data and perform personalized timbre unification for voice data of different languages.

[0048] The following describes exemplary applications of the electronic devices provided in the embodiments of this application. These electronic devices can be implemented as various types of terminals such as laptops, tablets, desktop computers, set-top boxes, smartphones, smart speakers, smartwatches, smart TVs, and in-vehicle terminals, or as servers. The following will describe exemplary applications when the device is implemented as a server.

[0049] See Figure 1 , Figure 1 This is a schematic diagram illustrating the application mode of the voice data processing method provided in the embodiments of this application, for example. Figure 1 The system involves server 200, network 300, and terminal 400. Terminal 400 is connected to server 200 through network 300, which can be a wide area network, a local area network, or a combination of both.

[0050] In this embodiment, terminal 400 is used as an intelligent question-answering robot for illustration. When terminal 400 performs voice interaction, in response to the received audio selection operation, terminal 400 determines the target audio data corresponding to the target timbre from the candidate audio data corresponding to each of the N candidate timbres. Terminal 400 collects the voice data of the dialogue partner and performs speech recognition on the voice data to obtain the question text. Terminal 400 sends the question text to server 200. Server 200 can determine the response text based on the question text and send the response text to terminal 400. Terminal 400 receives the response text and identifies it as the first text. It obtains the first audio data corresponding to the first text and encodes the first audio data and the first text to obtain an initial tag sequence. When the first text includes at least two languages, language conversion tags are added to the initial tag sequence to obtain a first target tag sequence. Based on the target audio data, feature conversion is performed on the first target tag sequence to obtain the first audio feature of the first target tag sequence. Speech conversion is performed on the first audio feature to obtain the second audio data corresponding to the first text. The timbre of the second audio data is the target timbre. The second audio data is the speech synthesis data, which is then output as speech synthesis data.

[0051] In some embodiments, when the voice data processing method provided in this application is implemented by the server 200, after the terminal 400 determines the target audio data, it can send the target audio data or the data identifier corresponding to the target audio data to the server 200 so that the server 200 can perform subsequent timbre unification processing based on the target audio data. After acquiring the voice data of the dialogue partner, terminal 400 can send the voice data to server 200. Server 200 performs speech recognition on the voice data to obtain the question text, then determines the response text based on the question text, and designates the response text as the first text. It then obtains the first audio data corresponding to the first text and encodes the first audio data and the first text to obtain an initial tag sequence. When the first text includes at least two languages, language conversion tags are added to the initial tag sequence to obtain a first target tag sequence. Based on the target audio data, feature conversion is performed on the first target tag sequence to obtain the first audio feature of the first target tag sequence. Speech conversion is performed on the first audio feature to obtain the second audio data corresponding to the first text. The timbre of the second audio data is the target timbre. The second audio data is the speech synthesis data. Server 200 sends the speech synthesis data to terminal 400, and terminal 400 outputs the speech synthesis data.

[0052] In some embodiments, the server (e.g., server 200) can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. The terminal 400 can be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, vehicle terminal, aircraft, intelligent robot, etc., but is not limited to these. The terminal and server can be directly or indirectly connected via wired or wireless communication, which is not limited in this embodiment.

[0053] See Figure 2 , Figure 2 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device may be a terminal or a server. Figure 2 The illustrated electronic device 500 includes at least one processor 410, a memory 450, and at least one network interface 420. The various components in the electronic device 500 are coupled together via a bus system 440. It is understood that the bus system 440 is used to implement communication between these components. In addition to a data bus, the bus system 440 also includes a power bus, a control bus, and a status signal bus. However, for clarity, ... Figure 2 The general labeled all buses as Bus System 440.

[0054] The processor 410 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc.

[0055] The memory 450 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state storage, hard disk drives, optical disk drives, etc. The memory 450 may optionally include one or more storage devices physically located away from the processor 410.

[0056] The memory 450 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), and the volatile memory may be random access memory (RAM). The memory 450 described in this application embodiment is intended to include any suitable type of memory.

[0057] In some embodiments, memory 450 is capable of storing data to support various operations, examples of which include programs, modules, and data structures or subsets or supersets thereof, as illustrated below.

[0058] Operating system 451 includes system programs for handling various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, and driver layer, for implementing various basic business functions and handling hardware-based tasks.

[0059] The network communication module 452 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 420, exemplary network interfaces 420 including Bluetooth, WiFi, and Universal Serial Bus (USB).

[0060] In some embodiments, the apparatus provided in this application can be implemented in software. Figure 2 A speech data processing device 455 stored in memory 450 is shown. It may be software in the form of programs and plug-ins, including the following software modules: timbre determination module 4551, first tag module 4552, second tag module 4553, feature conversion module 4554, and speech conversion module 4555. These modules are logically related and can therefore be arbitrarily combined or further split according to the functions they implement. The functions of each module will be described below.

[0061] The voice data processing method provided in this application will be described in conjunction with exemplary applications and implementations of the server device provided in the embodiments of this application.

[0062] The following describes the voice data processing method provided in the embodiments of this application. For example, to facilitate understanding of the voice data processing method provided in the embodiments of this application, the following describes the application scenarios of the voice data processing method provided in the embodiments of this application. When a robot performs voice interaction, it needs to convert text information into natural and fluent voice output so that the robot can communicate with the user in a way that is familiar to humans. The voice data processing method provided in the embodiments of this application can improve the timbre diversity of voice data and perform personalized timbre unification for voice data of different languages.

[0063] As mentioned above, the electronic device implementing the voice data processing method of the embodiments of this application can be a terminal, a server, or a combination of both. The following description uses an electronic device as a server as an example to illustrate the voice data processing method provided in the embodiments of this application. See also... Figure 3A , Figure 3A This is a first flowchart illustrating the voice data processing method provided in this application embodiment, which will be combined with... Figure 3A The steps shown are explained.

[0064] In step 301, the target audio data corresponding to the target timbre is determined from the candidate audio data corresponding to each of the N candidate timbres.

[0065] Here, each candidate timbre corresponds to a candidate audio data, and the target timbre refers to the timbre specified from the N candidate timbres.

[0066] In some embodiments, see Figure 3B , Figure 3B This is a schematic diagram of the second process of the voice data processing method provided in the embodiments of this application. Figure 3A Before step 301 shown, the following steps can also be performed: Figure 3B Steps 3001 to 3003 are explained in detail below.

[0067] In step 3001, the initial text is obtained, and control symbols are inserted and rewritten on the initial text to obtain candidate text.

[0068] Here, the initial text is used to determine candidate audio data, and can be obtained from the text dataset of the robot's historical voice interactions. The initial text can be processed by inserting control symbols and rewriting them to obtain candidate text. Control symbols are labels used to adjust the initial text; for example, control symbols are [laugh], [uv_break], etc., where [laugh] represents laughter and [uv_break] represents a pause. When control symbols are inserted into the initial text, for example, the initial text is "Hello, nice to meet you," and the text after control symbol insertion is "Hello [uv_break], nice to meet you [laugh]." In the speech synthesis result of the text after control symbol insertion, [laugh] will be replaced by laughter, and a pause will be added at [uv_break].

[0069] The system can insert control symbols or rewrite the initial text based on the number of characters in the initial text. When the number of characters in the initial text is greater than or equal to a preset threshold, control symbols are inserted. When the number of characters in the initial text is less than the preset threshold, the initial text is rewritten. When rewriting the initial text, the content is redefined, and the number of characters is increased or decreased. For example, if the initial text is "I want to sleep", the rewritten text would be "I just want to sleep".

[0070] In step 3002, N candidate timbre parameters are determined, and the N candidate timbre parameters are inserted into the candidate text respectively to obtain N target texts.

[0071] Here, candidate timbre parameters include the candidate timbre's spectrum, spatial coordinates, and timbre number. These parameters can also be understood as timbre seeds. Different candidate timbres correspond to different candidate timbre parameters. Cluster analysis is performed on different candidate timbres to obtain the category of each candidate timbre, and a candidate timbre parameter is assigned to each category. After determining N candidate timbre parameters, each parameter is inserted into the candidate text, resulting in N target texts. Each target text includes control symbols and the candidate timbre parameters.

[0072] In step 3003, speech conversion is performed on N target texts to obtain candidate audio data corresponding to each of the N candidate timbres.

[0073] Here, speech conversion is performed for each target text to obtain N candidate audio data, and the N candidate audio data correspond to N candidate timbres.

[0074] In some embodiments, speech conversion of N target texts to obtain candidate audio data corresponding to N candidate timbres can be achieved by performing the following steps: for each target text, feature conversion is performed on the target text to obtain a first audio tag sequence; feature extraction is performed on the first audio tag sequence to obtain a first audio feature; speech conversion is performed on the first audio feature to obtain candidate audio data.

[0075] Here, for each target text, feature transformation is performed on the target text based on the control symbols in the target text to obtain the first audio token sequence. For example, the target text includes the control symbol [empty_spk], which represents a pause, indicating that the target text contains pause features. Pause feature transformation is performed on the target text to generate a discrete sequence audio_token, which is the first audio token sequence.

[0076] Feature extraction is performed on the first audio tag sequence using an acoustic model to obtain Mel-spectral features, i.e., the first audio features. The acoustic model can be a Diffusion Variational Autoencoder (DVAE) model, which combines a variational autoencoder and a generative model based on diffusion processes. For each first audio tag sequence of the target text, different Mel-spectral features can be extracted. Using a vocoder model (Vocos), the different Mel-spectral features are converted into different waveforms, thus outputting candidate audio data corresponding to N candidate timbres.

[0077] In this embodiment, N candidate timbre parameters are inserted into candidate text to obtain N target texts, and speech conversion is performed on the N target texts to obtain candidate audio data corresponding to each of the N candidate timbres. This improves the timbre diversity of the speech data and provides a variety of personalized timbre options.

[0078] Continue to refer to Figure 3A In step 302, the first audio data and the first text corresponding to the first audio data are obtained, and the first audio data and the first text are encoded to obtain the initial marker sequence.

[0079] Here, the first audio data is the audio data to be converted in timbre, and the first text is the text corresponding to the first audio data. The first text can be the response text obtained by intelligent question-and-answer processing of the question text. The first audio data and the first text can be obtained from the robot's historical voice interaction dataset. The initial tag sequence includes the audio tag sequence corresponding to the first audio data and the text tag sequence corresponding to the first text.

[0080] In some embodiments, see Figure 3C , Figure 3C This is a schematic diagram of the third process of the voice data processing method provided in the embodiments of this application. Figure 3A In step 302 shown, "encoding the first audio data and the first text to obtain an initial marker sequence" can be achieved through... Figure 3C Steps 3021 to 3024 are implemented, and will be explained in detail below.

[0081] In step 3021, feature extraction is performed on the first audio data to obtain the first audio features, and position encoding is performed on the first audio features to obtain the audio representation vector.

[0082] Here, window functions such as Hamming windows or rectangular windows are used to divide the first audio data into a series of overlapping frames, resulting in multi-frame audio data. A Fast Fourier Transform is performed on each frame of audio data to convert the time-domain signal into a frequency-domain signal, obtaining the spectrum of each frame. The spectrum of each frame is then input into a visualization tool to create a two-dimensional image, resulting in a Mel spectrogram, which represents the first audio feature.

[0083] The first audio feature is positionally encoded using a first encoder. A fixed-size vector is assigned to each time frame of the first audio feature, representing the position of each time frame within the first audio feature. The vectors of each time frame are combined to form an audio representation vector. For example, the audio representation vector can be determined using the following formula (1):

[0084] H1 = Encoder1(posEnc(X)) (1)

[0085] Where H1 represents the audio representation vector, Encoder1 represents the first encoder, posEnc represents the positional encoding, and X represents the first audio feature (Mel spectrogram).

[0086] In step 3022, the audio representation vector is quantized to obtain the second audio tag sequence.

[0087] Here, the audio representation vector is quantized using a vector quantizer to obtain the speech markers for each time frame in the audio representation vector. The speech markers for each time frame are then combined into a second audio marker sequence. For example, the speech marker for the l-th time frame can be determined using the following formula (2):

[0088]

[0089] Where u represents the speech marker of the l-th time frame, c n Let C represent the speech vector of the l-th time frame in the audio representation vector, C represent the set of speech vectors, VQ represent the vector quantizer, and h represent the speech vector of the l-th time frame. l This represents the hidden layer representation of the l-th time frame.

[0090] In some embodiments, the speech vector of the l-th time frame can be updated using an exponential moving average method. For example, the updated speech vector of the l-th time frame can be determined using the following formula (3):

[0091] c ul =ac ul +(1-a)h l (3)

[0092] Among them, c ul This represents the updated speech vector of the l-th time frame, where a represents the attenuation coefficient, and h...l This represents the hidden layer representation of the l-th time frame.

[0093] The speech vector of each updated time frame is used as the quantized hidden layer representation, and combined with the positional encoding as the input to the second encoder. For example, the output of the second encoder can be determined by the following formula (4):

[0094] H3 = Encoder2(PosEnc(H2)) (4)

[0095] Where H3 represents the output of the second encoder, Encoder2 represents the second encoder, posEnc represents the position encoding, and H2 represents the quantized hidden layer representation.

[0096] In step 3023, the first text is encoded to obtain a text tag sequence.

[0097] Here, a byte pair encoding model is constructed based on the text dataset. The first text is then encoded using the byte pair encoding model to convert the character sequence in the first text into a tag sequence, thus obtaining the text tag sequence.

[0098] In step 3024, the second audio tag sequence and the text tag sequence are combined to obtain the initial tag sequence.

[0099] Here, the second audio tag sequence and the text tag sequence are combined to obtain the initial tag sequence, which is then used as the input sequence for the large language model.

[0100] In some embodiments, the second audio tag sequence and the text tag sequence are combined to obtain an initial tag sequence, which can be achieved by performing the following steps: concatenating the text tag sequence after the second audio tag sequence to obtain the initial tag sequence.

[0101] Here, the initial marker sequence includes a sequence start marker, a text marker sequence, a second audio marker sequence, additional information markers (including control symbols, etc.), and a sequence end marker. The sequence start marker is located at the beginning of the initial marker sequence, the second audio marker sequence is appended after the sequence start marker, the text marker sequence is appended after the second audio marker sequence, the additional information markers are appended after the text marker sequence, and the sequence end marker is located at the end of the initial marker sequence.

[0102] Continue to refer to Figure 3A In step 303, when the first text includes at least two languages, language conversion tags are added to the initial tag sequence to obtain the first target tag sequence.

[0103] Here, when the first text contains at least two languages, that is, when the first text contains different languages, the language conversion marker is added to the initial marker sequence to obtain the first target marker sequence.

[0104] In some embodiments, language conversion tags are added to the initial tag sequence to obtain a first target tag sequence, which can be achieved by performing the following steps: concatenating language conversion tags to the text tag sequence to obtain the first target tag sequence.

[0105] Here, the first target marker sequence includes a sequence start marker, a text marker sequence, a second audio marker sequence, a speech conversion marker, and a sequence end marker. The sequence start marker is located at the beginning of the first target marker sequence, the second audio marker sequence is concatenated after the sequence start marker, the text marker sequence is concatenated after the second audio marker sequence, the speech conversion marker is concatenated after the text marker sequence, and the sequence end marker is located at the end of the initial marker sequence.

[0106] In this embodiment of the application, when the first text includes at least two languages, a first target tag sequence is obtained by adding language conversion tags to the initial tag sequence. Based on the first target tag sequence, different languages ​​in the text can be converted, thereby achieving personalized timbre unification of speech data in different languages.

[0107] Continue to refer to Figure 3A In step 304, based on the target audio data, feature transformation is performed on the first target marker sequence to obtain the first audio feature of the first target marker sequence.

[0108] Here, the target audio data is encoded to obtain the target audio sequence. The optimal transmission path between the target audio sequence and the first target marker sequence is calculated using an optimal transmission condition flow matching model. Based on the optimal transmission path, the first target marker sequence is converted into a Mel spectrogram, i.e., the first audio feature.

[0109] Continue to refer to Figure 3A In step 305, the first audio feature is converted into speech to obtain the second audio data corresponding to the first text.

[0110] Here, the timbre of the second audio data is the target timbre.

[0111] In some embodiments, the process of converting the first audio features into speech to obtain the second audio data corresponding to the first text can be achieved by performing the following steps: obtaining a generative adversarial network model and calling the generative adversarial network model to extract waveforms from the first audio features to obtain a first audio waveform; determining audio data that matches the first audio waveform from an audio database, and identifying the matching audio data as the second audio data corresponding to the first text.

[0112] Here, the generative adversarial network model can be a vocoder network HiFiGAN. The generator in the generative adversarial network model is called to generate an audio waveform representing the first audio feature, i.e., the first audio waveform. After denoising and volume adjustment of the first audio waveform, it is matched with audio waveform data in the audio database to obtain matching audio data. The matching audio data is saved as a WAV file and identified as the second audio data corresponding to the first text.

[0113] In some embodiments, see Figure 3D , Figure 3D This is a schematic diagram of the fourth process of the voice data processing method provided in this application embodiment. When the first audio data and the target audio data are in the same language, it can also be processed through... Figure 3D Steps 310 to 314 are implemented, and will be explained in detail below.

[0114] In step 310, when the first audio data and the target audio data are in the same language, the target audio data is encoded to obtain the third audio tag sequence.

[0115] Here, when the first audio data and the target audio data are in the same language, the target audio data is divided into multiple time frames using a window function. The audio data of each time frame is processed by fast Fourier transform and Mel filtering to obtain Mel spectral features. The Mel spectral features are then converted into a series of audio vectors to obtain the third audio tag sequence.

[0116] In step 311, the prompt text corresponding to the target audio data is determined, and the prompt text is encoded to obtain the prompt mark sequence.

[0117] Here, speech recognition can be performed on the target audio data to obtain the corresponding prompt word text. Then, a pre-trained embedding model is used to encode the prompt word text corresponding to the target audio data, converting the prompt word text into a series of text vectors to obtain the prompt word tag sequence.

[0118] In step 312, the third audio tag sequence and the prompt word tag sequence are inserted into the initial tag sequence to obtain the second target tag sequence.

[0119] Here, the third audio tag sequence and the prompt word tag sequence are added to the initial tag sequence to obtain the second target tag sequence.

[0120] In some embodiments, the initial tag sequence includes a second audio tag sequence and a text tag sequence, with the second audio tag sequence preceding the text tag sequence. The second target tag sequence is obtained by inserting a third audio tag sequence and a cue word tag sequence into the initial tag sequence, which can be achieved by performing the following steps: inserting the cue word tag sequence after the second audio tag sequence and before the text tag sequence, and inserting the third audio tag sequence after the text tag sequence.

[0121] Here, the second target marker sequence includes a sequence start marker, a cue word marker sequence, a text marker sequence, a second audio marker sequence, a third audio marker sequence, an additional information marker, and a sequence end marker. The sequence start marker is located at the beginning of the second target marker sequence, the second audio marker sequence is appended after the sequence start marker, the cue word marker sequence is appended after the second audio marker sequence, the text marker sequence is appended after the cue word marker sequence, the additional information marker is appended after the text marker sequence, the third audio marker sequence is appended after the additional information marker, and the sequence end marker is located at the end of the second target marker sequence.

[0122] Continue to refer to Figure 3D In step 313, based on the target audio data, the second target marker sequence is subjected to feature transformation to obtain the second audio features of the second target marker sequence.

[0123] Here, the target audio data is encoded to obtain the target audio sequence. The optimal transmission path between the target audio sequence and the second target marker sequence is calculated using an optimal transmission condition flow matching model. Based on the optimal transmission path, the second target marker sequence is converted into a Mel spectrogram, which represents the second audio feature of the second target marker sequence.

[0124] In step 314, the second audio features are converted into speech to obtain the second audio data corresponding to the first text.

[0125] Here, a generative adversarial network model is invoked to extract the waveform of the second audio feature, resulting in the second audio waveform. Audio data matching the second audio waveform is then identified from the audio database, and this matching audio data is designated as the second audio data corresponding to the first text.

[0126] In this embodiment of the application, when the first audio data and the target audio data are in the same language, the third audio tag sequence and the prompt word tag sequence corresponding to the target audio data are inserted into the initial tag sequence to obtain the second target tag sequence. Based on the second target tag sequence, personalized timbre unification can be achieved for speech data in the same language.

[0127] In some embodiments, the map data processing method provided in this application can be applied to the field of robot voice interaction. From candidate audio data corresponding to N candidate timbres, target audio data corresponding to a target timbre is determined, realizing the determination of personalized timbre audio data from audio data with multiple timbres. Then, first audio data and the first text corresponding to the first audio data are obtained, and the first audio data and the first text are encoded to obtain an initial marker sequence. When the first text includes at least two languages, language conversion markers are added to the initial marker sequence to obtain a first target marker sequence, which can convert different languages ​​in the text. Then, based on the target audio data, feature conversion is performed on the first target marker sequence to obtain the first audio feature of the first target marker sequence. Furthermore, speech conversion is performed on the first audio feature to obtain the second audio data corresponding to the first text, where the timbre of the second audio data is the target timbre. Thus, by performing speech conversion on other audio data based on the personalized timbre audio data, the robot achieves personalized timbre uniformity for speech data of different languages.

[0128] The following will describe an exemplary application of the voice data processing method provided in the embodiments of this application in a robot voice interaction scenario.

[0129] With the development of robotics technology, the ability for robots to interact via voice has become a hot research and development topic, as it helps improve the usability and user experience of robots. In robot voice interaction systems, speech synthesis technology plays a crucial role. Speech synthesis technology aims to convert text information into natural and fluent speech, enabling robots to communicate with users in a way that is familiar to humans.

[0130] Early speech synthesis technologies primarily relied on rule-based synthesis and parametric synthesis. Rule-based synthesis synthesized speech using manually defined linguistic rules, but the resulting speech lacked naturalness and had limited adaptability. Parametric synthesis synthesized speech by modeling the parameters of the speech signal, which improved speech quality to some extent, but still suffered from limitations such as monotonous timbre and insufficient expressiveness. With the emergence of statistical parametric speech synthesis (SPS), which statistically analyzes and models large amounts of speech data, more natural and fluent speech synthesis can be achieved, but it still has limitations in handling complex text and emotional expression. In recent years, breakthroughs in speech synthesis technology have been driven by the development of deep learning. Speech synthesis based on deep neural networks can automatically learn the feature representations of speech, improving the naturalness and expressiveness of the synthesized speech. Essentially, it is a cascaded model composed of an acoustic model and a vocoder. For example, using deep learning models such as Recurrent Neural Networks (RNNs) and Long Short-Term Memory (LSTM) networks can better handle long sequences of speech data and train more coherent speech synthesis models. Furthermore, combining techniques such as Generative Adversarial Networks (GANs) as vocoders can further improve the quality and realism of synthesized speech.

[0131] The relevant technologies have the following problems in the speech synthesis process:

[0132] 1) Different users have different needs, such as different timbre preferences. Since the synthesized voice timbre is affected by the training data, namely the gender, timbre and speaking style of the training data, and there is also a problem of insufficient data to meet the timbre preferences.

[0133] 2) In robot voice interaction, multilingual support is also one of the important challenges faced by speech synthesis technology. The differences between training data of different languages ​​will lead to inconsistent timbre and unnatural transitions in the output of the trained speech synthesis model. As a result, the robot cannot communicate effectively with users from different regions. The consistency of timbre in mixed Chinese and English speech is also an urgent problem to be solved.

[0134] To address the limitations of training data and the lack of rich timbre in related technologies, which result in fewer timbres for robot speech synthesis, and the differences and lack of training data for mixed speech such as Chinese and English, leading to inconsistencies in the timbre of synthesized speech and unnatural transitions, this application proposes a speech data processing method that, compared to related technologies, includes the following improvements:

[0135] 1) By setting timbre seeds (the candidate timbre parameters mentioned above), the text-to-speech model chatTTS can produce diverse timbre data;

[0136] 2) Based on the personalized voice selected by the user, the speech is reproduced in rhythm or cloned across languages ​​through the speech synthesis model CosyVoice, thereby achieving the uniformity of personalized Chinese and English voice.

[0137] In this embodiment, by combining speech data with different timbres and a speech synthesis model, a variety of personalized timbres are obtained for users to choose from. Then, the personalized timbres are replicated or cloned across languages, and the resulting speech data is used in the speech synthesis model on the robot. This enables the robot to achieve a unified timbre of personalized Chinese and English voices during voice interaction, thereby improving the robot's voice interaction capabilities.

[0138] Example, reference Figure 4 , Figure 4 This is a schematic diagram of the voice data processing flow provided in an embodiment of this application. The electronic device implementing the voice data processing method of this embodiment can be a server, combined with... Figure 4 The steps shown illustrate the voice data processing flow provided in the embodiments of this application.

[0139] In step 401, it is determined whether to insert control symbols into the text.

[0140] You can choose whether to insert control symbols into the input text. Control symbols are used to adjust the content of the text. For example, the control symbols are [laugh] and [uv_break]. [laugh] represents laughter, and [uv_break] represents a pause. In the speech synthesis result of the text "Hello [uv_break], nice to meet you [la ugh]", [laugh] will be replaced by laughter, and a pause will be added at [uv_break].

[0141] When the input text is selected to have control symbols inserted, the Llama model is used to insert control symbols into the input text, such as [uv_break], [laugh], etc., resulting in text with control symbols. When the input text is selected not to have control symbols inserted, the process proceeds to step 402 to determine whether to redefine the text, that is, to regenerate the text content. For example, if the input text is "I want to sleep", the redefined text is "I just want to sleep".

[0142] In step 403, the text containing control symbols and timbre seed numbers is converted into a discrete sequence audio_token.

[0143] The Llama model generates a discrete sequence, audio_token, from text containing control symbols and timbre seed numbers. The Llama model determines whether to perform an operation based on the control symbols in the text. For example, if the control symbol [e mpty_spk] is inserted before or after the information representing speech in the text, it is equivalent to adding a pause symbol, indicating a pause in this speech information. The Llama model performs feature transformation on the text, turning the text content into a discrete sequence. The discrete sequence generated in this process is called audio_token.

[0144] In step 404, the discrete sequence audio_token is converted into Mel spectrum features.

[0145] The discrete sequence `audio_token` is converted into `mel_spec` using a diffusion variational encoder (DVAE) model, an acoustic model. The DVAE model combines a variational autoencoder (VAE) with a generative model of the diffusion process. By learning the latent distribution of speech data, it generates new speech data samples, including modeling the diffusion process and learning latent variables. This allows for the generation of speech data with various timbres. See the example in the reference. Figure 5 , Figure 5 This is a schematic diagram of the diffusion process of the Mel spectrum provided in an embodiment of this application. Mel spectrum X0 is obtained by diffusion sampling. t-1 Mel spectrum X t-1 Mel spectrum X was obtained after diffusion sampling. t .

[0146] Continue to refer to Figure 4 In step 405, the voice is output.

[0147] The Vocos vocoder model is used to perform speech conversion on the transformed Mel-spectrum features. Since the acoustic model generates multiple different Mel-spectrum features, the vocoder model then converts these features into waveforms, deriving speech with different timbres for the user to choose from. Personalized speech is customized based on the user's selected timbre, and speech data from other languages ​​can be replicated or cloned across languages ​​based on this personalized speech.

[0148] In step 406, voice tagging.

[0149] Example, reference Figure 6 , Figure 6This is a schematic diagram of the speech tagging process provided in the embodiments of this application. A vector quantization layer is inserted between the first encoder 601 and the second encoder 603 to extract Mel spectrogram features from the input speech, obtaining a Mel spectrogram. The Mel spectrogram is used as the input of the first encoder 601 and encoded in combination with position encoding to obtain a context-aware representation (the audio representation vector mentioned above). For example, the context-aware representation can be determined by the following formula (1):

[0150] H1 = Encoder1(posEnc(X)) (1)

[0151] Where H1 represents context-aware representation, Encoder1 represents the first encoder, posEnc represents position encoding, and X represents Mel spectrogram.

[0152] Discrete markers are obtained through vector quantizer 602, and combined with a set of audio representation vectors, the speech markers for the l-th time frame are obtained. For example, the speech markers for the l-th time frame can be determined by the following formula (2):

[0153]

[0154] Where u represents the speech marker of the l-th time frame, c n Let C represent the speech vector of the l-th time frame in the audio representation vector, C represent the set of speech vectors, VQ represent the vector quantizer, and h represent the speech vector of the l-th time frame. l This represents the hidden layer representation of the l-th time frame.

[0155] The speech vector will be updated during training. It can be updated using the exponential moving average method. For example, the updated speech vector of the l-th time frame can be determined by the following formula (3):

[0156] c ul =ac ul +(1-a)h l (3)

[0157] Among them, c ul This represents the updated speech vector of the l-th time frame, where a represents the attenuation coefficient, and h... l This represents the hidden layer representation of the l-th time frame.

[0158] The speech vector of each updated time frame is used as the quantized hidden layer representation H2={c u1 ,c u2 ,..}, and combined with position encoding as input to the second encoder, for example, the output of the second encoder can be determined by the following formula (4):

[0159] H3 = Encoder2(PosEnc(H2)) (4)

[0160] Where H3 represents the output of the second encoder, Encoder2 represents the second encoder, and posEnc represents the position encoder.

[0161] The output of the second encoder is used as the input to the Automatic Speech Recognition (ASR) decoder. In the ASR decoder 604, the output is the posterior probability P(Y|X) of the text label corresponding to the input speech. The posterior probability refers to the probability that a speech signal corresponds to a certain text label, that is, the probability of each text appearing in the content of the speech.

[0162] Continue to refer to Figure 4 In step 407, the text is encoded.

[0163] The input text is encoded using Byte Pair Encoding (BPE) to obtain a text vector.

[0164] In step 408, the sequence of text and speech tags is aligned.

[0165] Align each speech tag with its corresponding text vector to obtain a sequence of text and speech tags.

[0166] In step 409, the entire sequence of text and speech is learned.

[0167] The input sequence for constructing a large language model mainly consists of a sequence start marker, the extracted speech embedding vector, a text vector, and a sequence end marker. Since the text and speech markers are at different semantic levels, a start marker is inserted between them. During training, the original sequence is used as the expected output, and only the cross-entropy loss for the speech marker and the sequence end marker is considered during training, resulting in a sequence composed of speech markers and text vectors. For example, the cross-entropy loss can be determined using the following formula (5):

[0168]

[0169] Among them, L LM Denotes the cross-entropy loss, q(u l ) represents the posterior probability (predicted by the normalized softmax layer of the large language model).

[0170] In step 411, the markers in the input sequence are adjusted based on the speech of the reference timbre.

[0171] Using selected personalized speech as reference speech, when the reference speech and input speech are in the same language, the context learning zero-shot of the large language model can be utilized to replicate the timbre of the reference speech in a short sample. This process constructs the input sequence of the large language model according to actual needs, merging the reference speech token (Prompt SpeechToken, the third audio token sequence mentioned above) and the prompt speech text token (Prompt Text, the prompt word token sequence mentioned above) into a unified input sequence. The prompt speech token is pre-generated, and subsequent tokens are predicted by an autoregressive language model. For example, the reference... Figure 7 , Figure 7 This is a schematic diagram of the structure of an input sequence in the same language provided in an embodiment of this application. The input sequence in the same language includes input speech 701, prompt speech text 702, input text 703, and reference speech 704.

[0172] When the reference speech (the target audio data mentioned above) and the input speech (the first audio data mentioned above) are in different languages, i.e., the cross-language cloned sequence composition does not require the reference speech and prompt word text, but requires the insertion of a language conversion tag (LID). The remaining sequence structure is consistent with the input sequence in the same language. For example, the reference... Figure 8 , Figure 8 This is a schematic diagram of the structure of input sequences in different languages ​​provided in an embodiment of this application. The input sequences in different languages ​​include input speech 701, input text 703, and language conversion markers 705.

[0173] Continue to refer to Figure 4 In step 412, the speech tokens are converted into Mel spectrograms using a conditional flow matching model.

[0174] The principle of Conditional Flow Matching (OT-FCM) is to construct a probability density path from the prior distribution p(x) to the Mel spectrogram q(x) in a continuous-time normalized flow. The probability density path is represented by a time-dependent vector field. For example, the probability density path can be determined by the following formula (6):

[0175]

[0176] Where t ranges from 0 to 1, the initial value problem equation is solved. The Mel spectrum q(x) can be approximated by the prior distribution p(x) and sampled from it. The vector field is learned through the optimal transport condition flow. For example, the loss can be minimized by the following formulas (7) to (9):

[0177]

[0178] Speaker embedding vectors, speech tokens, and Mel spectra are all input into the neural network to match a vector field with learnable parameters. For example, the vector field can be determined by the following formula (10):

[0179]

[0180] Continue to refer to Figure 4 In step 413, speech conversion occurs.

[0181] Mel spectrograms are used as input to a vocoder to convert them into personalized, uniformly timbreed speech. The vocoder used here is the Generative Adversarial Network HiFiGAN, which achieves personalized speech output with the same timbre for both Chinese and English.

[0182] In the aforementioned robot voice interaction application scenarios, applying the chatTTS and CosyVoice models to the robot's speech synthesis model improves the speech synthesis effect, increases multi-timbre support, and unifies the timbre across different languages ​​(Chinese and English). This allows the robot to exhibit a more natural, realistic, and emotionally resonant personalized voice during voice interaction, and ensures consistent timbre and natural transitions when synthesizing mixed Chinese and English sentences. This provides a solution for personalized timbre customization and timbre unification in mixed languages ​​for robot voice interaction, significantly enhancing the robot's interactive capabilities.

[0183] The following continues to describe the exemplary structure of the voice data processing device 455 provided in the embodiments of this application as a software module. In some embodiments, such as Figure 2 As shown, the software modules in the map data processing device 455 stored in the memory 450 may include: a timbre determination module 4551, used to determine the target audio data corresponding to the target timbre from the candidate audio data corresponding to each of the N candidate timbres; a first tagging module 4552, used to acquire the first audio data and the first text corresponding to the first audio data, and to encode the first audio data and the first text to obtain an initial tagging sequence; a second tagging module 4553, used to add language conversion tags to the initial tagging sequence when the first text includes at least two languages, to obtain a first target tagging sequence; a feature conversion module 4554, used to perform feature conversion on the first target tagging sequence based on the target audio data to obtain the first audio feature of the first target tagging sequence; and a speech conversion module 4555, used to perform speech conversion on the first audio feature to obtain the second audio data corresponding to the first text, wherein the timbre of the second audio data is the target timbre.

[0184] In some embodiments, the timbre determination module 4551 is further configured to: obtain initial text, perform control symbol insertion and rewriting processing on the initial text to obtain candidate text, determine N candidate timbre parameters, insert the N candidate timbre parameters into the candidate text respectively to obtain N target text, and perform speech conversion on the N target text to obtain the candidate audio data corresponding to each of the N candidate timbres.

[0185] In some embodiments, the timbre determination module 4551 is further configured to perform feature transformation on each target text to obtain a first audio tag sequence; extract features from the first audio tag sequence to obtain a first audio feature; and perform speech conversion on the first audio feature to obtain candidate audio data.

[0186] In some embodiments, the first tagging module 4552 is further configured to extract features from the first audio data to obtain first audio features, and to perform position encoding on the first audio features to obtain an audio representation vector; to perform quantization processing on the audio representation vector to obtain a second audio tag sequence; to perform encoding processing on the first text to obtain a text tag sequence; and to perform combined processing on the second audio tag sequence and the text tag sequence to obtain an initial tag sequence.

[0187] In some embodiments, the first tagging module 4552 is further configured to concatenate the text tagging sequence after the second audio tagging sequence to obtain an initial tagging sequence; and to concatenate the language conversion tagging sequence after the text tagging sequence to obtain a first target tagging sequence.

[0188] In some embodiments, the speech conversion module 4555 is further configured to acquire a generative adversarial network model, and invoke the generative adversarial network model to extract waveforms from the first audio features to obtain a first audio waveform; determine audio data that matches the first audio waveform from the audio database, and determine the matching audio data as the second audio data corresponding to the first text.

[0189] In some embodiments, the speech conversion module 4555 is further configured to: encode the target audio data to obtain a third audio tag sequence when the first audio data and the target audio data are in the same language; determine the prompt word text corresponding to the target audio data and encode the prompt word text to obtain a prompt word tag sequence; insert the third audio tag sequence and the prompt word tag sequence into the initial tag sequence to obtain a second target tag sequence; perform feature conversion on the second target tag sequence based on the target audio data to obtain a second audio feature of the second target tag sequence; and perform speech conversion on the second audio feature to obtain the second audio data corresponding to the first text.

[0190] In some embodiments, the initial tag sequence includes a second audio tag sequence and a text tag sequence, with the second audio tag sequence preceding the text tag sequence. The speech conversion module 4555 is further configured to insert a prompt word tag sequence after the second audio tag sequence and before the text tag sequence, and to insert a third audio tag sequence after the text tag sequence to obtain the second target tag sequence.

[0191] This application provides a computer program product, which includes computer-executable instructions or a computer program stored in a computer-readable storage medium. The processor of an electronic device reads the computer-executable instructions or computer program from the computer-readable storage medium and executes the computer-executable instructions or computer program, causing the electronic device to perform the voice data processing method described above in this application.

[0192] This application provides a computer-readable storage medium storing computer-executable instructions or a computer program. When the computer-executable instructions or the computer program are executed by a processor, the processor will execute the voice data processing method provided in this application. For example, ... Figure 3A The speech data processing method is shown.

[0193] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or it may be a variety of devices including one or any combination of the above-mentioned memories.

[0194] In some embodiments, computer-executable instructions may take the form of programs, software, software modules, scripts, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as stand-alone programs or as modules, components, subroutines, or other units suitable for use in a computing environment.

[0195] As an example, computer-executable instructions may, but do not necessarily, correspond to files in a file system. They may be stored as part of a file that holds other programs or data, for example, in one or more scripts in a Hyper Text Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple co-located files (e.g., files that store one or more modules, subroutines, or code sections).

[0196] As an example, computer-executable instructions can be deployed to execute on a single electronic device, or on multiple electronic devices located in one location, or on multiple electronic devices distributed across multiple locations and interconnected via a communication network.

[0197] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, and improvements made within the spirit and scope of this application are included within the scope of protection of this application.

Claims

1. A voice data processing method, characterized in that, The method includes: From the candidate audio data corresponding to each of the N candidate timbres, determine the target audio data corresponding to the target timbre; Obtain first audio data and the first text corresponding to the first audio data, and encode the first audio data and the first text to obtain an initial tag sequence; When the first text contains at least two languages, language conversion markers are added to the initial marker sequence to obtain the first target marker sequence; Based on the target audio data, feature transformation is performed on the first target marker sequence to obtain the first audio feature of the first target marker sequence; The first audio feature is converted into speech to obtain the second audio data corresponding to the first text, and the timbre of the second audio data is the target timbre.

2. The method according to claim 1, characterized in that, Before determining the target audio data corresponding to the target timbre from the candidate audio data corresponding to each of the N candidate timbres, the method further includes: Obtain the initial text, and perform control symbol insertion and rewriting on the initial text to obtain candidate text; N candidate timbre parameters are determined, and the N candidate timbre parameters are respectively inserted into the candidate text to obtain N target texts; Speech conversion is performed on N target texts to obtain candidate audio data corresponding to each of the N candidate timbres.

3. The method according to claim 2, characterized in that, The process of performing speech conversion on N target texts to obtain candidate audio data corresponding to each of the N candidate timbres includes: For each target text, feature transformation is performed on the target text to obtain a first audio tag sequence; Feature extraction is performed on the first audio tag sequence to obtain the first audio features; The first audio feature is converted into speech to obtain candidate audio data.

4. The method according to claim 1, characterized in that, The encoding process of the first audio data and the first text to obtain the initial marker sequence includes: Feature extraction is performed on the first audio data to obtain the first audio feature, and the first audio feature is positionally encoded to obtain the audio representation vector; The audio representation vector is quantized to obtain a second audio tag sequence; The first text is encoded to obtain a text tag sequence; The second audio tag sequence and the text tag sequence are combined to obtain the initial tag sequence.

5. The method according to claim 4, characterized in that, The process of combining the second audio tag sequence and the text tag sequence to obtain the initial tag sequence includes: The text tag sequence is concatenated after the second audio tag sequence to obtain the initial tag sequence; The step of adding language conversion tokens to the initial token sequence to obtain the first target token sequence includes: The language conversion markers are concatenated to the text marker sequence to obtain the first target marker sequence.

6. The method according to claim 1, characterized in that, The step of performing speech conversion on the first audio features to obtain the second audio data corresponding to the first text includes: Obtain a generative adversarial network model, and call the generative adversarial network model to extract waveforms from the first audio features to obtain the first audio waveform; The audio data that matches the first audio waveform is determined from the audio database, and the matching audio data is determined as the second audio data corresponding to the first text.

7. The method according to any one of claims 1 to 6, characterized in that, The method further includes: When the first audio data and the target audio data are in the same language, the target audio data is encoded to obtain a third audio tag sequence; Determine the prompt word text corresponding to the target audio data, and encode the prompt word text to obtain a prompt word tag sequence; The third audio tag sequence and the cue word tag sequence are inserted into the initial tag sequence to obtain the second target tag sequence; Based on the target audio data, feature transformation is performed on the second target marker sequence to obtain the second audio features of the second target marker sequence; The second audio feature is converted into speech to obtain the second audio data corresponding to the first text.

8. The method according to claim 7, characterized in that, The initial tag sequence includes the second audio tag sequence and the text tag sequence, with the second audio tag sequence preceding the text tag sequence. The step of inserting the third audio tag sequence and the cue word tag sequence into the initial tag sequence to obtain the second target tag sequence includes: The prompt word sequence is inserted after the second audio mark sequence and before the text mark sequence, and the third audio mark sequence is inserted after the text mark sequence to obtain the second target mark sequence.

9. A voice data processing device, characterized in that, The device includes: The timbre determination module is used to determine the target audio data corresponding to the target timbre from the candidate audio data corresponding to each of the N candidate timbres. The first tagging module is used to acquire first audio data and first text corresponding to the first audio data, and to encode the first audio data and the first text to obtain an initial tagging sequence; The second tagging module is used to add language conversion tags to the initial tagging sequence when the first text includes at least two languages, to obtain a first target tagging sequence; The feature conversion module is used to perform feature conversion on the first target marker sequence based on the target audio data to obtain the first audio feature of the first target marker sequence; The speech conversion module is used to convert the first audio feature into speech to obtain the second audio data corresponding to the first text, wherein the timbre of the second audio data is the target timbre.

10. An electronic device, characterized in that, The electronic device includes: Memory is used to store executable instructions or computer programs. A processor, when executing computer-executable instructions or computer programs stored in the memory, implements the voice data processing method according to any one of claims 1 to 8.

11. A computer-readable storage medium storing computer-executable instructions or a computer program, characterized in that, When the computer-executable instructions or computer program are executed by a processor, the voice data processing method according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Speech synthesis method, speech synthesis device, storage medium and electronic equipment

    CN112185340A

  • Cross-language audio conversion method and device, computer equipment and storage medium

    CN112712789A