Audio synthesis method, audio synthesis model training method and related devices
By encoding and fusing audio data, high-quality synthesized audio that matches a specific speaker is generated, solving the problems of inconsistent voice features and high computing resource consumption in speech synthesis technology, and achieving more efficient audio synthesis.
Patent Information
- Application Number
- CN202411771674.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-03
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2044-12-03
AI Technical Summary
Existing speech synthesis technologies have difficulty generating unique speech features that match a specific speaker and consume high computing resources, which limits their application in resource-constrained environments.
By performing a first encoding and a second encoding on the audio data, extracting speech features and musical features, and performing attention processing and fusion processing, synthetic audio that conforms to the first object is generated.
The accuracy of audio synthesis is improved, the synthesized audio is made consistent with the voice features and text content of the first object, and the computing resource requirements are reduced.
Smart Images

Figure CN119763540B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to artificial intelligence technology, and in particular to an audio synthesis method, a training method for an audio synthesis model, and related devices. Background Art
[0002] Speech synthesis technology has made significant progress in recent years, particularly driven by deep learning models, which can generate synthesized speech that is rich in expressiveness and close to natural human speech. However, this technology still faces several challenges. For example, the synthesized speech may not conform to the unique voice characteristics of a specific speaker. Furthermore, achieving high-quality speech synthesis requires significant computational resources and relies on complex optimization algorithms, which can be difficult to implement in resource-constrained environments. Therefore, while speech synthesis technology has great potential in providing fast and flexible speech output, these technical barriers still need to be overcome for wider application. Summary of the Invention
[0003] The embodiments of the present application provide an audio synthesis method, an audio synthesis model training method and related devices, which can improve the accuracy of audio synthesis.
[0004] The technical solution of the embodiment of the present application is implemented as follows:
[0005] The present invention provides an audio synthesis method, which includes:
[0006] determining text features of the first text and determining audio data of the first object;
[0007] Performing a first encoding on the audio data to obtain a speech feature of the first object, and performing a second encoding on the audio data to obtain a musical feature of the audio data;
[0008] Performing attention processing on the musical feature and the text feature to obtain a first feature;
[0009] fusing the speech feature, the first feature, and the text feature to obtain a second feature;
[0010] The second feature is decoded to obtain synthesized audio.
[0011] The present invention provides a method for training an audio synthesis model, the method comprising:
[0012] fusing the plurality of first audio samples of the second object to obtain a second audio sample;
[0013] Performing audio synthesis on the second audio sample and the text sample to obtain a synthesized audio sample of the second object;
[0014] A model is trained based on the synthesized audio samples of the second object.
[0015] The present invention provides an audio synthesis device, including:
[0016] a data acquisition module, configured to determine text features of the first text and determine audio data of the first object;
[0017] a data processing module configured to perform a first encoding on the audio data to obtain a speech feature of the first object, and to perform a second encoding on the audio data to obtain a musical feature of the audio data; perform an attention process on the musical feature and the text feature to obtain a first feature; and perform a fusion process on the speech feature, the first feature, and the text feature to obtain a second feature;
[0018] The audio synthesis module is used to decode the second feature to obtain synthesized audio.
[0019] An embodiment of the present application provides an electronic device, including:
[0020] a memory for storing computer-executable instructions;
[0021] The processor is used to implement the audio synthesis method provided in the embodiment of the present application when executing the computer-executable instructions stored in the memory.
[0022] An embodiment of the present application provides a computer-readable storage medium storing a computer program or computer-executable instructions for implementing the audio synthesis method provided in the embodiment of the present application when executed by a processor.
[0023] An embodiment of the present application provides a computer program product, including a computer program or computer-executable instructions. When the computer program or computer-executable instructions are executed by a processor, the audio synthesis method provided in the embodiment of the present application is implemented.
[0024] The embodiments of the present application have the following beneficial effects:
[0025] The audio data is respectively subjected to a first encoding and a second encoding to obtain the speech features of the first object and the musical features of the audio data, and then the musical features and the text features are subjected to attention processing to obtain the first feature, and the speech features, the first features and the text features are fused to obtain the second feature. In this way, the speech features provide sound information related to the first object, and the text features provide semantic content. By integrating multimodal information such as musical features, speech features, and text features, the performance of audio synthesis is improved to obtain synthesized audio that conforms to the sound of the first object, and the text content of the synthesized audio is consistent with the text content of the first text, thereby improving the accuracy of audio synthesis. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 This is a schematic diagram of the architecture of the audio synthesis system provided in an embodiment of the present application;
[0027] Figure 2A This is a first structural diagram of an electronic device provided in an embodiment of the present application;
[0028] Figure 2B is a second structural diagram of an electronic device provided in an embodiment of the present application;
[0029] Figure 3A This is a schematic diagram of the first flow chart of the audio synthesis method provided in an embodiment of the present application;
[0030] Figure 3B This is a second flow chart of the audio synthesis method provided in an embodiment of the present application;
[0031] Figure 3C 3 is a schematic diagram of a third flow chart of the audio synthesis method provided in an embodiment of the present application;
[0032] Figure 3D 4 is a schematic diagram of a fourth flow chart of the audio synthesis method provided in an embodiment of the present application;
[0033] Figure 3E 5 is a schematic diagram of a fifth flow chart of the audio synthesis method provided in an embodiment of the present application;
[0034] Figure 4 1 is a flow chart of a method for training an audio synthesis model according to an embodiment of the present application;
[0035] Figure 5 This is a first structural diagram of the speech synthesis model provided in an embodiment of the present application;
[0036] Figure 6 This is a schematic diagram of the speaker recognition model structure provided by an embodiment of the present application;
[0037] Figure 7 This is a schematic diagram of the neural network model structure provided by the embodiment of the present application;
[0038] Figure 8 This is a schematic diagram of the structure of the scaled dot product attention model provided in an embodiment of the present application;
[0039] Figure 9 This is a second structural diagram of the speech synthesis model provided in an embodiment of the present application;
[0040] Figure 10 This is a third structural diagram of the speech synthesis model provided in an embodiment of the present application;
[0041] Figure 11 Schematic diagram of the vector quantization model structure provided in the embodiment of the present application;
[0042] Figure 12 This is a schematic diagram of audio synthesis provided in an embodiment of the present application.
[0043] It should be pointed out that the above-mentioned "first" and "second" are only used to distinguish different solutions, and do not represent the degree of distinction between the advantages and disadvantages of the solutions or the priority in the implementation process. DETAILED DESCRIPTION
[0044] In order to make the purpose, technical solutions and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limiting this application. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0045] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0046] In the following description, the terms "first\second\third" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0047] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.
[0048] The relevant data collection and processing in the embodiments of this application should be strictly in accordance with the requirements of relevant national laws and regulations when applied in examples, and the informed consent or separate consent of the personal information subject should be obtained. Subsequent data use and processing should be carried out within the scope of authorization of laws and regulations and the personal information subject.
[0049] Unless otherwise defined, all technical and scientific terms used in the embodiments of the present application have the same meanings as those commonly understood by those skilled in the art. The terms used in the embodiments of the present application are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.
[0050] Before further explaining the embodiments of the present application in detail, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are subject to the following interpretations.
[0051] 1) Audio data is sound stored in digital format. Audio data can be raw waveform data, compressed or encoded data.
[0052] 2) The first object is a specific object that sends audio data.
[0053] In the related art, the audio of different texts of the same object has inconsistent timbre, and the timbre and sound style of the object are ignored when synthesizing audio. To address the above problems, the embodiments of the present application provide an audio synthesis method, a training method for an audio synthesis model, an apparatus, an electronic device, a computer-readable storage medium, and a computer program product, which can improve the accuracy of audio synthesis and obtain synthesized audio consistent with the sound of the first object.
[0054] The audio synthesis method described in the embodiments of the present application can be applied to various fields, such as user interaction, digital human live broadcast (i.e., live broadcast through a digital human, the voice of the digital human live broadcast is generated by preset text and voice), etc., that is, the audio synthesis method in the embodiments of the present application is not limited to a certain field.
[0055] The following describes an exemplary application of the electronic device provided in the embodiment of the present application. The device provided in the embodiment of the present application can be implemented as a terminal or a server. The following describes an exemplary application when the device is implemented as a server.
[0056] See also Figure 1 , Figure 1 This is a schematic diagram of the architecture of the audio synthesis system 100 provided in an embodiment of the present application. In order to support an audio synthesis application, a terminal (terminal 400 is shown as an example) is connected to a server 200 via a network 300. The network 300 can be a wide area network or a local area network, or a combination of the two.
[0057] Terminal 400 is used to send the text features of the first text and the audio data of the first object to server 200 through network 300. Server 200 is used to perform a first encoding on the audio data through a trained audio synthesis model to obtain the voice features of the first object, and to perform a second encoding on the audio data to obtain the musical features of the audio data, perform attention processing on the musical features and text features to obtain the first feature, perform fusion processing on the voice features, the first feature and the text features to obtain the second feature, perform decoding processing on the second feature to obtain synthesized audio, and return the synthesized audio to terminal 400. Terminal 400 displays the synthesized audio through a graphical interface 410.
[0058] Next, an example of audio synthesis performed by terminal 400 will be described.
[0059] In some embodiments, the terminal 400 can independently complete the audio synthesis task. For example, the terminal 400 is used to determine the text features of the first text and determine the audio data of the first object, perform a first encoding on the audio data through a trained audio synthesis model to obtain the voice features of the first object, and perform a second encoding on the audio data to obtain the musical features of the audio data, perform attention processing on the musical features and text features to obtain the first feature, perform fusion processing on the voice features, the first feature and the text features to obtain the second feature, perform decoding processing on the second feature to obtain synthesized audio, and display the synthesized audio through the graphical interface 410.
[0060] In one implementation scenario, the server or terminal can generate information push audio for interaction based on the interactive text with the user, determine the text features of the interactive text, and determine the historical interactive audio of the agent, perform a first encoding on the historical interactive audio through a trained push audio synthesis model to obtain the voice features of the agent, and perform a second encoding on the historical interactive audio to obtain the musical features of the historical interactive audio, perform attention processing on the musical features and text features to obtain the first feature, perform fusion processing on the voice features, the first feature and the text features to obtain the second feature, perform decoding processing on the second feature to obtain information push audio, and push the information push audio to the user to facilitate interaction with the user.
[0061] In one implementation scenario, the server or terminal can generate the speech audio of the digital person in a speech scenario, determine the text features of the speech manuscript, and determine the historical speech audio of the user corresponding to the digital person, perform a first encoding on the historical speech audio through the trained audio synthesis model to obtain the voice features of the user corresponding to the digital person, and perform a second encoding on the historical speech audio to obtain the musical features of the historical speech audio, perform attention processing on the musical features and text features to obtain the first feature, perform fusion processing on the voice features, the first feature and the text features to obtain the second feature, and perform decoding processing on the second feature to obtain the speech audio.
[0062] In some embodiments, the server 200 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.
[0063] The terminal 400 may be a smartphone, tablet computer, laptop computer, desktop computer, smart speaker, smart watch, intelligent voice interaction device, smart home appliance, vehicle-mounted terminal, aircraft, etc., but is not limited thereto. The terminal and the server may be connected directly or indirectly via wired or wireless communication, which is not limited in the embodiments of the present application.
[0064] See also Figure 2A , Figure 2A is a first structural diagram of an electronic device provided in an embodiment of the present application. The electronic device 500 shown in FIG. 2 may be Figure 1 In the terminal 400 or server 200, the electronic device 500 includes at least one processor 510, a memory 550, and at least one network interface 520. The various components in the server 200 are coupled together via a bus system 540. It will be appreciated that the bus system 540 is used to enable communication between these components. In addition to a data bus, the bus system 540 also includes a power bus, a control bus, and a status signal bus. However, for clarity, in FIG2 , all of these buses are labeled as the bus system 540.
[0065] The processor 510 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., where the general-purpose processor can be a microprocessor or any conventional processor, etc.
[0066] The user interface 530 includes one or more output devices 531 that enable presentation of media content, including one or more speakers and / or one or more visual display screens. The user interface 530 also includes one or more input devices 532, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls;
[0067] In some embodiments, when the terminal 400 independently completes the audio synthesis task, the server 200 provided in the embodiment of the present application does not include the user interface 530 .
[0068] The memory 550 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, a hard drive, an optical drive, etc. The memory 550 may optionally include one or more storage devices physically located away from the processor 510.
[0069] The memory 550 includes volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 550 described in the embodiments of the present application is intended to include any suitable type of memory.
[0070] In some embodiments, the memory 550 can store data to support various operations, examples of which include programs, modules, and data structures, or a subset or superset thereof, as exemplified below.
[0071] Operating system 551, including system programs for processing various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, driver layer, etc., for implementing various basic services and processing hardware-based tasks;
[0072] A network communication module 552 for reaching other computing devices via one or more (wired or wireless) network interfaces 520 , exemplary network interfaces 520 including Bluetooth, WiFi, and USB;
[0073] a presentation module 553 for enabling presentation of information via one or more output devices 531 (e.g., a display screen, a speaker, etc.) associated with the user interface 530 (e.g., a user interface for operating peripheral devices and displaying content and information);
[0074] In some embodiments, when the audio synthesis task is completed independently by the terminal 400, the server 200 provided by the embodiment of the present application may not include the presentation module 553.
[0075] The input processing module 554 is used to detect one or more user inputs or interactions from one of the one or more input devices 532 and translate the detected inputs or interactions; in some embodiments, when the embodiment independently completes the audio synthesis task by the terminal 400, the server 200 provided in the embodiment of the present application may not include the presentation module 553.
[0076] In some embodiments, the audio synthesis device provided in the embodiments of the present application can be implemented in software. Figure 2A An audio synthesis device 555 stored in memory 550 is shown. This device can be software in the form of a program or plug-in, and includes the following software modules: a data acquisition module 5551, a data processing module 5552, and an audio synthesis module 5553. These modules are logical and can be arbitrarily combined or further separated according to the functions they implement. The functions of each module will be described below.
[0077] In some embodiments, the training device for the audio synthesis model provided in the embodiments of the present application can also be implemented in software, see Figure 2B , Figure 2B is a second structural diagram of an electronic device provided in an embodiment of the present application, Figure 2B Except for the training device 556 based on the audio synthesis model shown, the rest can be the same as Figure 2A The audio synthesis model training device 556 stored in the memory 550 can be software in the form of a program or plug-in, and includes the following software modules: a pre-processing module 5561 and a model training module 5562. These modules are logical and can be arbitrarily combined or further separated according to the functions they implement. The functions of each module will be described below.
[0078] It should be noted that in the following audio synthesis example, those skilled in the art can apply the audio synthesis method provided in the embodiment of the present application to synthesize audio based on their understanding of the following.
[0079] See also Figure 3A , Figure 3A This is a first flow chart of the audio synthesis method provided in the embodiment of the present application, which will be combined with Figure 3A The steps shown are used to illustrate the audio synthesis method provided in the embodiment of the present application. The audio synthesis method can be implemented by the server or the terminal alone, or by the server and the terminal in collaboration. The following will illustrate the collaborative implementation by the server and the terminal as an example.
[0080] In step 101 , text features of a first text and audio data of a first object are determined.
[0081] Here, the first text includes a series of symbols, such as letters, Chinese characters, numbers, punctuation marks, etc. These symbols are arranged according to specific grammatical rules to express specific meanings (such as knowledge, emotions, opinions, instructions, etc.). Grammatical rules are rules for constructing text and are used to characterize the arrangement order of symbols; text features are deep information extracted from the original text, and the deep information can be used to perform tasks such as text classification and sentiment analysis. The embodiment of the present application does not limit the text features. The text features can be word frequency, word embedding vectors, etc.; the first object is a specific object that emits audio data. Different objects have different object information (such as object identifiers, where the object identifier is used to identify a unique object, and each object corresponds to its own object identifier); audio data is sound stored in digital format. The embodiment of the present application does not limit the audio data. The audio data can be original waveform data, compressed or encoded data, etc. Audio data is used in multiple fields such as speech recognition, music analysis, and sound synthesis.
[0082] In some embodiments, "determining the text features of the first text" in step 101 can be achieved by: performing word segmentation processing on the first text to obtain text word segments; performing stem extraction on each text word segment to obtain target text word segments; screening the target text word segments, and vectorizing the screened target text word segments to obtain text features of the first text.
[0083] For example, the first text is segmented to obtain text segmentations, and the continuous character sequence in the first text is split into meaningful words or phrases; stem extraction is performed on each text segmentation to convert the word into its basic form to reduce the dimension of the data and eliminate the impact of word form changes. For example, "running", "runs" and "ran" can all be restored to "run"; stop words (such as "the") are removed; the target text segmentations obtained after stem extraction are vectorized to convert the text into a numerical vector. The vectorization process can be achieved in the following ways: a vocabulary is constructed based on the first text, and the vocabulary contains all non-repeating words; the first text is segmented and the text segmentations are matched with the vocabulary, and the frequency of each text segmentation in the first text is counted; the first text is converted into a vector, each element in the vector corresponds to a word in the vocabulary, and the value of the element is used to represent the probability of the corresponding word appearing in the first text.
[0084] Continuing with the above example, the above-mentioned "building a vocabulary based on the first text" can be implemented in the following way: extracting an initial vocabulary from historical text data, the vocabulary contains all non-repeated first words, taking two consecutive characters in the first text as the second word, and updating the initial vocabulary according to the probability of occurrence of the second word in the first text and the probability of occurrence of the first word in the first text to obtain a vocabulary.
[0085] Continuing with the above example, given the text "The quick brown fox" and the vocabulary ["the", "quick", "brown", "fox", "jumped", "over", "lazy", "dog", "outran"], the text features of the text are [0.25, 0.25, 0.25, 0, 0, 0, 0], where the value of each element represents the probability of the corresponding word appearing in the text, and 0 means that the word does not appear in the text.
[0086] In some embodiments, "determining the audio data of the first object" in step 101 can be achieved in the following manner: determining the initial audio data of the first object; when the duration of the initial audio data is less than the duration threshold, determining the first quantity based on the duration of the initial audio data and the duration threshold; splicing the initial audio data that meets the first quantity to obtain the audio data of the first object.
[0087] Continuing from the above embodiment, the above “determining the first quantity based on the duration of the initial audio data and the duration threshold” can be achieved in the following way: the first quantity is determined by adding the ratio of the duration threshold to the duration of the initial audio data and a preset quantity (such as 1).
[0088] It should be noted that the initial audio data is audio data with a duration less than the duration threshold. By splicing the initial audio data, audio data with a duration greater than or equal to the duration threshold is obtained, so that the audio synthesis model can learn rich sound information from the audio data.
[0089] For example, initial audio data (such as [1, -1, 2, -2]) that meets the first quantity (such as 2) is spliced to obtain audio data of the first object (such as [1, -1, 2, -2, 1, -1, 2, -2]).
[0090] In step 102, the audio data is first encoded to obtain the speech features of the first object, and the audio data is second encoded to obtain the musical features of the audio data.
[0091] Here, the first encoding and the second encoding are different encoding methods. The encoding is used to map audio data to a high-dimensional space and to convert audible audio data into feature vectors that cannot be perceived. The speech feature is manifested as a sound pattern unique to the object in the spectrum. The embodiment of the present application does not limit the speech feature. The speech feature can be used to characterize the timbre of the sound (such as the unique distribution of the timbre) and the style of the sound. The timbre of the sound distinguishes different objects through the specific frequency components in the spectrum and the relative intensity of the specific frequency components. The style of the sound is the unique characteristics and expression form of the sound, which is used to reflect the individual characteristics or emotional state of the sound source. The embodiment of the present application does not limit the speech feature. There is no restriction on the style of the voice, which can be the range, volume, etc., among which the range is the range from the lowest to the highest note that the object can produce, and the volume is the loudness of the sound, which is used to affect the perceived intensity and emotional expression of the sound; the musical characteristics include the basic components of speech reflected in the spectrum, such as the resonance peak distribution of vowels and consonants, which are used to distinguish the spectral characteristics of different phonemes in the language. Taking the sound as an example, when a vowel (such as a) is uttered, a specific set of resonance peaks will appear in the spectrum of each object, and the positions of the resonance peaks are roughly the same. When a consonant (such as p) is uttered, since the vocal cords do not vibrate, the spectrum is mainly composed of high-frequency noise components, and the resonance peaks are not obvious or missing.
[0092] It should be noted that the musical characteristics also include the frequency distribution of different punctuation marks in the frequency spectrum. Since the pause durations of punctuation marks of different objects are not exactly the same, the musical characteristics of different objects are not exactly the same.
[0093] In some embodiments, see Figure 3B , Figure 3B This is a second flow chart of the audio synthesis method provided in the embodiment of the present application, for Figure 3A Step 102 shown in the figure performs a first encoding on the audio data to obtain the speech features of the first object, which can be obtained by Figure 3B Steps 1021A to 1024A are implemented as described below.
[0094] In step 1021A, cepstrum processing is performed on the audio data to obtain the eighth feature.
[0095] Here, cepstrum processing extracts features by analyzing the frequency spectrum of audio data, and the eighth feature is used to capture important properties of the audio data while reducing the effects of noise and instability.
[0096] In some embodiments, step 1021A and the following step 1021B can be implemented by: performing frequency domain transformation processing on the audio data to obtain an audio spectrum; converting the audio spectrum into a first power spectrum, and filtering the first power spectrum to obtain a second power spectrum; performing cosine transformation on the logarithm of the second power spectrum to obtain the eighth feature.
[0097] It should be noted that the embodiments of the present application do not limit the frequency conversion method. The frequency conversion method can be fast Fourier transform, two-dimensional discrete Fourier transform, etc. The filtering process is used to modify or extract specific frequency components of the signal (such as the first power spectrum); the embodiments of the present application do not limit the filtering process method. The filtering process can be low-pass filtering, high-pass filtering, etc., among which low-pass filtering allows low-frequency signals to pass through and blocks high-frequency signals, and is used to eliminate high-frequency noise and smooth the signal; high-pass filtering allows high-frequency signals to pass through and blocks low-frequency signals, and is used to highlight the rapidly changing parts of the signal.
[0098] In some embodiments, the above-mentioned "performing frequency domain transform processing on the audio data to obtain an audio spectrum" can be achieved in the following way: performing Fourier transform on the audio data, transforming the audio data from the spatial domain to the frequency domain, and obtaining the audio spectrum of the audio data, wherein each spectrum point of the audio spectrum is the amplitude at each frequency, and each spectrum point corresponds to a specific frequency value. The amplitude of the spectrum point reflects the intensity or energy of the frequency component in the signal. Since different sounds have unique frequency distributions, the amplitude of the spectrum point is used to identify the timbre of different sounds.
[0099] For example, Fourier transform is performed on audio data (such as [1, -1, 2, -2]), and the audio data is transformed from the spatial domain to the frequency domain to obtain the audio spectrum of the audio data (such as [0, 0.6, 0.8, ..., 0.5], where the spectrum point 0 with the subscript 0 is used to represent the amplitude of the audio corresponding to the frequency when the frequency is 0), where the length of the audio data is used to represent the length of time, the time interval between two consecutive data points of the audio data is the same, and the value of the data point in the audio data is the amplitude of the audio (such as 1 and -1 are the amplitudes of audio data collected at different time points).
[0100] Following the above embodiment, the above “converting the audio spectrum into the first power spectrum” may be achieved in the following manner: each spectrum point in the audio spectrum is replaced by the square of each spectrum point to obtain the first power spectrum.
[0101] It should be noted that the power spectrum is the distribution of signal power at each frequency, and the power is proportional to the square of the amplitude.
[0102] For example, given an audio spectrum (such as the frequency of frequency 1 is 440 Hz, the value of the spectrum point A1 of the audio spectrum at frequency 1 is 1, the frequency of frequency 2 is 880 Hz, and the value of the spectrum point A2 of the audio spectrum at frequency 2 is 0.5), each spectrum point in the audio spectrum is replaced by the square of each spectrum point (the square of the value of the spectrum point A1 of the audio spectrum at frequency 1 is 1, and the square of the value of the spectrum point A2 of the audio spectrum at frequency 1 is 0.25) to obtain a first power spectrum.
[0103] Continuing with the above embodiment, taking high-pass filtering as an example, the above-mentioned "filtering the first power spectrum to obtain the second power spectrum" can be achieved in the following way: determine the power threshold, perform the following processing for each power in the first power spectrum, and when the power is less than the power threshold, replace the power in the first power spectrum with a preset power (such as 0) to obtain the second power spectrum.
[0104] For example, given a first power spectrum (such as [0.6, 0.8, 0.5], when the power in the first power spectrum is less than the power threshold (such as 0.6), the power in the first power spectrum (such as 0.5) is replaced with a preset power (such as 0) to obtain a second power spectrum (such as [0.6, 0.8, 0]).
[0105] For example, a cosine transform is performed on the logarithm (such as [-1, -2, 0]) of the second power spectrum (such as [0.5, 0.25, 1]), the logarithm is expressed as a set of coefficients of cosine functions, and each set of cosine functions is fused (cosine function A is "-1*cos(0*x)", cosine function B is "-2*cos(0.01*π*x)", and cosine function A is "0*cos(0.01*π*2*x)") to obtain the eighth feature, where the frequency of each set of cosine functions is the product of the index subscript of the logarithm and a preset frequency (such as the ratio of π to 100, 100 is used to represent the length of the second power spectrum).
[0106] Through the embodiments of the present application, filtering processing is performed to effectively reduce the impact of noise and improve the clarity of audio data. In addition, cepstrum processing can extract key frequency components in audio data, thereby improving the accuracy of audio synthesis. At the same time, it reduces the instability of audio data and improves the robustness of features.
[0107] In step 1022A, residual processing is performed on the eighth feature to obtain a ninth feature.
[0108] Here, residual processing solves the gradient vanishing problem in network training by adding skip connections, and the ninth feature is used to measure the difference between the outputs of different neural network layers.
[0109] In some embodiments, step 1022A may be implemented by performing convolution processing on the eighth feature to obtain a spectrum convolution feature, and fusing the spectrum convolution feature and the eighth feature to obtain a ninth feature.
[0110] It should be noted that the above-mentioned step of "performing convolution processing on the eighth feature to obtain a spectral convolution feature" is similar to the step of "performing convolution processing on the eighth feature to obtain a convolution feature" in the following step 1022B, and the above-mentioned step of "fusing the spectral convolution feature and the eighth feature to obtain the ninth feature" is similar to the step of "fusing the first normalized feature and the multi-head attention feature to obtain a fused feature", which will not be repeated here.
[0111] In step 1023A, context mask processing is performed on the ninth feature to obtain the tenth feature.
[0112] Here, context mask processing extracts context information of different scales through global and segment-level pooling operations, and the extracted context information is used to remove irrelevant noise in the features.
[0113] In some embodiments, step 1023A can be implemented by: performing global pooling on the ninth feature to obtain a first pooling feature; performing segment-level pooling on the ninth feature to obtain a second pooling feature, fusing the first pooling feature and the second pooling feature to obtain a third pooling feature; performing mapping processing on the third pooling feature to obtain a residual mapping feature, and performing convolution processing on the ninth feature to obtain a residual convolution feature; fusing the residual mapping feature and the residual convolution feature to obtain the tenth feature.
[0114] It should be noted that global pooling is used to compress the entire spatial dimension of each ninth feature into a single numerical value and convert the ninth feature into a vector of fixed length. The embodiment of the present application does not limit global pooling. Global pooling can be global average pooling, global maximum pooling, etc., wherein global average pooling is used to calculate the average value of all elements of the ninth feature, and global maximum pooling is used to select the maximum value from each element of the ninth feature; segment-level pooling is used to divide the ninth feature into multiple windows and then apply the pooling operation on each window.
[0115] For example, the ninth feature (such as [0.28, 0.16, 0, 0.14, 0.13, -0.06, 0.15, 0.03]) is subjected to global maximum pooling to obtain the first pooling feature (such as 0.28); the ninth feature is subjected to segment-level pooling to obtain the second pooling feature (such as [0.28, 0, -0.06, 0.15]), and the first pooling feature and the second pooling feature are fused to obtain the third pooling feature (such as [0.28, 0.28, 0, -0 .06, 0.15]); the third pooling feature is mapped to obtain the residual mapping feature (such as [0.14, 0.14, 0, -0.03, 0.075]), and the ninth feature is convolved to obtain the residual convolution feature (such as [0.19, 0.10, 0, 0.28, 0.16]); the residual mapping feature and the residual convolution feature are fused to obtain the tenth feature (such as [0.33, 0.24, 0, 0.25, 0.235]).
[0116] In step 1024A, dense convolution processing is performed on the tenth feature to obtain the speech feature of the first object.
[0117] Here, dense convolution processing is implemented through multiple dense convolution layers. The input features of each layer are obtained by concatenating the output features of the previous layer and the input features of the previous layer. The network structure of each dense convolution layer is consistent.
[0118] For example, take three layers of dense convolutional layers as an example to illustrate, the tenth feature (such as [0.33, 0.24, 0, 0.25, 0.235]) is convolved by dense convolutional layer A to obtain the first dense convolutional feature (such as [0.34, 0.23, 0, 0.24, 0.22]), the first dense convolutional feature and the tenth feature are concatenated to obtain the first spliced feature (such as [0.33, 0.24, 0, 0.25, 0.235, 0.34, 0.23, 0, 0.24, 0.22]), and the first spliced feature is concatenated by dense convolutional layer B. The second dense convolution feature is convolved with the first convolution feature to obtain a second dense convolution feature such as [0.23, 0.3, 0, 0.2, 0.5]), the second dense convolution feature and the first concatenation feature are concatenated to obtain a second concatenation feature (such as [0.33, 0.24, 0, 0.25, 0.235, 0.34, 0.23, 0, 0.24, 0.22, 0.23, 0.3, 0, 0.2, 0.5]), and the second concatenation feature is convolved through the dense convolution layer C to obtain the speech feature of the first object (such as [0.32, 0.3, 0, 0.6, 0.1]).
[0119] Through the embodiments of the present application, residual connections and context mask processing can improve the model's ability to learn complex speech features. The design of dense convolutional layers can extract deeper features to enhance the model's ability to capture sound details, thereby improving the accuracy of audio synthesis.
[0120] In some embodiments, see Figure 3C , Figure 3C This is a third flow chart of the audio synthesis method provided in the embodiment of the present application, for Figure 3A The second encoding of the audio data in step 102 is performed to obtain the musical characteristics of the audio data, which can be obtained by Figure 3C Steps 1021B to 1023B are implemented as described below.
[0121] In step 1021B, cepstrum processing is performed on the audio data to obtain the eighth feature.
[0122] Here, step 1021B is similar to step 1021A and will not be repeated here.
[0123] In step 1022B, convolution processing is performed on the eighth feature, and pooling processing is performed on the convolved eighth feature to obtain an eleventh feature.
[0124] Here, convolution processing is used to extract local information of the eighth feature (such as edge information, texture information, etc.), and pooling processing is used to reduce the spatial dimension of the data (such as the eighth feature after convolution) while retaining the most important features. The embodiment of the present application does not limit the method of pooling processing, and the pooling processing can be maximum pooling, average pooling, etc.
[0125] In some embodiments, the above-mentioned "convolution processing of the eighth feature" can be achieved in the following way: determine the convolution kernel, perform sliding processing on the eighth feature based on the convolution kernel to obtain multiple coverage areas, calculate the dot product of the convolution kernel and the eighth feature in each coverage area, combine each dot product, and obtain the convolved eighth feature, wherein the size of the coverage area is consistent with the size of the convolution kernel.
[0126] It should be noted that the embodiment of the present application does not limit the convolution kernel, and the convolution kernel can be a Gaussian filter, a sharpening filter, etc.
[0127] For example, the convolution kernel is determined (such as [1, 0, -1]), and the eighth feature (such as [1, 2, 3, 4, 5, 4]) is slidingly processed based on the convolution kernel to obtain multiple coverage areas (such as coverage area A is [1, 2, 3], coverage area B is [2, 3, 4], coverage area C is [3, 4, 5], and coverage area D is [4, 5, 4]). The dot product between the convolution kernel and the eighth feature in each coverage area is calculated (such as the dot product of coverage area A is -2, the dot product of coverage area B is -2, the dot product of coverage area C is -2, and the dot product of coverage area D is 0). Each dot product is combined to obtain the convolved eighth feature (such as [-2, -2, -2, 0]).
[0128] In some embodiments, the above-mentioned "pooling the eighth feature after convolution to obtain the eleventh feature" can be achieved by the following method: dividing the eighth feature after convolution into regions to obtain multiple pooling windows, determining the maximum value of the elements in each pooling window, and combining each maximum value to obtain the eleventh feature.
[0129] It should be noted that the pooling window is a window of a specific size.
[0130] For example, the eighth feature after convolution (such as [1, 2, 3, 4, 5, 4]) is divided into regions to obtain multiple pooling windows (such as pooling window A is [1, 2], pooling window B is [3, 4], and pooling window C is [5, 4]), and the maximum value of the elements in each pooling window is determined (such as the maximum value of the elements in pooling window A is 2, the maximum value of the elements in pooling window B is 4, and the maximum value of the elements in pooling window C is 5). Each maximum value is combined to obtain the eleventh feature (such as [2, 4, 5]).
[0131] In step 1023B, the eleventh feature is mapped and scaled to obtain the musicality feature of the audio data.
[0132] Here, mapping processing is used to map the fused features to a new feature space to obtain features for characterizing the potential structure of the eighth feature. The embodiment of the present application does not limit the mapping processing. The mapping processing can be linear mapping, nonlinear mapping, etc. The scaling processing is used to adjust the numerical range of the feature to a specific interval.
[0133] In some embodiments, taking the mapping process as a linear mapping process as an example, the above-mentioned "mapping process for the eleventh feature" can be achieved in the following way: determining a mapping parameter, and determining the product of the eleventh feature and the mapping parameter as the eleventh feature after mapping.
[0134] For example, a mapping parameter (such as 2) is determined, and the product of the eleventh feature (such as [1, 0, -1]) and the mapping parameter (such as [2, 0, -2]) is determined as the eleventh feature after mapping.
[0135] In some embodiments, the above-mentioned "scaling the mapped eleventh feature to obtain the musical characteristics of the audio data" can be achieved in the following way: convolving the mapped eleventh feature through one-dimensional convolution to obtain the musical characteristics of the audio data, wherein the dimension of the one-dimensional convolution is a specific dimension (such as 1), and the step size of the one-dimensional convolution is a specific step size (such as 10).
[0136] For example, the convolution kernel of the one-dimensional convolution is determined (such as 0.5), the step size of the one-dimensional convolution is 3, and the eleventh feature after mapping based on the convolution kernel (such as [1, 2, 3, 4, 5, 4]) is slidingly processed to obtain multiple coverage areas (such as coverage area A is [1], coverage area B is [4]), and the dot product of the convolution kernel and the mapped eleventh feature in each coverage area is calculated (such as the dot product of coverage area A is 0.5, and the dot product of coverage area B is 2). Each dot product is combined to obtain the musical characteristics of the audio data (such as [0.5, 2]).
[0137] Through the embodiments of the present application, local details and global structures of audio signals can be extracted through convolution and pooling processing, and scaling processing helps to adjust the numerical range and time scale of features, making the features more stable, and reducing the spatial dimension of the features, reducing the number of parameters and computational burden of the model, thereby improving the efficiency of audio synthesis.
[0138] Continue to see Figure 3A ,In step 103, attention processing is performed on the musical features and text features to obtain the first feature.
[0139] Here, step 103 can be implemented in the following manner: performing attention processing on the musical features and text features through the attention layer to obtain the first feature.
[0140] It should be noted that the attention layer includes an attention model that performs attention processing on the data. The embodiment of the present application does not limit the attention model. The attention model can be a multi-head self-attention network or a single-head self-attention network, etc. The attention processing is used to learn important information in the musical features and text features from different angles at the same time through the attention model.
[0141] For example, attention processing is performed on the musical features (such as [0.14, 0.13, -0.06]) and the text features (such as [0.15, 0.03, 0.06]) to obtain the first feature (such as [0.28, 0.16, 0]).
[0142] In some embodiments, see Figure 3D , Figure 3D This is a fourth flow chart of the audio synthesis method provided in the embodiment of the present application, for Figure 3A Step 103 shown can be performed by Figure 3D Steps 1031 to 1033 are implemented as described below.
[0143] In step 1031 , the musical feature is normalized to obtain the third feature, and the text feature is normalized to obtain the fourth feature.
[0144] Here, normalization is used to adjust the data to a uniform scale to avoid deviations caused by different features having different measurement units and numerical ranges. The embodiments of the present application do not limit the normalization method, and normalization can be batch normalization, layer normalization, etc.
[0145] In some embodiments, the above-mentioned "normalizing the musical features to obtain normalized features" can be achieved in the following ways: based on a preset norm parameter, performing a power operation on each element of the musical features to obtain a power element; summing up each power element to obtain a sum element; performing a power operation on the sum element based on the inverse of the norm parameter to obtain a target element; determining the ratio of each element of the musical features to the target element as a normalized feature, wherein the norm parameter (i.e., the exponent) is used to characterize the number of times the element itself is multiplied, and the norm parameter is a real number greater than 0.
[0146] For example, given a musical feature (such as [2, 4]), a norm parameter (such as 2), calculate the square of each element of the musical feature to obtain a power element (such as 2 squared is 4, 2 squared is 4), calculate the sum of each power element to obtain a sum element (such as 20), use the reciprocal of the norm parameter (such as 0.5, that is, the square root) to perform a power operation on the sum element to obtain a target element (such as 4.5), and determine the ratio of each element of the musical feature to the target element (such as the ratio of 2 to 4.5 is 0.44, and the ratio of 2 to 4.5 is 0.88) as a normalized feature (such as [0.44, 0.88]).
[0147] In some embodiments, the above step of "normalizing the text feature to obtain the fourth feature" is similar to the above step of "normalizing the musical feature to obtain the third feature", and will not be repeated here.
[0148] In step 1032, multi-head attention processing is performed on the third feature and the fourth feature to obtain the fifth feature.
[0149] Here, multi-head attention processing is the process of performing data processing on the third feature and the fourth feature through a multi-head attention network, which is used to simultaneously focus on information at different positions when processing data (such as the third feature and the fourth feature), thereby improving the ability to capture information.
[0150] In some embodiments, step 1032 can be implemented by: determining a key vector and a value vector of the audio data based on the third feature, and determining a query vector for the first text based on the fourth feature; determining a similarity matrix for the first text based on the query vector and the key vector; and obtaining a fifth feature based on the similarity matrix and the value vector. In this embodiment, the product of the similarity matrix and the value vector can be determined as the fifth feature.
[0151] Here, the query vector is a vector representation used for comparison with the key vector, and the query vector is used to find the correlation between any element in the text and any element in the audio data; the key vector is a vector representation that matches the query vector, and the key vector of each element in the audio data determines the importance of the element in the attention weight. The higher the degree of match between the key vector and the query vector, the greater the attention weight of the corresponding element; after calculating the attention weight of each element, the value vector is used to perform weighted summation based on the attention weight to generate the fifth feature; the similarity matrix is obtained by calculating the similarity between the query vector and the key vector. Each element of the similarity matrix represents the degree of match between the corresponding query vector and the key vector. The embodiment of the present application does not limit the similarity between features. The similarity can be the dot product of features, the cosine similarity between features, etc.
[0152] In some embodiments, the above-mentioned "determining the key vector of the audio data and the value vector of the audio data based on the third feature, and determining the query vector of the first text based on the fourth feature" can be implemented in at least one of the following ways: determining the third feature as the key vector and the value vector, and determining the fourth feature as the query vector; or, mapping the third feature to obtain the key vector and the value vector, and mapping the fourth feature to obtain the query vector.
[0153] It should be noted that the above-mentioned "mapping the third feature to obtain a key vector and a value vector, and mapping the fourth feature to obtain a query vector" is similar to mapping the third feature to obtain a key vector and a value vector, and mapping the fourth feature to obtain a query vector in step 1023B, and will not be repeated here.
[0154] Following the above embodiment, the above “determining a similarity matrix based on the query vector and the key vector” can be implemented in the following manner: determining the dot product of the query vector and the key vector as the similarity matrix.
[0155] For example, a similarity matrix is determined based on the dot product (e.g., [[1, 0.6, 1.2], [1.5, 0.9, 1.8], [-0.5, -0.3, -0.6]]) of the query vector (e.g., [2, 3, -1]) and the key vector (e.g., [0.5, 0.3, 0.6]); and the product (e.g., [1.4, 2.05, -0.7]) of the similarity matrix and the value vector (e.g., [0.5, 0.3, 0.6]) is determined as the fifth feature.
[0156] Through the embodiments of the present application, the model can pay attention to different parts of the input data at the same time, which helps to capture local and global contextual information. By allocating attention weights, the model can automatically focus on important features, thereby optimizing the feature selection process. The model learns different spatial representations, increases the expressive power of the model, and can understand complex data structures.
[0157] In step 1033 , the first feature is determined based on the third feature and the fifth feature.
[0158] In some embodiments, step 1032 can be implemented by: fusing the third feature and the fifth feature to obtain the sixth feature; mapping the sixth feature to obtain the seventh feature; and fusing the sixth feature and the seventh feature to obtain the first feature.
[0159] Continuing from the above embodiment, the above “fusing the third feature and the fifth feature to obtain the sixth feature” can be achieved by concatenating the third feature and the fifth feature to obtain the sixth feature, or performing weighted summation of the third feature and the fifth feature to obtain the sixth feature.
[0160] For example, the third feature (such as [0.14, 0.13, -0.06]) and the fifth feature (such as [0.15, 0.03, 0.06]) are concatenated to obtain the sixth feature (such as [0.14, 0.13, -0.06, 0.15, 0.03, 0.06]); or, the third feature and the fifth feature are weighted and summed to obtain the sixth feature (taking the weights of the third feature and the fifth feature as 0.5 as an example, the sixth feature (such as [0.28, 0.16, 0]) is obtained.
[0161] Continuing with the above embodiment, mapping is used to map the sixth feature to a new feature space to obtain a seventh feature used to characterize the latent structure of the sixth feature. This embodiment of the present application does not limit the mapping method, and the mapping method can be linear mapping, nonlinear mapping, etc. The above step of "fusing the seventh feature and the sixth feature to obtain the first feature" is similar to the above step of "fusing the third feature and the fifth feature to obtain the sixth feature" and is not repeated here.
[0162] For example, the third feature and the fifth feature are fused to obtain the sixth feature, and the sixth feature (such as [0.28, 0.16, 0]) is mapped to obtain the seventh feature (such as [0.14, 0.08, 0]); the seventh feature and the sixth feature are fused to obtain the first feature (such as [0.42, 0.24, 0]).
[0163] Continue to see Figure 3A In step 104, the speech feature, the first feature and the text feature are fused to obtain the second feature.
[0164] Here, in some embodiments, the above-mentioned "step of fusing speech features, first features and text features to obtain the sixth feature" can be implemented by splicing the speech features, the first features and the text features to obtain the sixth feature, or performing weighted summation of the speech features, the first features and the text features to obtain the sixth feature.
[0165] For example, the speech feature (such as [0.28, 0.16, 0]), the first feature (such as [0.14, 0.13, -0.06]) and the text feature (such as [0.15, 0.03, 0.06]) are spliced to obtain the sixth feature (such as [0.28, 0.16, 0, 0.14, 0.13, -0.06, 0.15, 0.03, 0.06]); or, the speech feature, the first feature and the text feature are weighted and summed to obtain the sixth feature (taking the weights of the speech feature, the first feature and the text feature as 0.33 as an example, the sixth feature is obtained as [0.19, 0.10, 0]).
[0166] Through the embodiments of the present application, the speech features, the first features and the text features are integrated so that the synthesized audio contains not only the text features but also the speech features of the object, so as to ensure that the synthesized audio conforms to the voice of the object and improve the accuracy of audio synthesis.
[0167] In step 105, the second feature is decoded to obtain synthesized audio.
[0168] The synthesized audio matches the voice of the first object, and the text content of the synthesized audio is consistent with the text content of the first text.
[0169] Here, the decoding process is used to convert the feature vectors of the high-dimensional space back into the form of original data or a form that can be understood by the object (such as audio data). The synthesized audio is audio data obtained by performing audio synthesis processing on the audio data and the first text through an audio synthesis method. The sound of the first object includes the timbre and pronunciation style unique to the first object. Taking different ages as an example, the pronunciation style can be a juvenile voice, a young voice, an old voice, etc. The text content of the first text is each character in the first text. The text content of the synthesized audio can be the text content obtained by performing speech recognition processing on the synthesized audio.
[0170] It should be noted that the melody of the synthesized audio conforms to the melody of the audio data of the first object.
[0171] In some embodiments, see Figure 3E , Figure 3E This is a fifth flow chart of the audio synthesis method provided in the embodiment of the present application, for Figure 3A Step 105 shown can be performed by Figure 3E Steps 1051 to 1053 are implemented as described below.
[0172] In step 1051, the speech feature and the second feature are predicted to obtain the twelfth feature.
[0173] In some embodiments, step 1051 can be implemented by normalizing the speech feature to obtain a first normalized feature, normalizing the second feature to obtain a second normalized feature, performing multi-head attention processing on the first normalized feature and the second normalized feature to obtain a twelfth feature.
[0174] It should be noted that the above-mentioned step of "normalizing the speech features to obtain the first normalized feature, and normalizing the second feature to obtain the second normalized feature" is similar to step 1031, and the above-mentioned step of "performing multi-head attention processing on the first normalized feature and the second normalized feature to obtain the twelfth feature" is similar to step 1032, and will not be repeated here.
[0175] In step 1052, mapping processing is performed on the twelfth feature, and upsampling processing is performed on the mapped twelfth feature to obtain the thirteenth feature.
[0176] Here, step 1052 is similar to the step of mapping the eleventh feature in step 1023B and will not be repeated here. The embodiment of the present application does not limit the upsampling process. The upsampling process can be a nearest neighbor interpolation method, a bilinear interpolation method, etc. The upsampling process is used to increase the spatial dimension of the feature map. The nearest neighbor interpolation method is used as an example to illustrate that the nearest neighbor pixel point is selected to fill the new spatial position.
[0177] For example, the mapped twelfth feature (such as [[1, 2], [3, 4]]) is upsampled to obtain the thirteenth feature (such as [[1, 1, 2, 2], [1, 1, 2, 2], [3, 3, 4, 4], [3, 3, 4, 4]]).
[0178] In step 1053, deconvolution processing is performed on the thirteenth feature to obtain synthesized audio.
[0179] In some embodiments, step 1053 can be implemented by determining a transposed convolution kernel, sliding the transposed convolution kernel on the thirteenth feature through a sliding window, calculating the dot product of each element of the thirteenth feature and the transposed convolution kernel, fusing each dot product to obtain a synthetic audio feature, and activating the audio feature to obtain a synthetic audio.
[0180] It should be noted that deconvolution processing is the inverse process of convolution processing, and activation processing introduces nonlinearity through the activation function to solve the nonlinear problem. The embodiment of the present application does not limit the activation processing. The activation function can be a linear mapping unit (Rectified Linear Unit, ReLU), a leaky linear mapping unit (Leaky Rectified Linear Unit, Leaky-ReLU), etc., wherein the linear mapping unit is used to characterize that when the input is greater than 0, the input value is output, otherwise 0 is output. The linear mapping unit is simple to calculate and has a fast training speed.
[0181] For example, determine the transposed convolution kernel (such as [[1, 2], [3, 4]]), slide the transposed convolution kernel on the thirteenth feature (such as [[1, 2], [3, 4]]) through a sliding window, calculate the dot product of each element of the thirteenth feature and the transposed convolution kernel (such as the dot product corresponding to element 1 [[1, 2], [3, 4]], the dot product corresponding to element 2 [[2, 4], [6, 8]], the dot product corresponding to element 3 [[3, 6], [9, 12]], the dot product corresponding to element 4 [[4, 8], [12, 16]]), fuse each dot product to obtain a synthetic audio feature (such as [[1, 4, 4], [6, 20, 16], [9, 24, 16]]), activate the audio feature, and obtain the synthetic audio (such as [1, 3, 4, -5, -2, 4]).
[0182] See also Figure 4 , Figure 4 This is a flow chart of the training method of the audio synthesis model provided in the embodiment of the present application, which will be combined with Figure 4 The steps shown are used to illustrate that the training method of the model provided in the embodiment of the present application can be implemented by the server or the terminal alone, or by the server and the terminal in collaboration. The following will illustrate the collaborative implementation by the server and the terminal as an example.
[0183] In step 201 , a plurality of first audio samples of a second object are fused to obtain a second audio sample.
[0184] Here, the embodiment of the present application does not impose any restrictions on the second object. The second object can be the first object or any object other than the first object. The first audio sample and the second object are audio samples emitted for different texts. The text content of each first audio sample is different. The embodiment of the present application does not impose any restrictions on the duration of the first audio sample. The duration of the first audio sample can be the same or different.
[0185] In some embodiments, step 201 may be implemented by performing a splicing process on a plurality of first audio samples of the second object to obtain a second audio sample.
[0186] For example, multiple first audio samples of the second object (such as audio sample A of the second object for the text "I eat" such as [0.28, 0.16, 0, 0.14, 0.13, -0.06, 0.15, 0.03, 0.06], and audio sample A of the second object for the text "delicious" such as [0, 0.14, -0.06, 0.15]) are spliced to obtain a second audio sample (such as [0.28, 0.16, 0, 0.14, 0.13, -0.06, 0.15, 0.03, 0.06, 0, 0.14, -0.06, 0.15]).
[0187] In step 202, audio synthesis is performed on the second audio sample and the text sample to obtain a synthesized audio sample of the second object.
[0188] Here, this application does not limit the model. The model is used to implement the audio synthesis method. The model can be an audio synthesis model to be trained. The audio synthesis model to be trained includes a first encoding layer, a second encoding layer, an attention layer, and a decoding layer.
[0189] In some embodiments, the above-mentioned "performing audio synthesis on the second audio sample and the text sample to obtain a synthesized audio sample of the second object" can be achieved in the following ways: performing a first encoding on the second audio sample through the first encoding layer to obtain the sound sample features of the second object; performing a second encoding on the second audio sample through the second encoding layer to obtain the musicality sample features of the second audio sample; performing attention processing on the musicality sample features and the text sample features through the attention layer to obtain the first sample features; performing fusion processing on the sound sample features, the first sample features and the text sample features to obtain the second sample features; and performing decoding processing on the second sample features through the decoding layer to obtain a synthesized audio sample.
[0190] In step 203, a model is trained based on the synthesized audio sample of the second object.
[0191] Here, the model is an audio synthesis model to be trained as an example for explanation. The process of training the model can be achieved in the following way: updating the parameters of the audio synthesis model to be trained, and using the updated parameters of the audio synthesis model to be trained as the parameters of the audio synthesis model.
[0192] In some embodiments, the above-mentioned "updating the parameters of the audio synthesis model to be trained based on the synthesized audio samples of the second object, and using the updated parameters of the audio synthesis model to be trained as the parameters of the audio synthesis model" can be achieved in the following way: constructing a loss function based on the synthesized audio samples and the audio sample labels; updating the parameters of the audio synthesis model to be trained until the loss function converges, and using the parameters of the audio synthesis model to be trained when the loss function converges as the parameters of the audio synthesis model.
[0193] Continuing from the above embodiment, the above-mentioned “constructing a loss function based on synthetic audio samples and audio sample labels” can be implemented in the following manner: encode the synthetic audio sample (such as [0.25, 0.25, 0.25, 0, 0, 0, 0, 1, -1, 2, -2]) to obtain a first encoding feature; encode the audio sample label (such as [1, -1, 2, -2, 1, -1, 2, -2]) to obtain a second encoding feature; construct a loss function based on the first encoding feature and the second encoding feature.
[0194] It should be noted that the audio sample label is the audio sample of the second object for the text sample. The embodiment of the present application does not limit the method of constructing the loss function. The loss function can be the mean absolute error, mean square error, etc. of the two encoding vectors.
[0195] Continuing from the above embodiment, the above-mentioned "updating the parameters of the audio synthesis model to be trained" can be achieved in the following way: performing backpropagation in the audio synthesis model to be trained based on the loss function to obtain the gradient; and updating the parameters of the audio synthesis model to be trained based on the gradient.
[0196] It should be noted that backpropagation is implemented through the backpropagation algorithm, and the gradient of the loss function with respect to the parameters of the audio synthesis model to be trained is calculated by the chain rule of derivatives.
[0197] Through the embodiments of the present application, multiple first audio samples of the second subject are trained to learn the speaking style of the first audio samples of the second subject at different times, ensuring that the synthesized audio samples have timbre consistency.
[0198] Below, an exemplary application of the audio synthesis method provided in an embodiment of the present application in a practical application scenario will be described.
[0199] In the related art, the audio of different texts of the same speaker has inconsistent timbre, and the speaker's speaking style is ignored when customizing the timbre of the target speaker.
[0200] In order to solve the above problems, the embodiment of the present application proposes an audio synthesis method that deeply integrates the text and acoustic information of the target speaker, incorporates the speaker's speaking style into the synthesized audio, and makes full use of the data, filtering out the data set with only one speaker audio and using multiple audio data for encoding to obtain a unified coding feature, ensuring timbre consistency and enriching the speaker's timbre.
[0201] Taking audio synthesis as an example, see Figure 5 , Figure 5 This is a first structural diagram of the speech synthesis model provided in an embodiment of the present application. The first process of training the speech synthesis model provided in an embodiment of the present application is described in detail below.
[0202] The model training method provided in the embodiment of the present application first filters the initial training data, retaining only data with more than or equal to 2 audios of the same speaker (i.e., the second object), and directly filtering out data with less than 2 audios. Then, the filtered training data audio (i.e., audio samples) are all processed into Mel features. The dimension of the processed Mel feature is [batch_size, mel_length, 80] as an example for illustration, where batch_size is the number of batches, which is used to represent the number of training data processed simultaneously, mel_length is the length of the Mel feature, and 80 is used to represent the dimension (dim) of the dimensional space in which the Mel feature is located is 80. The Mel feature can be obtained by processing the audio data using Mel-frequency cepstral coefficients (MFCC). The Mel feature is a feature widely used in speaker segmentation, voiceprint recognition, speech recognition, and speech synthesis. The Mel frequency is extracted based on the auditory characteristics of the human ear. The auditory characteristics of the human ear have a nonlinear correspondence with the frequency of the audio. The spectral feature (i.e., the eighth feature) is calculated from the nonlinear correspondence based on the Mel-frequency cepstral coefficients. Multiple mel features of the same speaker are concatenated to obtain a concatenated mel feature 501. The dimension of the concatenated mel feature 501 is [batch_size, 2000, 80]. The mel features used for concatenation can be randomly selected from the multiple mel features. Because the speaker (i.e., the subject, including the first subject and the second subject) may vary in mood, environment, and other factors during recording or conversation, it is difficult to maintain the uniformity of the speaker's timbre. Using the random mel features of the speaker greatly maintains the uniformity of the speaker's timbre and maintains the stability of training. When concatenating the 2000-frame mel feature 501, the mel feature 501 has a specific length (e.g., 2000). Mel features 501 with more than 2000 frames are directly truncated to 2000 frames. For those with less than 2000 frames, the concatenated mel feature 501 is repeatedly concatenated multiple times until 2000 frames are reached, and then the concatenation is stopped to obtain the final mel feature 501.
[0203] The training data provided in the embodiment of the present application is presented in the form of (<text, mel feature label, mel concatenation feature>), wherein the text corresponds to the mel feature label one-to-one, and the mel concatenation feature does not include the mel feature label (i.e., the mel feature corresponding to the input audio whose text is consistent with the input text) to prevent information leakage. The mel concatenation feature is input into the pre-trained speaker recognition model 502 for feature extraction to obtain the speaker feature, thereby ensuring the consistency of the timbre of the same speaker during the training process, wherein the dimension of the speaker feature is [batch_size, 1, dim], batch_size is the number of batches, which is used to characterize the number of speaker features processed simultaneously by the pre-trained speaker recognition model 502, and dim is used to characterize the dimension of the speaker feature in the high-dimensional space (such as 80), see Figure 6 , Figure 6 This is a schematic diagram of the speaker recognition model structure provided in an embodiment of the present application. The reasoning process of the speaker recognition model is described in detail below.
[0204] First, the Mel-valued concatenation feature is used as the input feature of the pre-trained speaker recognition model 502, and the input feature is subjected to residual processing through the residual module 5011 to obtain the residual feature, wherein the residual module 5011 includes multiple residual layers. Taking the number of residual layers as 4 as an example, the residual module 5011 includes residual layer 50111, residual layer 50112, residual layer 50113, residual layer 50114, etc. The features input to the residual layer 50111 are input features, and the features input to other residual layers are output features of the previous residual layer. The residual features are densely convolved through multiple dense convolution layers. The number of dense convolution layers is 3 for illustration. The multiple dense convolution layers include dense convolution layer 5012, dense convolution layer 5013 and dense convolution layer 5014. The dense convolution layer 5012 includes multiple neurons. The neuron 50120 is used as an example for illustration. The input feature input to the neuron 50120 is encoded through the feedforward neural network 50121 to obtain a first feedforward encoding feature. The first feedforward encoding feature is encoded through the time delay neural network 50122 to obtain a second feedforward encoding feature. The first feedforward encoding feature is respectively subjected to global pooling and segment-level pooling to obtain a global pooling feature and a segment-level pooling feature. The global pooling features and the segment-level pooling features are fused to obtain fused pooling features, the fused pooling features are activated by the activation network 50123 to obtain activation features, and the fused pooling features are mapped by the mapping network 50124 to obtain mapping features, and the mapping features and the second feedforward encoding features are fused to obtain the output features of the neuron 50120, wherein the activation network 50123 includes a first feedforward neural network and an activation layer (such as a linear mapping unit), the mapping network 50124 includes a second feedforward neural network and a mapping layer (such as a Sigmod function), and the model structures of the first feedforward neural network and the second feedforward neural network are the same as the model structure of the feedforward neural network 50121.
[0205] In the embodiment of the present application, a pre-trained speaker recognition model is used. Typically, the pre-trained speaker recognition model has training data of hundreds of thousands of speaker audio clips, and the extracted speaker features are highly accurate. Mel features are a common and frequently used speech feature in the field of speech synthesis. In the embodiment of the present application, multiple random Mel features of the same speaker are input and spliced to extract speaker features to ensure timbre consistency. The pre-trained speaker recognition model 502 takes into account both recognition performance and inference efficiency. On public Chinese and English data sets, the recognition accuracy is improved, and at the same time, it has a faster inference speed. The speaker recognition model is used to infer the input acoustic features to obtain speaker features (i.e., speech features, speaker embedding).
[0206] Continue to see Figure 5The linear mapping layer 503 performs linear mapping on the speaker features, maps the speaker features to a specific dimension, and obtains linear mapping features (ie, speech features). The dimension of the linear mapping features is [batch_size, 1, dim].
[0207] Secondly, the Mel-shaped concatenation feature 501 is encoded by a neural network model 504 (such as a residual network) to obtain acoustic feature information (i.e., musical features). The dimension of the acoustic feature information is [batch_size, 2000, dim1], where dim1 is the number of convolution kernels in the last layer of the neural network in the speaker recognition model. Figure 7 , Figure 7 This is a schematic diagram of the neural network model structure provided in the embodiment of the present application. Figure 7 As shown, the neural network model 504 includes a convolution layer, a maximum pooling layer and a fully connected layer, which performs multi-layer convolution and pooling operations on the Mel features, and the output is a rhythmic feature (i.e., a musical feature), wherein the convolution layer contains multiple convolution kernels, and each convolution layer corresponds to a number of convolution kernels.
[0208] Continue to see Figure 5 , the speaker's acoustic features are convolved through one-dimensional convolution 505, and the convolution features after convolution (i.e., the third features) are used as key features (i.e., key vectors) and value features (i.e., value vectors) of the scaled dot product attention mechanism 508, wherein the step size of the convolution kernel in the one-dimensional convolution 505 is a specific step size (e.g., 10) to reduce the amount of calculation. The one-dimensional convolution learns the local feature information of the acoustic feature information and reduces the length by a specific ratio (e.g., one tenth) to obtain the convolution feature. The one-dimensional convolution is used to scale the length of the concatenated Mel feature to 1 / 10. Since the computational complexity of the scaled dot product attention is the square of the length of the feature, scaling the length to 1 / 10 greatly reduces the amount of calculation. The dimension of the convolution feature is [batch_size, 200, dim2], and dim2 is the number of convolution kernels in the last layer of the neural network in the one-dimensional convolution.
[0209] Then, text features are obtained, wherein the text features are obtained by encoding the text (i.e., text samples) through a Byte Pair Encoder (BPE) and an embedding layer, and the text features are encoded through the encoder module 506 to learn the emotional information of the text. In the embodiment of the present application, after BPE encoding the text (i.e., the first text), the text features obtained after BPE encoding are input into the embedding layer, and the semantic information, emotional information, etc. contained in the text are learned through the embedding layer to obtain the encoded text features (i.e., the text features of the first text) to guide the content and emotion of the synthesized audio. The text can be Chinese characters, words, etc. In addition, the text is not a phoneme.
[0210] The encoded text features are convolved through the dense convolution layer 507, and the encoded convolution features are mapped to a specific dimension as the query features (i.e., query vectors) of the scaled dot product attention mechanism 508. The output dimension of the attention features output by the scaled dot product attention mechanism 508 is [batch_size, text_length, dim]. The scaled dot product attention mechanism 508 makes the acoustic information and text information pay attention to each other and influence each other, so that the text features pay attention to the related acoustic features (such as rhythm), thereby maintaining the speaker's rhythmic style characteristics when generating audio. It should be noted that rhythm is present in the speech features, but it is also closely related to the text features. See Figure 8 , Figure 8 This is a schematic diagram of the structure of the scaled dot product attention model provided in the embodiment of the present application. Figure 8 As shown, the scaled dot product attention model 508 includes N multi-head attention layers 801, a residual normalization module 802, a feedforward neural network 803, a residual normalization module 804, a linear layer 805 and a mapping layer 806. The multi-head attention layers 801 perform multi-head attention processing on the key features (i.e., key vectors), the value features (i.e., value vectors) and the query features (i.e., query vectors) to obtain multi-head attention features. The residual normalization module 802 performs residual processing and normalization processing on the multi-head attention features and the query features to obtain the first residual features. The first residual feature is encoded by the feedforward neural network 803 to obtain the feedforward encoding feature, the feedforward encoding feature and the first residual feature are residually processed and normalized by the residual normalization module 804 to obtain the second residual feature, and the second residual feature is data-processed by the linear layer 805 and the mapping layer 806 to obtain the attention feature, wherein the residual normalization module 802 includes a residual layer and a normalization layer, the residual layer is used to prevent network degradation, and the normalization layer is used to normalize the activation value obtained after the feature output by the residual layer is activated.
[0211] Finally, continue to see Figure 5 , the attention feature is added to the speaker feature to obtain the first fusion feature. The first fusion feature is related to the speaker's speaking style and rhythm. The first fusion feature is rich in feature information such as the speaker's speaking style and rhythm. The dimension of the first fusion feature is [batch_size, text_length, dim]. Then the first fusion feature is added to the text feature to obtain the second fusion feature. The dimension of the second fusion feature is [batch_size, text_length, dim]. Compared with the first fusion feature, the second fusion feature contains more text information.
[0212] See also Figure 9 , Figure 9This is a schematic diagram of the second structure of the speech synthesis model provided in an embodiment of the present application. The following is a detailed description of the process of reasoning through the second structure of the speech synthesis model provided in an embodiment of the present application.
[0213] The obtained second fusion feature is spliced with the text feature, and the spliced feature is input into the autoregressive module 901 (such as the GPT model) to obtain the hidden feature. The hidden feature is decoded by the Mel decoder 902 to obtain the predicted Mel feature. The predicted Mel audio (i.e., audio sample data) and the target Mel audio (the speech of the speaker corresponding to the text, i.e., the audio sample label) are subjected to loss calculation, and the parameters of the audio synthesis model are iteratively updated for multiple rounds based on the calculated loss until the loss converges. When the loss converges, the audio synthesis model is saved.
[0214] See also Figure 10 , Figure 10 This is a third structural diagram of the speech synthesis model provided in an embodiment of the present application. The following describes in detail the process of training the autoregressive module 901 and the Mel decoder 902 provided in an embodiment of the present application.
[0215] The true Mel feature (i.e., the audio sample label) is encoded by the quantization encoder 9031 and the quantization module 9032 in the vector quantization model 903 to obtain the quantized feature. Figure 11 , Figure 11 This is a schematic diagram of the vector quantization model structure provided in an embodiment of the present application. The vector quantization model includes a quantization encoder 9031, a quantization module 9032, and a quantization decoder 9033. The quantization encoder 9031 is used to learn useful information in the real Mel feature and obtain an intermediate feature (continuous vector feature). The quantization module 9032 converts this intermediate feature into a discrete feature and a latent intermediate feature (continuous vector feature), wherein the latent intermediate feature and the discretized feature can be converted into each other without the need for an additional network module. The quantization decoder 9033 is used to restore the latent intermediate feature to a Mel feature.
[0216] The obtained second fusion feature, text feature and quantization feature are spliced, and the spliced feature is input into the autoregressive module 901 (such as the GPT model) to obtain the hidden feature. The hidden feature is decoded by the Mel decoder 902 to obtain the first predicted Mel feature, or the hidden feature is decoded by the quantization decoder 9033 to obtain the second predicted Mel feature. The predicted Mel feature includes the first predicted Mel feature and the second predicted Mel feature. The predicted Mel audio (i.e., audio sample data) obtained by processing the predicted Mel feature through the reverberator is subjected to loss calculation with the target Mel audio (the speech of the speaker corresponding to the text, i.e., the audio sample label) to obtain the first loss, and the hidden feature is subjected to cross entropy loss with the true Mel feature to obtain the second loss. Based on the calculated first loss and second loss, the parameters of the autoregressive module 901 and the Mel decoder 902 are iteratively updated for multiple rounds until the loss converges. When the loss converges, the autoregressive module 901 and the Mel decoder 902 are saved.
[0217] See also Figure 12 , Figure 12 This is a schematic diagram of audio synthesis provided in an embodiment of the present application. The following explains the process of audio synthesis provided in an embodiment of the present application.
[0218] Load the saved model and input a 10-20 second audio clip of the target speaker and any text to synthesize audio that matches the target speaker's timbre. During inference, if the audio clip is too short, repeatedly concatenate the audio clips to obtain an audio clip of appropriate length. Text features are the features obtained by encoding any text.
[0219] First, the audio 601 of the target speaker is processed into mel features. The mel features can be repeatedly spliced multiple times to obtain mel-spliced features 602. The specific number of splicing times is determined by the length of the mel features. The pre-trained speaker recognition model 603 extracts features from the mel-spliced features 602 to obtain speaker features.
[0220] Secondly, the Mel-scale concatenation feature 602 is encoded through a neural network model 604 (such as a residual network) to obtain the acoustic features of the speaker, and the acoustic features of the speaker are convolved through a one-dimensional convolution 605 to obtain the key features and value features of the scaled dot product attention mechanism 606, wherein the step size of the convolution kernel in the one-dimensional convolution 605 is a specific step size (such as 10).
[0221] Then, the text features are obtained, and the text features are encoded through the dense convolution layer 607. The encoded convolution text features are used as query features of the scaled dot product attention mechanism 606. The query features, key features and value features are paid attention to by the scaled dot product attention mechanism 606 to obtain attention example features. The attention features are added to the speaker features to obtain the first fusion features, and then the enriched second fusion features are obtained by adding the attention features to the text features. The third fusion features obtained by fusing the second fusion features with the text features are input into the autoregressive module 608 (such as the GPT model, the GPT model can mask the multi-head attention module) to obtain hidden features. The hidden features are decoded by the Mel decoder 609 to obtain predicted Mel features. The predicted Mel features are restored to sound by the vocoder 610 to obtain predicted Mel audio.
[0222] When synthesizing audio, the effective autoregressive module 608 is applied to the user interaction field. By optimizing the text features input to the autoregressive module 608, the synthesized speech not only maintains the target speaker's speaking style, such as rhythm, but also improves the speaker similarity, reaching a level that is indistinguishable from the real thing.
[0223] When applied in the field of user interaction, audio 601 can be a 10-20 second audio recording of an agent whose voice is suitable for the scene and has strong interactive ability. By inputting audio 601 into the saved model, the agent's voice can be synthesized. After obtaining the right to use the agent's voice, corresponding synthesized audio is generated for different interactive texts through the audio recorded by the agent, so as to facilitate interaction with the user based on the synthesized audio, thereby effectively improving the similarity between the timbre and speaking style of the target speaker and the synthesized audio.
[0224] In practical applications, a vector quantization model is first trained, wherein the encoder and quantization module of the vector quantization model are used to extract quantitative features. Then, the fused features obtained by fusing text features and Mel-shaped splicing features using the solution proposed in the embodiment of the present application are spliced with the quantization features, and the spliced features are used to train the autoregressive module 608. When used, the input text (i.e., the first text) and a 10-20s audio of the target speaker (i.e., audio data) can be used to synthesize speech (i.e., synthesized audio data) that is highly similar to the timbre and speaking style of the target speaker. This can be used in user interaction scenarios. In the user interaction scenario, a voice stream signal is first received in real time from the telephone user end, and then the voice stream signal is input into the speech recognition model to obtain the user's speech content. After speech recognition is performed by the natural language processing model, the user's intention is obtained and the text content to be fed back (i.e., the first text) is obtained. The fed-back text content is subjected to audio synthesis by the audio synthesis method provided in the embodiment of the present application to obtain synthesized speech (i.e., synthesized audio data). The synthesized speech is fed back to the customer via the phone for response, completing the interaction with the user.
[0225] In summary, the similarity between the timbre and speaking style of the target speaker and the synthesized audio in the zero-sample voice cloning scenario can be effectively improved. The timbre of the agent can be customized in the outbound call scenario. With the consent of the agent, a 10-20s audio of the agent is recorded and input into the trained model to create a cloned timbre model of the agent. The agent's timbre can then be used for outbound calls. By deeply integrating the speaker text information and acoustic information, the two interact and influence each other, and the speaker's speaking style and other acoustic information are integrated into the speaker text. Considering that the speaker's speaking style may be inconsistent at different times, multiple Mel features are processed during training to obtain speaker features to ensure the consistency of the timbre, so that when used, a fake-to-real effect can be achieved.
[0226] The following continues to describe the exemplary structure of the audio synthesis device 555 provided in the embodiment of the present application implemented as a software module. In some embodiments, such as Figure 2A As shown, the software modules stored in the audio synthesis device 555 of the memory 550 may include:
[0227] The data collection module 5551 is used to determine the text features of the first text and the audio data of the first object.
[0228] The data processing module 5552 is used to perform a first encoding on the audio data to obtain the speech features of the first object, and to perform a second encoding on the audio data to obtain the musical features of the audio data; to perform attention processing on the musical features and text features to obtain the first features; and to perform fusion processing on the speech features, the first features and the text features to obtain the second features.
[0229] The audio synthesis module 5553 is used to decode the second feature to obtain synthesized audio.
[0230] In some embodiments, the data processing module 5552 is also used to normalize the musical features to obtain the third feature, and normalize the text features to obtain the fourth feature; perform multi-head attention processing on the third feature and the fourth feature to obtain the fifth feature; and determine the first feature based on the third feature and the fifth feature.
[0231] In some embodiments, the data processing module 5552 is also used to determine the key vector of the audio data and the value vector of the audio data based on the third feature, and determine the query vector of the first text based on the fourth feature; determine the similarity matrix based on the query vector and the key vector; and determine the product of the similarity matrix and the value vector as the fifth feature.
[0232] In some embodiments, the data processing module 5552 is further used to fuse the third feature and the fifth feature to obtain the sixth feature, and map the sixth feature to obtain the seventh feature; and fuse the sixth feature and the seventh feature to obtain the first feature.
[0233] In some embodiments, the data processing module 5552 is also used to perform cepstrum processing on the audio data to obtain the eighth feature; perform residual processing on the eighth feature to obtain the ninth feature; perform context mask processing on the ninth feature to obtain the tenth feature; and perform dense convolution processing on the tenth feature to obtain the speech feature of the first object.
[0234] In some embodiments, the data processing module 5552 is also used to perform cepstrum processing on the audio data to obtain an eighth feature; perform convolution processing on the eighth feature, perform pooling processing on the convolved eighth feature to obtain an eleventh feature; perform mapping processing on the eleventh feature, and perform scaling processing on the mapped eleventh feature to obtain the musical characteristics of the audio data.
[0235] In some embodiments, the data processing module 5552 is also used to perform frequency domain transformation on the audio data to obtain an audio spectrum; convert the audio spectrum into a first power spectrum, and filter the first power spectrum to obtain a second power spectrum; perform cosine transformation on the logarithm of the second power spectrum to obtain the eighth feature.
[0236] In some embodiments, the audio synthesis module 5553 is also used to perform prediction processing on the speech feature and the second feature to obtain a twelfth feature; perform mapping processing on the twelfth feature and up-sample the mapped twelfth feature to obtain a thirteenth feature; perform deconvolution processing on the thirteenth feature to obtain synthesized audio.
[0237] The following continues to describe the exemplary structure of the audio synthesis model training device 556 provided in the embodiment of the present application implemented as a software module. In some embodiments, such as Figure 2B As shown, the software modules stored in the audio synthesis model training device 556 of the memory 550 may include:
[0238] The pre-processing module 5561 is configured to fuse multiple first audio samples of the second object to obtain a second audio sample.
[0239] The model training module 5562 is used to perform audio synthesis on the second audio sample and the text sample based on the model to obtain a synthesized audio sample of the second object; and train the model based on the synthesized audio sample of the second object.
[0240] An embodiment of the present application provides a computer program product, which includes computer-executable instructions. The computer-executable instructions are stored in a computer-readable storage medium. A processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium, and the processor executes the computer-executable instructions, so that the electronic device executes the audio synthesis method described above in the embodiment of the present application.
[0241] The embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions or computer programs are stored. When the computer-executable instructions or computer programs are executed by a processor, the processor will be caused to execute the audio synthesis method provided in the embodiment of the present application, for example, Figures 3A to 3E The audio synthesis method shown or Figure 4 The training method of the audio synthesis model is shown.
[0242] In some embodiments, the computer-readable storage medium may be a memory such as RAM, ROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or may be various devices including one or any combination of the above memories.
[0243] In some embodiments, computer-executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.
[0244] As an example, computer-executable instructions may, but need not, correspond to a file in a file system, may be stored as part of a file that stores other programs or data, such as in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple coordinating files (e.g., files storing one or more modules, subroutines, or code portions).
[0245] By way of example, computer-executable instructions may be deployed to be executed on one electronic device, or on multiple electronic devices located at one site, or on multiple electronic devices distributed across multiple sites and interconnected by a communication network.
[0246] To sum up, the audio data is respectively subjected to the first encoding and the second encoding to obtain the speech features of the first object and the musical features of the audio data, and then the musical features and the text features are subjected to attention processing to obtain the first features, and the speech features, the first features and the text features are fused to obtain the second features. In this way, the speech features provide sound information related to the first object, and the text features provide semantic content. By integrating multimodal information such as musical features, speech features, and text features, the performance of audio synthesis is improved to obtain synthesized audio that conforms to the sound of the first object, and the text content of the synthesized audio is consistent with the text content of the first text, thereby improving the accuracy of audio synthesis.
[0247] The above description is merely an embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are included in the scope of protection of the present application.
Claims
1. An audio synthesis method, characterized in that: The method comprises: determining text features of a first text and audio data of a first object; Performing a first encoding on the audio data to obtain a speech feature of the first object, and performing a second encoding on the audio data to obtain a musical feature of the audio data; Normalizing the musicality feature to obtain a third feature, and normalizing the text feature to obtain a fourth feature; Performing multi-head attention processing on the third feature and the fourth feature to obtain a fifth feature; Determining a first feature based on the third feature and the fifth feature; fusing the speech feature, the first feature, and the text feature to obtain a second feature; The second feature is decoded to obtain synthesized audio.
2. The method according to claim 1, characterized in that The multi-head attention processing is performed on the third feature and the fourth feature to obtain the fifth feature, including: determining a key vector of the audio data and a value vector of the audio data based on the third feature, and determining a query vector of the first text based on the fourth feature; determining a similarity matrix of the first text based on the query vector and the key vector; The fifth feature is obtained according to the similarity matrix and the value vector.
3. The method according to claim 1, characterized in that The determining the first feature based on the third feature and the fifth feature includes: Fusing the third feature and the fifth feature to obtain a sixth feature; Mapping the sixth feature to obtain a seventh feature; The sixth feature and the seventh feature are fused to obtain the first feature.
4. The method according to claim 1, wherein The performing a first encoding on the audio data to obtain the speech feature of the first object includes: Performing cepstrum processing on the audio data to obtain an eighth feature; Performing residual processing on the eighth feature to obtain a ninth feature; Performing context mask processing on the ninth feature to obtain a tenth feature; Perform dense convolution processing on the tenth feature to obtain a speech feature of the first object.
5. The method according to claim 1, wherein The performing a second encoding on the audio data to obtain the musical characteristics of the audio data includes: Performing cepstrum processing on the audio data to obtain an eighth feature; Performing convolution processing on the eighth feature, and performing pooling processing on the convolved eighth feature to obtain an eleventh feature; Mapping processing is performed on the eleventh feature, and scaling processing is performed on the mapped eleventh feature to obtain the musicality feature of the audio data.
6. The method according to claim 4 or 5, characterized in that The audio data is subjected to cepstrum processing to obtain an eighth feature, including Performing frequency domain transformation on the audio data to obtain an audio spectrum; Converting the audio spectrum into a first power spectrum, and filtering the first power spectrum to obtain a second power spectrum; Perform cosine transform on the logarithm of the second power spectrum to obtain the eighth feature.
7. The method according to claim 1, characterized in that The decoding process of the second feature to obtain synthesized audio includes: performing prediction processing on the speech feature and the second feature to obtain a twelfth feature; Performing mapping processing on the twelfth feature, and performing upsampling processing on the mapped twelfth feature to obtain a thirteenth feature; Deconvolution processing is performed on the thirteenth feature to obtain the synthesized audio.
8. An electronic device, characterized in that: The electronic device comprises: a memory for storing computer-executable instructions; The processor is configured to implement the audio synthesis method according to any one of claims 1 to 7 when executing the computer executable instructions or computer program stored in the memory.
9. A computer-readable storage medium storing computer-executable instructions or a computer program, characterized in that: When the computer executable instructions or computer program are executed by a processor, the audio synthesis method according to any one of claims 1 to 7 is implemented.
10. A computer program product comprising computer executable instructions or a computer program, characterized in that When the computer executable instructions or computer program are executed by a processor, the audio synthesis method according to any one of claims 1 to 7 is implemented.