Audio timbre conversion method and device, electronic equipment and storage medium

By reversely generating target audio data of target tones, the problem of poor audio generation quality of target tones in the target language in the prior art is solved, and the consistency and naturalness of tone are improved.

CN120032664APending Publication Date: 2025-05-23TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311573139.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-22
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

When the existing speech synthesis model generates audio of the target tone in the specified language, there are problems such as inconsistent mixing effects, inconsistent tone and poor audio quality.

Method used

By determining the target language and target tone, the tone characteristics of the target tone are obtained, and based on the audio data of the target language and reference tone, the content features, fundamental frequency features and phonological characteristics are extracted, and the target audio data of the target tone is reversely generated.

Benefits of technology

The audio generation quality of the speech synthesis model in the target language is improved, ensuring the consistency and naturalness of tone, and solving the problem of scarcity of audio data in the target language.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032664A_ABST
    Figure CN120032664A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, in particular to an audio timbre conversion method and device, electronic equipment and a storage medium, and the method comprises the steps: determining a target language and a target timbre, obtaining the timbre feature of the target timbre, and obtaining the reference audio data based on the target language and recorded by adopting the reference timbre; based on the reference audio data, extracting content features, and based on the reference audio data, extracting corresponding fundamental frequency features and phonetic rhyme features; and reversely generating target audio data of the target timbre based on the timbre feature, the content feature, the fundamental frequency feature and the rhyme feature of the target timbre. Thus, in the process of reversely generating the target audio data, the pronunciation tone and rhythm in the reference audio data can be restored with high fidelity, the pronunciation effect based on the target language in the target audio data is guaranteed, and the generation quality of the audio data is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0002] With the development of information technology, with the help of speech synthesis technology, when the target object with the target timbre has difficulty in reading the content in the target language, audio reading the content in the target language with the target timbre can be generated.

[0003] At present, in order to generate audio in the target language according to the target timbre, a speech synthesis model is usually trained based on the audio data of the target object in its native language and the audio data of other objects whose native language is the target language, so that the speech synthesis model can learn the target timbre of the target object and the reading ability in the target language.

[0004] However, since the speech synthesis model is trained with the help of audio data of other subjects whose native language is the specified language when learning to generate audio of the target timbre in a specified language, the generated audio has an obvious inharmonious mixed reading effect for the content in different languages ​​read aloud with the target timbre, the audio content in different languages ​​has obvious inconsistencies in timbre, the played audio has a strong sense of unnaturalness, and the audio quality is very poor.

[0005] Based on this, it can be seen that it is necessary to transform the audio data of the target timbre in the target language to better assist the training of the speech synthesis model, thereby improving the audio generation effect of the speech synthesis model. Summary of the invention

[0006] The embodiments of the present application provide a method, device, electronic device and storage medium for transforming audio timbre, which are used to generate audio data that meets the needs of model learning.

[0007] In a first aspect, a method for changing audio timbre is proposed, comprising:

[0008] Determining a target language and a target timbre, and obtaining timbre characteristics of the target timbre;

[0009] Acquiring reference audio data based on the target language and recorded using a reference timbre;

[0010] Extracting content features based on the text content identified from the reference audio data, and extracting corresponding fundamental frequency features and phonological features based on the fundamental frequency information and phonological identification information identified from the reference audio data;

[0011] Based on the timbre feature of the target timbre, the content feature, the fundamental frequency feature, and the phonological feature, target audio data of the target timbre is inversely generated.

[0012] In a second aspect, a device for changing audio timbre is provided, comprising:

[0013] A determination unit, used to determine a target language and a target timbre, and obtain a timbre feature of the target timbre;

[0014] An acquisition unit, configured to acquire reference audio data based on the target language and recorded using a reference timbre;

[0015] An extraction unit, configured to extract content features based on the text content identified from the reference audio data, and to extract corresponding fundamental frequency features and phonological features based on the fundamental frequency information and phonological identification information identified from the reference audio data;

[0016] A generating unit is used to reversely generate target audio data of the target timbre based on the timbre feature of the target timbre, the content feature, the fundamental frequency feature, and the phonological feature.

[0017] Optionally, when extracting corresponding fundamental frequency features and phonological features based on the fundamental frequency information and phonological identification information obtained by identifying the reference audio data, the extraction unit is used to:

[0018] Identifying fundamental frequency information in the reference audio data, and based on the fundamental frequency information, obtaining phonological identification information for indicating unvoiced sounds and voiced sounds in the reference audio data, and extracting phonological features of the parameter audio data based on the phonological identification information;

[0019] Determine a comprehensive value range corresponding to each information content in the baseband information, and re-assign a value to the information content associated with the unvoiced sound in the baseband information based on the principle of information content continuity, and reset the content of each information content in the re-assigned baseband information according to the comprehensive value range and the total number of preset value categories;

[0020] According to resetting the baseband information of each information content, the corresponding baseband feature is extracted.

[0021] Optionally, when re-assigning information content associated with unvoiced sound in the fundamental frequency information, the extraction unit is used to:

[0022] For the endpoint region associated with the unvoiced sound in the fundamental frequency information, the information content in the endpoint region is reassigned by using the information content that is closest to the endpoint region and corresponds to the voiced sound;

[0023] For the middle area associated with the unvoiced sound in the fundamental frequency information, the content positions covered by the middle area and the information contents of the voiced sounds corresponding to the two ends are determined, and the information contents at the content positions are reassigned according to the information contents of the voiced sounds corresponding to the two ends.

[0024] Optionally, when resetting the content according to the comprehensive value range corresponding to each information content and the total number of preset value categories, the extraction unit is used to:

[0025] According to the total number of preset value categories, the comprehensive value range is divided into a corresponding number of sub-value ranges, wherein one sub-value range corresponds to one category label;

[0026] For each of the information contents, the following operations are performed respectively: a sub-value range corresponding to an information content is determined, and the information content is reset to a classification label corresponding to the sub-value range.

[0027] Optionally, the timbre feature of the target timbre is extracted by a training unit in the device in the following manner:

[0028] Acquire a sample data set; a sample data includes: a sample audio data based on any language and recorded using a sample timbre; each sample timbre covered by the sample data set includes a target timbre;

[0029] The sample data set is used to perform multiple rounds of iterative training on the initial timbre characterization network and the initial audio generation network in the initial timbre conversion model to obtain the trained target timbre conversion model and the timbre features extracted for each sample timbre within the target timbre conversion model.

[0030] Optionally, during a round of iterative training, the training unit is used to perform the following operations:

[0031] Extracting sample content features based on the text content identified from the sample audio data read, extracting corresponding sample fundamental frequency features and sample phonological features based on the fundamental frequency information and phonological identification information identified from the sample audio data, and extracting sample timbre features based on the timbre information identified from the sample audio data using the initial timbre representation network;

[0032] Using the initial audio generation network, based on the sample timbre feature, the sample content feature, the sample fundamental frequency feature, and the sample phonological feature, reversely synthesize and predict the audio data;

[0033] Based on the frequency spectrum difference between the predicted audio data and the sample audio data, network parameters of the initial audio generation network and the initial timbre characterization network are adjusted.

[0034] Optionally, when the initial timbre characterization network is used to extract sample timbre features based on timbre information obtained by identifying the sample audio data, the training unit is used to:

[0035] The initial timbre characterization network is used to obtain a corresponding sample spectrogram based on the sample audio data, and the spectrum content in the sample spectrogram is reorganized with reference to the time dimension, and the sample timbre features are extracted from the processed sample spectrogram.

[0036] Optionally, after obtaining the trained target timbre conversion model, the training unit is further used to:

[0037] Acquire a fine-tuning sample set, wherein a fine-tuning sample includes: a fine-tuning audio data recorded based on the target language and using the sample timbre, or a fine-tuning audio data recorded based on another language and using the target timbre;

[0038] The fine-tuning sample set is used to perform multiple rounds of fine-tuning training on the target timbre representation network and the target audio generation network in the target timbre conversion model to obtain the fine-tuned final timbre conversion model and the timbre features extracted for the target timbre in the final timbre conversion model.

[0039] Optionally, when acquiring reference audio data based on the target language and recorded with a reference timbre, the acquiring unit is used to:

[0040] Arbitrarily acquiring audio data recorded using the target timbre, and obtaining a target average pitch corresponding to the target timbre for the audio data;

[0041] In a preset audio database, select candidate audio data based on the target language, and obtain corresponding candidate average pitches for each candidate audio data;

[0042] Among the candidate average pitches, a reference average pitch whose similarity with the target average pitch meets the set conditions is screened out, and the candidate audio data corresponding to the reference average pitch is determined as the obtained reference audio data recorded using the reference timbre.

[0043] Optionally, after the target audio data of the target timbre is inversely generated, the training unit in the device is further used to:

[0044] Acquire each target audio data in the target language generated for the target timbre, and acquire each historical audio data based on other languages ​​and corresponding to the target timbre;

[0045] Based on the target audio data and the historical audio data, generate speech synthesis samples respectively, wherein a speech synthesis sample includes: the timbre feature of the target timbre, a target audio data and a corresponding sample text, or the timbre feature of the target timbre, a historical audio data and a corresponding sample text;

[0046] The speech synthesis samples are used to perform multiple rounds of iterative training on the constructed initial speech synthesis model to obtain a trained target speech synthesis model.

[0047] Optionally, after obtaining the trained target speech synthesis model, the training unit is further used to:

[0048] In response to a request for reading aloud the target text data, synthesizing corresponding audio to be played using the target speech synthesis model based on text features extracted from the target text data and timbre features of the target timbre;

[0049] The audio to be played is played to provide the target text data read aloud by the target tone.

[0050] In a third aspect, an electronic device is proposed, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above method when executing the computer program.

[0051] In a fourth aspect, a computer-readable storage medium is proposed, on which a computer program is stored, and the computer program implements the above method when executed by a processor.

[0052] In a fifth aspect, a computer program product is proposed, comprising a computer program, wherein the computer program implements the above method when executed by a processor.

[0053] The beneficial effects of this application are as follows:

[0054] In the embodiments of the present application, a method, device, electronic device and storage medium for transforming audio timbre are proposed, a target language and a target timbre are determined, and the timbre features of the target timbre are obtained; then reference audio data based on the target language and recorded with a reference timbre are obtained; then, content features are extracted based on the text content identified with respect to the reference audio data, and corresponding fundamental frequency features and phonological features are extracted based on the fundamental frequency information and phonological identification information identified with respect to the reference audio data; then, based on the timbre features, content features, fundamental frequency features, and phonological features of the target timbre, the target audio data of the target timbre are reversely generated.

[0055] In this way, in the process of generating the target audio data of the target timbre, by extracting the text features, fundamental frequency features and phonological features from the reference audio data of the reference timbre, the reference audio data can be analyzed from two levels: the audio text and the pronunciation situation, and the pronunciation method when reading the audio text based on the target language can be analyzed, so that in the subsequent reverse generation of the target audio data, the pronunciation tone and rhythm in the reference audio data can be restored with high fidelity, the pronunciation effect based on the target language in the target audio data can be guaranteed, and the generation quality of the audio data can be improved;

[0056] Furthermore, since the target audio data generated in reverse is generated under the comprehensive effect of timbre features, content features, fundamental frequency features, and phonological features, this is equivalent to applying the influence of the timbre features of the target timbre to the pronunciation method represented by the content features, fundamental frequency features, and phonological features, so that the generated target audio data can express the effect of reading the audio text based on the target language using the target timbre; it is equivalent to transforming the audio timbre of the reference audio data, that is, realizing the conversion of any person's reference audio data based on the target language into target audio data of the target timbre; in addition, since the target audio data in the target language can be generated for the target timbre, the problem of poor speech synthesis quality caused by the target timbre having less audio data in the target language is solved, and data that meets the needs of model training can be generated. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] Figure 1 A schematic diagram of possible application scenarios in the embodiments of the present application;

[0058] Figure 2A Schematic diagram of the audio timbre transformation process in the embodiment of the present application;

[0059] Figure 2B Schematic diagram of the process of training and obtaining the target timbre conversion model in the embodiment of the present application;

[0060] Figure 2C A schematic diagram of a process of splitting a plurality of sample audio data from one audio data in an embodiment of the present application;

[0061] Figure 3A A schematic diagram of the structure of an initial timbre conversion model constructed in an embodiment of the present application;

[0062] Figure 3B A schematic diagram of the structure of another initial timbre conversion model constructed in an embodiment of the present application;

[0063] Figure 4A This is a schematic diagram of the process of a round of iterative training in an embodiment of the present application;

[0064] Figure 4BSchematic diagram of the process of reassigning information content related to voiceless sounds in the embodiments of the present application;

[0065] Figure 4C Schematic diagram of the process of resetting the content of fundamental frequency information in the embodiments of the present application;

[0066] Figure 5 Schematic diagram of the process of training a target speech synthesis model in the embodiments of the present application;

[0067] Figure 6 Schematic diagram of the process of training an initial speech synthesis model in the embodiments of the present application;

[0068] Figure 7 Schematic diagram of the logical structure of the audio timbre transformation device in the embodiments of the present application;

[0069] Figure 8 Schematic diagram of the hardware composition structure of an electronic device applying the embodiments of the present application. Detailed implementation manners

[0070] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some but not all of the embodiments of the technical solutions of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments recorded in this application document without creative efforts belong to the scope protected by the technical solutions of the present application.

[0071] The terms "first", "second", etc. in the specification, claims, and the above-mentioned drawings of the present application are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here.

[0072] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the function of the module or unit.

[0073] The following explains some terms in the embodiments of the present application to facilitate the understanding of those skilled in the art.

[0074] Voice Conversion (VC): also known as timbre conversion, voice conversion, or voice change. In the embodiment of the present application, by performing voice conversion, the audio of any given speaker can be converted into the voice of another speaker, and the tone, emotion and other information in the audio can be retained.

[0075] Intelligent voice changing: In the embodiments of the present application, it refers to the technology based on artificial intelligence to achieve the effect of voice changing.

[0076] Speech synthesis: refers to text-to-speech technology. Given a piece of text, the text is broadcast out through an artificial intelligence (AI) algorithm.

[0077] Reference audio data: also known as source speaker audio, refers to the input audio of intelligent voice changing. In the embodiment of the present application, it refers to audio data selected from a preset audio database and recorded based on the target language using a reference timbre, which can specifically be real recorded corpus.

[0078] Target audio data: in the embodiment of the present application, it refers to the audio data of the target timbre obtained after the audio timbre is transformed.

[0079] End-to-end training: In the embodiment of the present application, it refers to a training method in which the input and output of the model are audio data respectively, and the training process does not need to be divided into multiple stages.

[0080] Fundamental Frequency (F0): refers to the frequency of vocal cord vibration when making voiced sounds, measured in Hertz (Hz). Usually the fundamental frequency ranges from 80 to 450 Hz, and the fundamental frequency of boys' speech is lower than that of girls and children's speech.

[0081] Mel filterbanks: used to filter the power spectrum after Fourier transforming the audio data to obtain the power spectrum.

[0082] Speaker Embedding: refers to the vector used to distinguish the identity or timbre of different speakers.

[0083] Pitch Parameters: refers to the features extracted based on fundamental frequency information.

[0084] Zero-sample multi-language mixed reading: refers to the training corpus in only one language for the target voice in speech synthesis. In this case, the target voice needs to be trained to have the ability to read another language.

[0085] The following is a brief introduction to the design concept of the embodiment of the present application:

[0086] In related technologies, when processing zero-sample multi-language mixed reading tasks, speech synthesis technology is usually used to synthesize corresponding audio data based on given text through AI algorithms.

[0087] For example, a speech synthesis system based on the VITS architecture, where VITS is an end-to-end speech synthesis system based on an adversarial learning framework, which can directly synthesize text into corresponding audio. The training of the speech synthesis model requires a certain amount of recorded corpus.

[0088] Regarding the speech synthesis process, the applicant has found that mixed reading of multiple languages ​​is a major difficulty of the speech synthesis system, especially when the target voice does not have the corpus of the target language. How to make the target voice have fluent and natural multi-language reading ability is an urgent problem to be solved.

[0089] For example, for a target timbre X, the target object with the target timbre X can only read Chinese but not English or the cost of reading English is relatively high; in this case, if a speech synthesis system for the target timbre X is required so that the target timbre X can read both Chinese and English, it is difficult to obtain training samples that meet the requirements.

[0090] During the process of conception, the applicant thought of obtaining a reference voice Y, where the reference voice Y has a large amount of recorded corpus in the target language; at this time, the native language corpus of the target voice X and the target language corpus of the reference voice Y can be mixed together for training, and their respective speaker identities are given in the model. When the model training is completed, the target voice X will have a certain ability to read the target language.

[0091] However, since the reference timbre Y is generally that of a speaker whose native language is the target language, this will lead to obvious inconsistency in the synthesized multi-language mixed reading effect. When switching between the two languages, the audio content of different languages ​​will have obvious inconsistencies in timbre and lack of timbre similarity, giving off a very unnatural feeling and very poor audio quality.

[0092] In view of this, in the embodiments of the present application, a method, device, electronic device and storage medium for transforming audio timbre are proposed, a target language and a target timbre are determined, and the timbre features of the target timbre are obtained; then reference audio data based on the target language and recorded with a reference timbre is obtained; thereafter, content features are extracted based on the text content identified for the reference audio data, and corresponding fundamental frequency features and phonological features are extracted based on the fundamental frequency information and phonological identification information identified for the reference audio data; then, based on the timbre features, content features, fundamental frequency features, and phonological features of the target timbre, the target audio data of the target timbre is reversely generated.

[0093] In this way, in the process of generating the target audio data of the target timbre, by extracting the text features, fundamental frequency features and phonological features from the reference audio data of the reference timbre, the reference audio data can be analyzed from two levels: the audio text and the pronunciation situation, and the pronunciation method when reading the audio text based on the target language can be analyzed, so that in the subsequent reverse generation of the target audio data, the pronunciation tone and rhythm in the reference audio data can be restored with high fidelity, the pronunciation effect based on the target language in the target audio data can be guaranteed, and the generation quality of the audio data can be improved;

[0094] Furthermore, since the target audio data generated in reverse is generated under the comprehensive effect of timbre features, content features, fundamental frequency features, and phonological features, this is equivalent to applying the influence of the timbre features of the target timbre to the pronunciation method represented by the content features, fundamental frequency features, and phonological features, so that the generated target audio data can express the effect of reading the audio text based on the target language using the target timbre; it is equivalent to transforming the audio timbre of the reference audio data, that is, realizing the conversion of any person's reference audio data based on the target language into target audio data of the target timbre; in addition, since the target audio data in the target language can be generated for the target timbre, the problem of poor speech synthesis quality caused by the target timbre having less audio data in the target language is solved, and data that meets the needs of model training can be generated.

[0095] The preferred embodiments of the present application are described below in conjunction with the drawings in the specification. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present application, and are not used to limit the present application. In addition, the embodiments of the present application and the features in the embodiments may be combined with each other if there is no conflict.

[0096] See also Figure 1 As shown, it is a schematic diagram of a possible application scenario in an embodiment of the present application. The schematic diagram of the application scenario includes a client device 110 and a processing device 120.

[0097] In some feasible embodiments of the present application, the processing device 120 can determine the target language and audio text selected by the target object in response to the audio generation request triggered by the target object on the client device 110; thereafter, the processing device 120 obtains the timbre characteristics of the target object, and then obtains reference audio data recorded based on the target language using a reference timbre, wherein the reference audio data can express the effect of reading the audio text based on the target language using the reference timbre; and then, based on the features extracted from the reference audio data and the timbre characteristics of the target timbre, reversely generate the target audio data, wherein the generated target audio data can express the effect of reading the audio text based on the target language using the target timbre.

[0098] In some other feasible embodiments of the present application, the processing device 120 can obtain the timbre characteristics and target language of the target timbre according to actual business processing needs, and then obtain reference audio data of the reference timbre recorded based on the target language; then, based on the features extracted from the reference audio data and the timbre characteristics of the target timbre, reversely generate the target audio data of the target timbre based on the target language.

[0099] The client device 110 includes but is not limited to mobile phones, tablet computers, notebooks, e-book readers, intelligent voice interaction devices, smart home appliances, car terminals, etc.

[0100] The processing device 120 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), as well as big data and artificial intelligence platforms. In a feasible implementation method, the processing device can be a terminal device with certain processing capabilities, such as a tablet computer and a notebook.

[0101] In the embodiment of the present application, the client device 110 and the processing device 120 can communicate through a wired network or a wireless network. The following description only describes the audio timbre conversion process from the perspective of the processing device 120.

[0102] The following describes the audio timbre transformation process in combination with possible application scenarios:

[0103] Application scenario 1: In the text reading scenario, generate target audio data with target timbre in the target language.

[0104] Specifically, in scenarios where text reading is required, such as news reading and novel reading, the processing device first determines the target timbre used for text reading, then determines the native language of the target object who has the target timbre, and determines the target language corresponding to the text content that needs to be read aloud; further, when the target language is not the native language of the target object, there may be no audio recorded in the target language with the target timbre, or there may be less audio recorded in the target language with the target timbre.

[0105] Based on this, the processing device obtains reference audio data recorded based on the target language of the reference timbre, and analyzes the content features of the corresponding text content based on the reference audio data, as well as the fundamental frequency features and phonetic features used to characterize the pronunciation method, and then combines the extracted content features, fundamental frequency features and phonetic features, as well as the timbre features of the target timbre to reversely generate the target audio data of the target language based on the target timbre.

[0106] Application scenario 2: In the voice broadcast scenario, generate target audio data with target timbre in the target language.

[0107] Specifically, in text broadcasting scenarios such as smart assistants, smart motorcycles, telephone customer service, digital humans, virtual assistants, etc., the processing device first determines the target timbre used for the text broadcast, then determines the native language of the target object with the target timbre, and determines the target language corresponding to the text content to be read aloud; further, when the target language is not the native language of the target object, there may be no audio recorded in the target language with the target timbre, or there may be less audio recorded in the target language with the target timbre.

[0108] Based on this, the processing device obtains reference audio data recorded based on the target language of the reference timbre, and analyzes the content features of the corresponding text content based on the reference audio data, as well as the fundamental frequency features and phonetic features used to characterize the pronunciation method, and then combines the extracted content features, fundamental frequency features and phonetic features, as well as the timbre features of the target timbre to reversely generate the target audio data of the target language based on the target timbre.

[0109] Application scenario three: in the audio voice change scenario, generate target audio data with target timbre in target language.

[0110] Specifically, for the purpose of increasing the interest of video generation, the processing device can respond to the target object's demand for timbre transformation between different languages, obtain timbre features extracted for the target object's target timbre, and then determine the target language that the target object expects to read aloud.

[0111] Furthermore, the processing device obtains reference audio data recorded based on the target language using a reference timbre, and analyzes the content features of the corresponding text content based on the reference audio data, as well as the fundamental frequency features and phonological features used to characterize the pronunciation method, and then combines the extracted content features, fundamental frequency features and phonological features, as well as the timbre features of the target timbre to reversely generate the target audio data based on the target language using the target timbre.

[0112] In addition, it should be understood that in the specific implementation of the present application, it involves the transformation and processing of the target timbre. When the embodiments recorded in the present application are applied to specific products or technologies, it is necessary to obtain the permission or consent of the relevant objects that own the target timbre, and the collection, use and processing of relevant data need to comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0113] The following is an explanation of the audio timbre transformation process from the perspective of the processing device in conjunction with the accompanying drawings:

[0114] See also Figure 2AAs shown in the figure, it is a schematic diagram of the audio timbre transformation process in the embodiment of the present application. Figure 2A , the audio timbre transformation process is described:

[0115] Step 201: The processing device determines a target language and a target timbre, and obtains a timbre feature of the target timbre.

[0116] In the embodiment of the present application, when implementing the transformation of the audio timbre, the processing device needs to first determine the target timbre after the timbre transformation, the target language on which the timbre transformation process is based, and obtain the timbre features extracted for the target timbre.

[0117] It should be noted that in different business scenarios, there are different relationships between the target language and the target timbre. For example, the target language may be a language that the target object with the target timbre cannot read fluently; or the target language may be a language that the target object with the target timbre has a high reading cost, where the target object with the target timbre refers to a target object that can produce the target timbre.

[0118] Moreover, when the processing device specifically determines the target timbre, it may first determine the target object, and then determine the sound timbre of the target object as the target timbre; or, it may determine the sound timbre suitable for business processing as the target timbre.

[0119] For example, assuming that object A is selected as the reader of electronic book X, the timbre of object A may be determined as the target timbre.

[0120] For another example, assuming that after screening, it is determined that timbre X is suitable for reading electronic book X, timbre X can be determined as the target timbre.

[0121] In the embodiments of the present application, in some feasible implementations, when the processing device obtains the timbre features of the target timbre, it can train the model to learn the ability to extract timbre features from audio data and learn the ability to reversely regenerate audio data, so that the timbre features extracted by the trained model for the target timbre can be finally obtained; in other feasible implementations, the processing device can obtain the timbre features extracted for the target timbre from other devices.

[0122] The following takes the example of a processing device obtaining the timbre characteristics of a target timbre through model training to illustrate the relevant acquisition process:

[0123] See also Figure 2B As shown in the figure, it is a schematic diagram of the process of training to obtain the target timbre conversion model in the embodiment of the present application. Figure 2B , the relevant model training process is explained:

[0124] Step 211: The processing device obtains a sample data set, wherein a piece of sample data includes: a piece of sample audio data based on any language and recorded using a sample timbre; each sample timbre covered by the sample data set includes a target timbre.

[0125] In the embodiment of the present application, according to actual processing needs, the processing device can construct the sample data set by itself, or obtain the sample data set from other devices.

[0126] In the process of constructing a sample data set, the processing device can determine each timbre that has a timbre change requirement as a sample timbre, and then for each sample timbre, obtain sample audio data recorded using the sample timbre; then, take a sample audio data of a sample timbre as a sample audio data, and obtain a sample data set including multiple sample audio data.

[0127] It should be noted that in the embodiment of the present application, since the audio data recorded using different sample timbres may have different durations, when generating sample audio data, the processing device can segment the acquired audio data according to the preset duration of the sample audio data, so that multiple sample audio data can be obtained based on one audio data.

[0128] For example, see Figure 2C As shown, it is a schematic diagram of the process of splitting multiple sample audio data from one audio data in an embodiment of the present application. It is assumed that each sample audio data includes T frames of audio frames, and the duration corresponding to each audio frame is n; then, the duration corresponding to each sample audio data is T*n, then based on Figure 2C The audio data shown in the figure can be split into 4 sample audio data, namely sample audio data 1-4, and the duration of each sample audio data is T*n.

[0129] In addition, in the embodiment of the present application, for a target timbre that has a timbre change requirement, during the model training process, the target timbre needs to be used as a selected sample timbre to construct corresponding sample audio data; in particular, assuming that there is a timbre change requirement for multiple timbres, multiple timbres can be used as selected partial sample timbres. In addition, the processing device can select some other timbres as sample timbres according to actual processing needs, and then determine multiple sample timbres participating in the model training.

[0130] Furthermore, when the processing device constructs sample audio data based on the selected multiple sample timbres, it can split the sample audio data for each sample timbre from the audio data based on any language recorded using the sample timbre; in particular, it can obtain audio data recorded based on the native language of the corresponding sample object using the sample timbre, and obtain the sample audio data based on the splitting of the acquired audio data, wherein the native language of an object generally refers to the language that the object is most familiar with and can read most fluently.

[0131] Afterwards, the processing device uses the obtained multiple sample audio data as sample data respectively to obtain a sample data set including multiple sample data, wherein one sample data includes one sample audio data.

[0132] Step 2012: The processing device uses the sample data set to perform multiple rounds of iterative training on the initial timbre representation network and the initial audio generation network in the initial timbre conversion model to obtain the trained target timbre conversion model and the timbre features extracted for each sample timbre within the target timbre conversion model.

[0133] In an embodiment of the present application, according to actual processing needs, the initial timbre conversion model constructed by the processing device may only include an initial timbre characterization network and an initial audio generation network, or the initial timbre conversion model constructed by the processing device may include: an initial timbre characterization network and an initial audio generation network that need to be trained, and a content feature extraction network and a pitch feature extraction network that do not need to be trained.

[0134] In other words, in the embodiment of the present application, when the intelligent sound changing system for realizing the timbre transformation of audio is functionally divided as a whole, it can be divided into a content recognition module, a target timbre characterization module (corresponding to the function of the initial timbre characterization network), a waveform generation module (corresponding to the function of the initial audio generation network), and a pitch characterization module. In some feasible implementations, an initial timbre conversion model can be constructed based on the implementation network corresponding to the target timbre characterization module and the implementation network corresponding to the waveform generation module; in other feasible implementations, an initial timbre conversion model can be constructed based on the implementation network of the content recognition module, the implementation network of the target timbre characterization module, the implementation network of the pitch characterization module, and the implementation network of the waveform generation module.

[0135] See also Figure 3A As shown, it is a schematic diagram of the structure of an initial timbre conversion model constructed in the embodiment of the present application. Figure 3AAs can be seen from the illustrated content, the constructed initial timbre conversion model includes an initial timbre representation network and an initial audio generation network that need to be trained; in this case, it is necessary to use a pitch representation module for extracting fundamental frequency features and phonetic features, and a content recognition module for extracting content features, to assist in generating the input content of the initial timbre conversion model.

[0136] Below Figure 3A The functional modules and networks involved in the intelligent voice changing system shown in the figure are described as follows:

[0137] Content recognition module: The corresponding implementation network includes a subnetwork for implementing automatic speech recognition (ASR) and an N-layer transformer for extracting hidden content features from text content; after the sample audio data is processed by the content recognition module, the content features obtained are recorded as linguistic embeddings. According to actual processing needs, the size can be T*1024, where T represents the total number of frames of the input sample audio data.

[0138] Initial timbre characterization network: used to implement the function of the target timbre characterization module, where the core of the initial timbre characterization network is a deep neural network based on convolution, and the core network for implementing feature extraction specifically uses the processed spectrogram as input. Specifically, in the processing of the initial timbre characterization network, the spectrogram of the sample audio data is first extracted to obtain the spectrogram, and then the spectrogram is randomly shuffled in the time dimension to remove the content information and only retain the timbre information. After that, the timbre information is processed through multiple layers of convolution and maximum pooling layers to obtain a timbre characterization vector. After the training is completed, the timbre characterization vector extracted for the target timbre is the timbre feature of the target timbre.

[0139] Pitch characterization module: The corresponding implementation network includes a subnetwork for extracting fundamental frequency information, a subnetwork for implementing interpolation discretization processing, a subnetwork capable of extracting phonological identification information based on fundamental frequency information, each feature embedding subnetwork, and each convolution subnetwork, wherein the subnetwork for extracting fundamental frequency information may specifically be a subnetwork for implementing the audio analysis function of the Librosa tool. In the processing of the pitch characterization module, the fundamental frequency information (F0) of the sample audio data is first extracted by the librosa tool, and F0 is linearly interpolated and discretized, the information content of F0 is divided into a specified number (such as 264) of gears, and the information content is discretized according to the corresponding gears; at the same time, the phonological identification information is obtained based on F0, recorded as (voiced / unvoiced, v / uv); then, the discretized F0 and v / uv are processed by the corresponding feature embedding (embedding) subnetwork to obtain F0 embedding and v / uv embedding, and then the two sets of vectors are respectively input into the corresponding convolution network for encoding to obtain fundamental frequency features and phonological features.

[0140] For example, after the F0 of T frames of audio is processed by interpolation discretization and feature embedding, a T*512-dimensional F0 embedding feature is obtained. At the same time, after feature embedding processing of the phonetic identification information, a T*128-dimensional v / uv embedding feature is obtained; then, the F0 embedding feature is encoded by a convolutional subnetwork to obtain a T*1024-dimensional fundamental frequency feature (denoted as pitch embedding), and the v / uv embedding feature is encoded by a convolutional subnetwork to obtain a T*1024-dimensional phonetic feature (denoted as v / uvembedding), where the fundamental frequency feature and phonetic feature can be collectively referred to as pitch parameters.

[0141] Initial audio generation network: used to implement the function of the waveform generation module. The initial audio generation network is composed of a multi-layer deconvolution neural network (referred to as a multi-layer deconvolution subnetwork) and incorporates the function of the neural source filter (NSF) component. The initial audio generation network takes the superposition result of fundamental frequency features and content features, the superposition result of phonological features and content features, and timbre features as input, and can provide stable and reliable excitation for the waveform generation of the target timbre.

[0142] See also Figure 3B As shown, it is a structural diagram of another initial timbre conversion model constructed in the embodiment of the present application, combined with the attached Figure 3B As can be seen from the illustrated content, the constructed initial timbre conversion model includes an initial timbre representation network and an initial audio generation network that need to be trained, as well as a pitch representation network and a content recognition network that do not need to be trained.

[0143] Below Figure 3B The functional modules and networks involved in the intelligent voice changing system shown in the figure are described as follows:

[0144] Content recognition network: used to implement the functions of the content recognition module. The implementation network corresponding to the content recognition network includes a subnetwork for implementing automatic speech recognition (ASR) and an N-layer transformer subnetwork for extracting content features of the hidden layer based on text content. After the sample audio data is processed by the content recognition network, the content features obtained are recorded as linguistic embeddings. According to actual processing needs, the size can be T*1024, where T represents the total number of frames of the input audio data.

[0145] Initial timbre characterization network: used to implement the function of the target timbre characterization module, wherein the core of the initial timbre characterization network is a deep neural network based on convolution, and the core network for implementing feature extraction specifically uses the processed spectrogram as input. Specifically, in the processing process of the initial timbre characterization network, the sample audio data is firstly extracted from the spectrogram to obtain the spectrogram, and then the spectrogram is randomly shuffled in the time dimension to remove the content information and retain only the timbre information; then, the timbre information is processed through multiple layers of convolution and maximum pooling layers to obtain a timbre characterization vector, wherein after the training is completed, the timbre characterization vector extracted for the target timbre is the timbre feature of the target timbre. Pitch characterization network: used to implement the function of the pitch characterization module, wherein the implementation network corresponding to the pitch characterization network includes a subnetwork for implementing fundamental frequency information extraction, a subnetwork for implementing interpolation discretization processing, a subnetwork capable of extracting phonetic identification information based on fundamental frequency information, each feature embedding subnetwork, and each convolution subnetwork, wherein the subnetwork for implementing fundamental frequency information extraction can specifically be a subnetwork for implementing the audio analysis function of the Librosa tool. In the processing of the pitch representation module, the fundamental frequency information (F0) of the sample audio data is first extracted through the librosa tool, and F0 is linearly interpolated and discretized, the information content of F0 is divided into a specified number (such as 264) gears, and the information content is discretized according to the corresponding gears; at the same time, the phonetic identification information is obtained based on F0, recorded as (voiced / unvoiced, v / uv); then, the discretized F0 and v / uv are processed through the corresponding feature embedding (embedding) sub-network to obtain F0 embedding and v / uv embedding, and then the two sets of vectors are respectively input into the corresponding convolutional network for encoding to obtain the fundamental frequency feature and phonetic feature.

[0146] For example, after the F0 of T frames of audio is processed by interpolation discretization and feature embedding, a T*512-dimensional F0 embedding feature is obtained. At the same time, after feature embedding processing of the phonetic identification information, a T*128-dimensional v / uv embedding feature is obtained; then, the F0 embedding feature is encoded by a convolutional subnetwork to obtain a T*1024-dimensional fundamental frequency feature (denoted as pitch embedding), and the v / uv embedding feature is encoded by a convolutional subnetwork to obtain a T*1024-dimensional phonetic feature (denoted as v / uvembedding), where the fundamental frequency feature and phonetic feature can be collectively referred to as pitch parameters.

[0147] Initial audio generation network: used to implement the function of the waveform generation module. The initial audio generation network is composed of a multi-layer deconvolution neural network (remembered as a multi-layer deconvolution subnetwork) and incorporates the function of the NSF component. The initial audio generation network takes the superposition result of fundamental frequency characteristics and content characteristics, the superposition result of phonological characteristics and content characteristics, and timbre characteristics as input, and can provide stable and reliable excitation for the waveform generation of the target timbre.

[0148] Furthermore, after constructing the initial timbre conversion model, the processing device uses a sample data set to perform multiple rounds of iterative training on the initial timbre representation network and the initial audio generation network in the initial timbre conversion model until the preset convergence conditions are met, thereby obtaining the trained target timbre conversion model and the timbre features extracted for each sample timbre within the target timbre conversion model.

[0149] For example, in a feasible implementation, the training sample set includes 50 sample timbres, each sample timbre has 1,000 sample audio data, and the number of iterations of model training is 3,000,000 rounds, wherein repeated training samples are allowed in different training rounds.

[0150] It should be noted that, since both Figure 3A The model structure shown in Figure 3B The model structure shown in FIG. 1 does not change the functional modules involved in the model training process, but the functional scope covered by the initial audio conversion model changes. Therefore, in the following description of this application, only the initial audio conversion model is used. Figure 3B Taking the illustrated model structure as an example, the initial round of iterative training process for the initial audio transformation model is schematically explained:

[0151] See also Figure 4A As shown, it is a schematic diagram of the process of a round of iterative training in the embodiment of the present application. Figure 4A The following is a description of the relevant processing procedures:

[0152] It should be noted that, during a round of iterative training in the embodiment of the present application, the number of sample audio data simultaneously input into the initial timbre conversion model, that is, the value of batchsize, is set according to actual processing needs, and the present application does not impose any specific restrictions on this; the following only takes the batchsize value of 1 as an example to illustrate the relevant processing process. It should be understood that when the batchsize value is greater than 1, the processing device adopts the initial timbre conversion model and executes the following processing process in parallel for each sample audio data read.

[0153] Step 401: The processing device extracts sample content features based on the text content obtained by recognizing the read sample audio data, and extracts corresponding sample fundamental frequency features and sample phonetic features based on the fundamental frequency information and phonetic identification information obtained by recognizing the sample audio data, and uses an initial timbre characterization network to extract sample timbre features based on the timbre information obtained by recognizing the sample audio data.

[0154] Specifically, after the processing device reads the sample audio data required for a round of iterative training, it uses the initial timbre change model to perform automatic speech recognition on the read sample audio data, identifies the text content expressed by the sample audio data, and then extracts features from the text content to obtain sample content features.

[0155] At the same time, in the process of extracting the corresponding sample fundamental frequency features and sample phonological features based on the fundamental frequency information and phonological identification information obtained by identifying the sample audio data, the processing device identifies the fundamental frequency information in the sample audio data, and obtains phonological identification information for indicating the unvoiced and voiced sounds in the sample audio data based on the identified fundamental frequency information, and extracts the phonological features of the sample audio data based on the phonological identification information; then determines the comprehensive value range corresponding to each information content in the fundamental frequency information, and based on the principle of continuity of the information content, reassigns the information content associated with the unvoiced sound in the fundamental frequency information, and resets the content of each information content in the reassigned fundamental frequency information according to the comprehensive value range and the preset total number of value categories; thereafter, extracts the corresponding sample fundamental frequency features according to the reset fundamental frequency information of each information content.

[0156] Specifically, after extracting the fundamental frequency information based on the sample audio data, the processing device can identify the position of the unvoiced and voiced sounds according to the vibration frequency of the vocal cords when producing voiced sounds represented by the fundamental frequency information, wherein the position of the unvoiced and voiced sounds is represented by the phonological identification information. Furthermore, in order to ensure the continuity of the information content at different positions in the fundamental frequency information, the processing device reassigns the information content corresponding to the unvoiced sounds in the fundamental frequency information, and discretizes the information content of the reassigned fundamental frequency information to obtain the fundamental frequency information of each information content reset; thereafter, based on the fundamental frequency information of each information content reset, the corresponding fundamental frequency features can be extracted; at the same time, the processing device extracts the corresponding phonological features based on the phonological identification information.

[0157] It should be noted that when re-assigning the information content associated with the unvoiced sound in the baseband information, considering that the information content associated with the unvoiced sound may be in the end area or the middle area of ​​the baseband information, there are the following two processing methods when re-assigning the information content associated with the unvoiced sound:

[0158] Method 1: Reassign the information content in the endpoint area.

[0159] Specifically, in the processing process corresponding to the first mode, the processing device re-assigns the information content in the endpoint area associated with the unvoiced sound in the fundamental frequency information by using the information content that is closest to the endpoint area and corresponds to the voiced sound.

[0160] It should be noted that, in the embodiment of the present application, since the information content associated with the unvoiced sound and the information content associated with the voiced sound in the baseband information are completely different, the unvoiced sound and the voiced sound can be distinguished according to the information content in the baseband information; based on this, the information content area corresponding to the audio frame of the unvoiced sound in the first audio frame position or the last audio frame position in the sample audio data is the endpoint area, wherein the amount of information content covered in the endpoint area is not fixed; the total number of endpoint areas is 1 or 2.

[0161] For example, see Figure 4B As shown, it is a schematic diagram of the process of re-assigning information content associated with unvoiced sounds in an embodiment of the present application, combined with the attached Figure 4B As shown in the figure, assuming that the sample audio data 1 includes 10 audio frames, after extracting the baseband information of the sample audio data 1, we can get Figure 4B The baseband information shown in FIG. 1 includes 10 pieces of information, one piece of information corresponds to one audio frame. Furthermore, since the first audio frame: audio frame 1, and the last audio frame: audio frames 9 and 10, have a corresponding information content of 0, that is, they correspond to unvoiced sounds. Therefore, it can be indicated that Figure 4BThe endpoint area in the audio frame; furthermore, the information content in the endpoint area can be reassigned to the information content of the closest corresponding voiced sound, so the information content corresponding to audio frame 1 is reassigned to 87, and the information content corresponding to audio frames 9 and 10 is reassigned to 115.

[0162] Method 2: Reassign the information content in the middle area.

[0163] Specifically, in the processing process corresponding to the second method, the processing device determines, for the middle area in the fundamental frequency information associated with the unvoiced sound, the content positions covered by the middle area and the information content corresponding to the voiced sounds at both ends, and reassigns the information content at each content position according to the information content corresponding to the voiced sounds at both ends.

[0164] It should be noted that, in the embodiment of the present application, since the information content associated with the unvoiced sound and the information content associated with the voiced sound in the baseband information are completely different, the unvoiced sound and the voiced sound can be distinguished according to the information content in the baseband information; based on this, in the non-endpoint audio frames within the sample audio data, the information content area corresponding to the audio frame of the unvoiced sound in the baseband information is the middle area, wherein an middle area covers continuous information content and the number of information contents covered is not fixed; the total number of determined middle areas is also not fixed.

[0165] In addition, when re-assigning the information content in the middle area, the information content at each content position in the middle area can be re-assigned according to the information content corresponding to the voiced sounds at both ends of the middle area, so that the value of the information content changes smoothly in the middle area.

[0166] Continue to combine Figure 4B The content shown is explained. In the non-endpoint area corresponding to the sample audio data, it is possible to determine that the area in the baseband information corresponding to audio frames 5 and 6 is the middle area, that is, the middle area includes two content positions, corresponding to audio frames 5 and 6 in the sample audio data; at the same time, it is possible to determine that the information content corresponding to the voiced sound at both ends of the middle area is 120 and 90. Combined with the content positions included in the middle area, it can be seen that when the information content is smoothly reduced from 120 to 90, it can be reduced three times, with an average reduction of 10 each time. Therefore, the information content corresponding to audio frame 5 can be reassigned to 120-10=110, and the information content corresponding to audio frame 6 can be reassigned to 110-10=100.

[0167] In this way, through the processing of method one and method two, the information content associated with the unvoiced sound in the baseband information can be reassigned, so that the various information contents in the reassigned baseband information can transition smoothly, and the continuity of the baseband information can be guaranteed, which helps to improve the subsequent feature extraction effect for the baseband information.

[0168] Furthermore, the processing device, based on the comprehensive value range corresponding to each information content before reassignment and the total number of preset value categories, divides the comprehensive value range into a corresponding number of sub-value ranges according to the total number of preset value categories when resetting the content for each information content in the baseband information after reassignment, wherein one sub-value range corresponds to one classification label; and performs the following operations for each information content: determining the sub-value range corresponding to one information content, and resetting one information content to the classification label corresponding to the sub-value range.

[0169] For example, see Figure 4C As shown, it is a schematic diagram of the process of resetting the content of the baseband information in an embodiment of the present application, combined with the attached Figure 4C It can be seen from the public content that after re-assigning the information content corresponding to the unvoiced sound in the baseband information (or the information content associated with the unvoiced sound), the processed baseband information is discretized. Figure 4C As shown, the comprehensive value range corresponding to the information content in the audio information before re-assignment is 0 to 132. When the total number of preset classification gears is 33, the value span of the sub-value range corresponding to each classification gear is (132-0) / 33=4; therefore, the corresponding classification gears can be determined for the values ​​of different information contents in the current baseband information, and the corresponding information content can be reset to the classification label corresponding to the corresponding classification gear; when 33 classification gears are identified by 1-33, it can be discretized as Figure 4C Results shown.

[0170] In this way, by discretizing the information content in the baseband information, it is equivalent to converting the numerical value into a classification value, which can reduce the complexity of the data and improve the data processing efficiency.

[0171] Then, the processing device extracts the corresponding sample fundamental frequency features according to the fundamental frequency information after resetting each information content.

[0172] In this way, in the process of extracting the fundamental frequency features of the samples based on the fundamental frequency information, the processing device first reassigns the information content of the associated unvoiced sound for each information content in the fundamental frequency information, so that the information content in the processed fundamental frequency information changes smoothly, and then discretizes the processed fundamental frequency information, so that the value range of each information content in the fundamental frequency information can be adjusted, thereby improving the data processing efficiency, thereby providing a guarantee for the extraction effect of the fundamental frequency features and ensuring the efficient execution of the fundamental frequency feature extraction process.

[0173] In addition, in an embodiment of the present application, the processing device uses an initial timbre characterization network to extract sample timbre features based on the timbre information obtained by identifying the sample audio data. The initial timbre characterization network is used to obtain the corresponding sample spectrogram based on the sample audio data, and the spectrum content in the sample spectrogram is reorganized with reference to the time dimension, and the sample timbre features are extracted from the processed sample spectrogram.

[0174] Specifically, the processing device adopts an initial timbre characterization network, and first extracts a sample spectrogram from the sample audio data according to a spectrogram extraction method, wherein the spectrogram extraction method includes but is not limited to Fourier transform and fast Fourier transform; then, for the spectral content in the spectrogram, the content is reorganized with reference to the time dimension, and the sample timbre features are extracted based on the sample spectrogram after the content reorganization.

[0175] It should be noted that in the embodiment of the present application, when reorganizing the content with reference to the time dimension, the processing device can shuffle and reorganize it from the time dimension. For example, in the original sample spectrum diagram, the spectrum content is presented according to the normal time dimension; when reorganizing the content with reference to the time dimension, the time is shuffled, and then the spectrum content is reorganized according to the shuffled time order, and finally the processed sample spectrum diagram is obtained.

[0176] In this way, by reorganizing the content of the sample spectrogram with reference to the time dimension, it is possible to avoid the influence of the content in the sample audio data when extracting features based on the sample spectrogram, so that the extraction of sample timbre features and sample content features are completely decoupled and do not interfere with each other. This helps to better train the model to learn the ability to reversely generate audio based on various features during the model training process.

[0177] Step 402: The processing device uses an initial audio generation network to reversely synthesize and predict audio data based on sample timbre features, sample content features, sample fundamental frequency features, and sample phonological features.

[0178] After the processing device extracts the sample timbre features, sample content features, sample fundamental frequency features, and sample phonological features from the sample audio data, it uses an initial audio generation network to output the waveform of the predicted target timbre based on the influence of the input of the sample timbre features, under the reliable excitation of the superposition results of the sample content features and the sample fundamental frequency features, and the superposition results of the sample content features and the sample phonological features, to obtain the predicted audio data output by the model.

[0179] Step 403: The processing device adjusts network parameters of the initial audio generation network and the initial timbre characterization network based on the frequency spectrum difference between the predicted audio data and the sample audio data.

[0180] After the processing device obtains the predicted audio data output by the model, it calculates the loss value according to the spectral difference between the predicted audio data and the sample audio data, and adjusts the network parameters of the initial audio generation network and the initial timbre characterization network based on the loss value obtained in the current round (i.e., a batch).

[0181] In a feasible implementation, when the processing device calculates the loss value based on the spectral difference between the predicted audio data and the sample audio data, it can use a Mel filter group to convert the predicted audio data and the sample audio data into Mel spectra respectively, and then obtain the corresponding loss value by calculating the L1 distance between the Mel spectrum corresponding to the predicted audio data and the Mel spectrum corresponding to the sample audio data.

[0182] In this way, by executing the processing process of steps 401-403, an end-to-end model training can be performed for the initial timbre conversion model, the input to the initial timbre conversion model is the sample audio data, the output from the initial timbre conversion model is the predicted audio data, and the model parameters are adjusted according to the spectral difference between the predicted audio data and the sample audio data.

[0183] Similarly, the processing of steps 401-403 can be repeated to implement multiple rounds of iterative training until a preset number of training rounds is reached to obtain a trained target timbre conversion model and timbre features extracted for each sample object in the target timbre conversion model.

[0184] In this way, with the help of multiple rounds of iterative training for the initial timbre transformation model, the model can be trained to learn the ability to extract timbre features, as well as the ability to reversely generate audio data based on the features that characterize the content, rhythm, and timbre of the audio data.

[0185] In a feasible implementation, the processing device may directly use the timbre features extracted for the target timbre within the trained target timbre transformation model as the acquired timbre features of the target timbre.

[0186] In particular, in an optional implementation, when the target timbre conversion model is trained, the existing recording of the target timbre can be used to fine-tune the target timbre conversion model. Moreover, in order to make the subsequent customized cross-language timbre pronunciation more robust, other corpora of the target language can be introduced.

[0187] Specifically, after the processing device obtains the trained target voice conversion model, it acquires a fine-tuning sample set. One fine-tuning sample includes: recording one fine-tuning audio data based on the target language and using the sample voice, or recording one fine-tuning audio data based on other languages and using the target voice. Then, using the fine-tuning sample set, perform multiple rounds of fine-tuning training on the target voice representation network and the target audio generation network in the target voice conversion model to obtain the final voice conversion model after fine-tuning, and the voice features extracted for the target voice in the final voice conversion model.

[0188] For example, during fine-tuning training, the voices of a male and a female, as well as the target voice, can be selected as each sample voice, and fine-tuning samples are generated based on the audio data of the target voice, and fine-tuning samples are generated based on the audio data recorded by a male and a female in the target language. One fine-tuning sample includes one fine-tuning audio data of one sample voice. Combining the invention background of the present application, generally, there is no audio data of the target voice in the target language.

[0189] It should be noted that the fine-tuning training process has the same processing logic as the training process illustrated in the above steps 401-403. The difference between the two trainings lies in the training samples used. Therefore, the present application will not elaborate on the processing procedures executed during the fine-tuning training process.

[0190] After performing multiple rounds of iterative training on the target voice transformation model, the final voice transformation model after fine-tuning training can be obtained. At the same time, the voice features finally extracted by the final voice transformation model based on the target voice can be obtained. Furthermore, the voice features extracted by the final voice transformation model based on the target voice can be used as the voice features of the obtained target voice.

[0191] It should be noted that in the embodiments of the present application, during the model training process, the voice features extracted for the sample voices can be saved, and the saved voice features are updated as the model parameters are adjusted until the training result is obtained, and the final voice features saved for each sample voice can be obtained.

[0192] In this way, after training the initial voice transformation model using the sample audio data of each sample voice in any language to obtain the target voice transformation model, by performing fine-tuning training on the target voice transformation model, the process of training the target voice transformation model is equivalent to the pre-training process of the model. With the help of fine-tuning training, the influence of the target voice and the influence of the target language can be exerted on the target voice transformation model, which is equivalent to providing learning guidance for the target voice transformation model under specific voice transformation requirements, and helps to improve the processing effect of the model.

[0193] Step 202: The processing device acquires reference audio data recorded based on the target language and using the reference voice.

[0194] In a feasible implementation manner of this application, when performing step 202, the processing device can arbitrarily obtain audio data recorded in the target timbre, and for the audio data, obtain the target average pitch corresponding to the target timbre; then, in the preset audio database, screen out each candidate audio data based on the target language, and for each candidate audio data, respectively obtain the corresponding candidate average pitch; afterwards, among the candidate average pitches, screen out the reference average pitch whose similarity with the target average pitch meets the set conditions, and determine the candidate audio data corresponding to the reference average pitch as the obtained reference audio data recorded in the reference timbre.

[0195] Specifically, in order to improve the transformation effect of the audio timbre, the processing device can perform data screening from the preset audio database, and screen out the reference timbre whose similarity between the average pitch of the timbre and the average pitch of the target timbre meets the set conditions among the audio data that can fluently read the target language, and determine the audio data of this reference timbre as the reference audio data, where the set condition can be that the difference in average pitch does not exceed the set value; the value of the set value is set according to the actual processing needs, and this application does not make specific restrictions on this.

[0196] Preferably, since this application needs to generate audio data of the target timbre based on the target language, then, the reference audio data of the reference timbre that has a mother tongue of another language but can fluently read the target language can be screened out from the audio database, where the average pitch of the reference timbre is preferably as close as possible to the average pitch of the target timbre, the pitch cannot deviate too much, and the object genders corresponding to the target timbre and the reference timbre are the same.

[0197] In this way, by screening the reference audio data in the audio database, a reference timbre whose pitch is adapted to the target timbre can be screened out, which can reduce the processing difficulty of the model in the timbre transformation process, ensure the generation effect of the audio data after the timbre transformation, and better generate the audio data that can be used for speech synthesis.

[0198] In other feasible implementation manners of this application, the processing device can directly select the reference audio data from the audio data recorded based on the target language.

[0199] Step 203: The processing device extracts the content features based on the text content recognized for the reference audio data, and extracts the corresponding fundamental frequency features and phonetic features based on the fundamental frequency information and phonetic symbol information recognized for the reference audio data.

[0200] In the embodiment of the present application, when only the target timbre transformation model is trained, when executing step 203, the target timbre transformation model can be used to obtain content features, fundamental frequency features and phonological features extracted based on the reference audio data; when fine-tuning training is performed based on the target timbre transformation model to obtain the final-state timbre transformation model, when executing step 203, the final-state timbre transformation model can be used to obtain content features, fundamental frequency features and phonological features extracted based on the reference audio data.

[0201] Therefore, the following only takes the processing logic implemented by the model when executing step 203 as an example to illustrate the relevant feature extraction process:

[0202] When extracting fundamental frequency features and phonological features, the processing device identifies fundamental frequency information in the reference audio data, and based on the fundamental frequency information, obtains phonological identification information for indicating unvoiced and voiced sounds in the reference audio data, and extracts phonological features of the parameter audio data based on the phonological identification information; then determines the comprehensive value range corresponding to each information content in the fundamental frequency information, and based on the principle of continuity of information content, reassigns information content associated with unvoiced sounds in the fundamental frequency information, and resets the content of each information content in the reassigned fundamental frequency information according to the comprehensive value range and the preset total number of value categories; and then, extracts the corresponding fundamental frequency features according to the reset fundamental frequency information of each information content.

[0203] It should be noted that, since the relevant processing logic has been described in detail in step 401, it will not be described in detail here.

[0204] In this way, by interpolating and discretizing the information content of the fundamental frequency information, the difficulty of data processing can be simplified, the data processing efficiency can be improved, and the efficiency of generating fundamental frequency features can be improved; moreover, by extracting phonetic features based on certain phonetic identification information, feature extraction can be performed from the perspective of the pronunciation position of unvoiced and voiced sounds; the extracted fundamental frequency features and phonetic features can describe the pitch of the audio data from different perspectives, providing a high-quality reference basis for the subsequent reverse generation of audio data.

[0205] Similar to the processing process described in the above training process, when reassigning the information content associated with the unvoiced sound in the fundamental frequency information, the processing device can, for the endpoint area associated with the unvoiced sound in the fundamental frequency information, use the information content that is closest to the endpoint area and corresponds to the voiced sound to reassign the information content in the endpoint area; then, for the middle area associated with the unvoiced sound in the fundamental frequency information, determine the content positions covered by the middle area and the information content corresponding to the voiced sound at both ends, and reassign the information content at each content position according to the information content corresponding to the voiced sound at both ends.

[0206] In this way, the information content in the baseband information can be processed smoothly, so that the information content in the baseband information after processing is more coherent than before processing, which helps to improve the feature extraction effect.

[0207] Similarly, when resetting the content according to the comprehensive value range corresponding to each information content and the preset total number of value categories, the processing device divides the comprehensive value range into a corresponding number of sub-value ranges according to the preset total number of value categories, wherein one sub-value range corresponds to one classification label; and then performs the following operations for each information content: determines the sub-value range corresponding to an information content, and resets an information content to the classification label corresponding to the sub-value range.

[0208] In this way, by discretizing the information content in the baseband information, it is equivalent to converting the numerical value into a classification value, which can reduce the complexity of the data and improve the data processing efficiency.

[0209] Step 204: The processing device reversely generates target audio data of the target timbre based on the timbre characteristics, content characteristics, fundamental frequency characteristics, and phonological characteristics of the target timbre.

[0210] Similar to the implementation method corresponding to step 203, in the embodiment of the present application, when only the target timbre transformation model is trained, when executing step 204, the target timbre transformation model can be used to achieve reverse generation of target audio data; when fine-tuning training is performed based on the target timbre transformation model to obtain the final timbre transformation model, when executing step 204, the final timbre transformation model can be used to achieve reverse generation of target audio data.

[0211] Therefore, judging only from the processing logic realized with the help of the model, the processing device uses the audio generation network trained in the model to reversely generate the target audio data under the influence of the superposition results of phonetic features and content features, the superposition results of fundamental frequency features and content features, and the timbre characteristics of the target timbre.

[0212] In this way, with the help of the trained timbre conversion model, target audio data that uses the target timbre to read the audio text based on the target language can be generated, which is equivalent to converting the audio timbre of the reference audio data.

[0213] Furthermore, after generating target audio data based on the target language with the target timbre, the processing device can train a speech synthesis model based on the generated target audio data and audio data recorded in a language that can be mastered proficiently with the target timbre, so that multi-language mixed reading based on the target timbre can be achieved.

[0214] Specifically, the processing device obtains each target audio data in the target language generated for the target timbre, and obtains each historical audio data based on other languages ​​and corresponding to the target timbre; then, based on each target audio data and each historical audio data, each speech synthesis sample is generated respectively, wherein a speech synthesis sample includes: the timbre feature of the target timbre, and a target audio data and a corresponding sample text, or the timbre feature of the target timbre, and a historical audio data and a corresponding sample text; thereafter, each speech synthesis sample is used to perform multiple rounds of iterative training on the constructed initial speech synthesis model to obtain a trained target speech synthesis model.

[0215] In addition, it should be noted that each historical audio data includes any one or combination of the following data contents: historical audio data recorded based on other languages ​​using the target timbre; historical audio data in other languages ​​generated for the target timbre.

[0216] In a feasible implementation of the present application, the determined target timbre and target language serve the speech synthesis model. In other words, after clarifying the multi-language mixed reading requirements of the speech synthesis model and the target timbre for multi-language mixed reading, the target language is determined based on the languages ​​that need to be mixed and the languages ​​covered in the audio data of the existing target timbre; based on this, the languages ​​covered in the speech synthesis sample are the languages ​​that need to be mixed.

[0217] It should be understood that the technical solution proposed in the embodiments of the present application can be applied to any model that has a need for speech synthesis, and can provide suitable training samples for the training of the initial speech synthesis model. The audio timbre transformation process implemented in the present application is equivalent to data augmentation for the training of the initial speech synthesis model, thereby improving the naturalness and similarity of zero-sample multi-language mixed reading.

[0218] For example, the initial speech synthesis model in the present application can be a speech synthesis system based on the VITS architecture. After obtaining the audio of the target timbre A in language X and the audio of the modified language Y, the two audios can be mixed together and share the same speaker vector; and then by fine-tuning the VITS speech synthesis system, the timbre A that can speak both language X and language Y can be obtained.

[0219] For example, see Figure 5As shown, it is a schematic diagram of the process of training a target speech synthesis model in an embodiment of the present application. Assuming that the speech synthesis model ultimately achieves mixed reading of Chinese and English, the selected target timbre is target timbre 1 and the native speaker of target timbre 1 is Chinese, and there is historical audio data of target timbre 1 recorded based on Chinese, then the constructed speech synthesis samples include two categories, one is speech synthesis samples generated based on the historical audio data of target timbre 1; the other is speech synthesis samples generated based on the target audio data after English is generated for the target timbre using English as the target language.

[0220] Furthermore, continue to combine Figure 5 To illustrate, the processing device uses the sample text and the timbre features of the target timbre as model input to obtain the predicted audio output by the initial speech synthesis model; then based on the spectral difference between the predicted audio and the corresponding sample audio, the model loss is calculated, and the model parameters of the speech synthesis model are adjusted based on the model loss.

[0221] In this way, with the help of the generated target audio data, training samples can be provided for the training of the speech synthesis model, effectively solving the problem of missing corpus for zero-sample multi-language mixed reading. Moreover, with the help of the intelligent voice change scheme proposed in this application, cross-language augmented corpus can be provided for the text-to-speech (TTS) task. Based on this augmented corpus, a TTS voice with excellent multi-language mixed reading effect can be trained.

[0222] Furthermore, after the trained target speech synthesis model is obtained, in a specific application process, the processing device can respond to a reading request for the target text data, adopt the target speech synthesis model, and synthesize the corresponding audio to be played based on the text features extracted from the target text data and the timbre features of the target timbre; play the audio to be played to provide the target text data read by the target timbre.

[0223] Specifically, with the aid of the trained target speech synthesis model, the speech to be played when the target text data is read aloud with the target timbre can be synthesized according to the input target text data and the timbre characteristics of the target timbre.

[0224] In a feasible implementation, the processing device can send the trained intelligent voice change system to the terminal device, so that the audio to be played can be synthesized on the terminal device; or, the processing device can respond to the reading request of the relevant object, adopt the target speech synthesis model, synthesize the video to be played for the target text data, and feed back the video to be played to the terminal device.

[0225] In addition, the processing device can provide the user with a function of selecting a target timbre, so that the processing device can generate audio to be played in a targeted manner according to the target timbre selected by the user and the target text data to be played.

[0226] In this way, in specific business scenarios, audio to be played that meets the needs can be generated according to actual playback needs. Moreover, since the target speech synthesis model used to generate the video to be played is trained with appropriate training samples, the naturalness and timbre consistency of the audio data generated when multiple languages ​​are mixed can be guaranteed, thereby greatly improving the user experience and effectively increasing the user retention rate of the product.

[0227] The following is a schematic illustration of the audio timbre transformation process proposed in this application in conjunction with a specific application scenario:

[0228] Assume that there is a need to perform mixed reading of multi-language contents for a target timbre, and the languages ​​that need to be mixed read include Chinese, English, and Korean, and currently there is only a large amount of audio data recorded based on Chinese using the target timbre.

[0229] Then it is necessary to take English and Korean as the target timbre respectively, and train to obtain the final-state timbre conversion model corresponding to English and the final-state timbre conversion model corresponding to Korean, wherein different final-state timbre conversion models can be obtained by fine-tuning training based on a target timbre conversion model; or, different final-state timbre conversion models can be obtained after fine-tuning training based on different target timbre conversion models.

[0230] Furthermore, based on the final-state timbre conversion model, audio data that speaks Korean based on a reference timbre can be converted into target video data that speaks the same content with a target timbre, and audio data that speaks English based on a reference timbre can be converted into target video data that speaks the same content with a target timbre.

[0231] Furthermore, see Figure 6 As shown, it is a schematic diagram of the process of training the initial speech synthesis model in an embodiment of the present application. After obtaining audio data of Chinese spoken with the target timbre, and obtaining audio data of Korean spoken with the target timbre and audio data of English spoken with the target timbre generated by the model, the audio data of the target timbre in the three languages ​​are used to generate each speech synthesis sample; then, each speech synthesis sample is used to perform multiple rounds of iterative training on the constructed initial speech synthesis model to obtain a trained target speech synthesis model.

[0232] In summary, the present application proposes a solution that can significantly improve the effect of zero-sample multi-language mixed reading in speech synthesis. With the help of an end-to-end intelligent voice change system and existing real recording corpus, the timbre of the corpus is replaced with the target timbre. Even if the target timbre does not originally have any corpus in the target language, it can easily give the timbre the ability to read the target language. Moreover, it should be understood that the present application proposes an algorithm system framework and timbre customization scheme for end-to-end intelligent voice change. The algorithm framework has the ability to convert the timbre from any person to any person, and can restore the tone and rhythm of the reference audio data with high fidelity to ensure the fluency of the expression. The present application can accurately convert the target timbre to the target language while ensuring the high similarity of the target timbre; in addition, the present application proposes a method of applying the converted audio data to a common speech synthesis system (such as the VITS system), so that through the corpus fusion scheme, the effect of zero-sample multi-language mixed reading in speech synthesis can be significantly improved, and the unnaturalness and low timbre similarity problems in zero-sample multi-language mixed reading can be solved.

[0233] Based on the same inventive concept, refer to Figure 7 As shown, it is a schematic diagram of the logical structure of the audio timbre conversion device in the embodiment of the present application. The audio timbre conversion device 700 includes a determination unit 701, an acquisition unit 702, an extraction unit 703, and a generation unit 704, wherein:

[0234] A determination unit 701 is used to determine a target language and a target timbre, and obtain a timbre feature of the target timbre;

[0235] An acquisition unit 702 is used to acquire reference audio data based on the target language and recorded using a reference timbre;

[0236] An extraction unit 703 is used to extract content features based on the text content identified with respect to the reference audio data, and to extract corresponding fundamental frequency features and phonological features based on the fundamental frequency information and phonological identification information identified with respect to the reference audio data;

[0237] The generating unit 704 is used to reversely generate target audio data of the target timbre based on the timbre characteristics, content characteristics, fundamental frequency characteristics, and phonological characteristics of the target timbre.

[0238] Optionally, when extracting corresponding fundamental frequency features and phonological features based on fundamental frequency information and phonological identification information obtained by identifying reference audio data, the extraction unit 703 is used to:

[0239] Identifying fundamental frequency information in the reference audio data, and based on the fundamental frequency information, obtaining phonological identification information for indicating unvoiced and voiced sounds in the reference audio data, and extracting phonological features of the parameter audio data based on the phonological identification information;

[0240] Determine the comprehensive value range corresponding to each information content in the baseband information, and re-assign the information content associated with the unvoiced sound in the baseband information based on the principle of information content continuity, and reset the content of each information content in the re-assigned baseband information according to the comprehensive value range and the total number of preset value categories;

[0241] According to resetting the fundamental frequency information of each information content, the corresponding fundamental frequency features are extracted.

[0242] Optionally, when re-assigning information content associated with the unvoiced sound in the fundamental frequency information, the extraction unit 703 is configured to:

[0243] For the endpoint area associated with the unvoiced sound in the fundamental frequency information, the information content in the endpoint area is reassigned by using the information content that is closest to the endpoint area and corresponds to the voiced sound;

[0244] For the middle area associated with the unvoiced sound in the fundamental frequency information, the content positions covered by the middle area and the information content of the corresponding voiced sounds at both ends are determined, and the information content at each content position is reassigned according to the information content of the corresponding voiced sounds at both ends.

[0245] Optionally, when resetting the content according to the comprehensive value range corresponding to each information content and the total number of preset value categories, the extraction unit 703 is used to:

[0246] According to the total number of preset value categories, the comprehensive value range is divided into a corresponding number of sub-value ranges, wherein one sub-value range corresponds to one category label;

[0247] For each information content, the following operations are performed respectively: a sub-value range corresponding to an information content is determined, and an information content is reset to a classification label corresponding to the sub-value range.

[0248] Optionally, the timbre feature of the target timbre is extracted by the training unit 705 in the device in the following manner:

[0249] Obtain a sample data set; a sample data includes: a sample audio data based on any language and recorded using a sample timbre; each sample timbre covered by the sample data set includes a target timbre;

[0250] Using a sample data set, the initial timbre representation network and the initial audio generation network in the initial timbre conversion model are trained for multiple rounds of iterative training to obtain the trained target timbre conversion model and the timbre features extracted for each sample timbre within the target timbre conversion model.

[0251] Optionally, during a round of iterative training, the training unit 705 is configured to perform the following operations:

[0252] Extracting sample content features based on the text content identified from the sample audio data read, extracting corresponding sample fundamental frequency features and sample phonological features based on the fundamental frequency information and phonological identification information identified from the sample audio data, and extracting sample timbre features based on the timbre information identified from the sample audio data using an initial timbre representation network;

[0253] Using an initial audio generation network, based on sample timbre features, sample content features, sample fundamental frequency features, and sample phonological features, reverse synthesis predicts audio data;

[0254] Based on the spectral difference between the predicted audio data and the sample audio data, the network parameters of the initial audio generation network and the initial timbre characterization network are adjusted.

[0255] Optionally, when the initial timbre characterization network is used to extract sample timbre features based on timbre information obtained by identifying sample audio data, the training unit 705 is used to:

[0256] An initial timbre characterization network is used to obtain the corresponding sample spectrogram based on the sample audio data. The spectral content in the sample spectrogram is reorganized with reference to the time dimension, and the sample timbre features are extracted from the processed sample spectrogram.

[0257] Optionally, after obtaining the trained target timbre conversion model, the training unit 705 is further used to:

[0258] Acquire a fine-tuning sample set, wherein a fine-tuning sample includes: a fine-tuning audio data recorded based on the target language and using the sample timbre, or a fine-tuning audio data recorded based on another language and using the target timbre;

[0259] Using the fine-tuning sample set, multiple rounds of fine-tuning training are performed on the target timbre representation network and the target audio generation network in the target timbre conversion model to obtain the fine-tuned final timbre conversion model and the timbre features extracted for the target timbre in the final timbre conversion model.

[0260] Optionally, when obtaining reference audio data based on the target language and recorded with a reference timbre, the obtaining unit 702 is used to:

[0261] Arbitrarily obtain audio data recorded with a target timbre, and obtain a target average pitch corresponding to the target timbre for the audio data;

[0262] In a preset audio database, candidate audio data based on the target language are screened out, and for each candidate audio data, corresponding candidate average pitch is obtained respectively;

[0263] Among the candidate average pitches, a reference average pitch whose similarity to the target average pitch meets the set conditions is selected, and the candidate audio data corresponding to the reference average pitch is determined as the obtained reference audio data recorded using the reference timbre.

[0264] Optionally, after the target audio data of the target timbre is reversely generated, the training unit 705 in the device is further configured to:

[0265] Obtain each target audio data in the target language generated for the target timbre, and obtain each historical audio data based on other languages and corresponding to the target timbre;

[0266] Based on each target audio data and each historical audio data, respectively generate each speech synthesis sample, where one speech synthesis sample includes: the timbre feature of the target timbre, and one target audio data and the corresponding sample text, or, the timbre feature of the target timbre, and one historical audio data and the corresponding sample text;

[0267] Use each speech synthesis sample to perform multiple rounds of iterative training on the constructed initial speech synthesis model to obtain the trained target speech synthesis model.

[0268] Optionally, after the trained target speech synthesis model is obtained, the training unit 705 is further configured to:

[0269] In response to a reading request for the target text data, use the target speech synthesis model to synthesize the corresponding audio to be played based on the text feature extracted from the target text data and the timbre feature of the target timbre;

[0270] Play the audio to be played to provide the target text data read by the target timbre.

[0271] After introducing the method and device for transforming the audio timbre in the exemplary embodiments of the present application, next, an electronic device according to another exemplary embodiment of the present application is introduced.

[0272] Those skilled in the art can understand that various aspects of the present application can be implemented as a system, method, or program product. Therefore, various aspects of the present application can be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, which can be collectively referred to as "circuit", "module", or "system" here.

[0273] Based on the same inventive concept as the above method embodiment, in the case where the electronic device in the embodiment of the present application corresponds to a processing device, refer to Figure 8As shown, it is a schematic diagram of the hardware composition structure of an electronic device using an embodiment of the present application, and the electronic device 800 may include at least a processor 801 and a memory 802. Among them, the memory 802 stores a computer program, and when the computer program is executed by the processor 801, the processor 801 performs any of the above-mentioned steps of audio timbre transformation.

[0274] In some possible implementations, the electronic device according to the present application may include at least one processor and at least one memory. The memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of audio timbre transformation according to various exemplary implementations of the present application described above in this specification. For example, the processor may execute the following steps: Figure 2A Follow the steps shown in .

[0275] Based on the same inventive concept as the above method embodiment, various aspects of the audio timbre transformation provided by the present application can also be implemented in the form of a program product, which includes a program code. When the program product is run on an electronic device, the program code is used to enable the electronic device to execute the steps of the audio timbre transformation according to various exemplary embodiments of the present application described above in this specification. For example, the electronic device can execute the following steps: Figure 2A Follow the steps shown in .

[0276] The program product may use any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0277] Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications falling within the scope of the present application.

[0278] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is also intended to include these modifications and variations.

Claims

1. A method for changing audio timbre, It is characterized in that include: Determining a target language and a target timbre, and obtaining timbre characteristics of the target timbre; Acquiring reference audio data based on the target language and recorded using a reference timbre; Extracting content features based on the text content identified from the reference audio data, and extracting corresponding fundamental frequency features and phonological features based on the fundamental frequency information and phonological identification information identified from the reference audio data; Based on the timbre feature of the target timbre, the content feature, the fundamental frequency feature, and the phonological feature, target audio data of the target timbre is inversely generated.

2. The method according to claim 1, It is characterized in that The extracting corresponding fundamental frequency features and phonological features based on the fundamental frequency information and phonological identification information obtained from the reference audio data includes: Identifying fundamental frequency information in the reference audio data, and based on the fundamental frequency information, obtaining phonological identification information for indicating unvoiced sounds and voiced sounds in the reference audio data, and extracting phonological features of the parameter audio data based on the phonological identification information; Determine a comprehensive value range corresponding to each information content in the baseband information, and re-assign a value to the information content associated with the unvoiced sound in the baseband information based on the principle of information content continuity, and reset the content of each information content in the re-assigned baseband information according to the comprehensive value range and the total number of preset value categories; According to resetting the baseband information of each information content, the corresponding baseband feature is extracted.

3. The method according to claim 2, It is characterized in that The re-assigning the information content associated with the unvoiced sound in the baseband information includes: For the endpoint region associated with the unvoiced sound in the fundamental frequency information, the information content in the endpoint region is reassigned by using the information content that is closest to the endpoint region and corresponds to the voiced sound; For the middle area associated with the unvoiced sound in the fundamental frequency information, the content positions covered by the middle area and the information contents of the voiced sounds corresponding to the two ends are determined, and the information contents at the content positions are reassigned according to the information contents of the voiced sounds corresponding to the two ends.

4. The method according to claim 2, It is characterized in that The resetting of the content according to the comprehensive value range corresponding to each information content and the total number of preset value categories includes: According to the total number of preset value categories, the comprehensive value range is divided into a corresponding number of sub-value ranges, wherein one sub-value range corresponds to one category label; For each of the information contents, the following operations are performed respectively: a sub-value range corresponding to an information content is determined, and the information content is reset to a classification label corresponding to the sub-value range.

5. The method according to claim 1, It is characterized in that The timbre characteristics of the target timbre are extracted in the following manner: Acquire a sample data set; a sample data includes: a sample audio data based on any language and recorded using a sample timbre; each sample timbre covered by the sample data set includes a target timbre; The sample data set is used to perform multiple rounds of iterative training on the initial timbre characterization network and the initial audio generation network in the initial timbre conversion model to obtain the trained target timbre conversion model and the timbre features extracted for each sample timbre within the target timbre conversion model.

6. The method according to claim 5, It is characterized in that During one round of training iterations, the following operations are performed: Extracting sample content features based on the text content identified from the sample audio data read, extracting corresponding sample fundamental frequency features and sample phonological features based on the fundamental frequency information and phonological identification information identified from the sample audio data, and extracting sample timbre features based on the timbre information identified from the sample audio data using the initial timbre representation network; Using the initial audio generation network, based on the sample timbre feature, the sample content feature, the sample fundamental frequency feature, and the sample phonological feature, reversely synthesize and predict the audio data; Based on the frequency spectrum difference between the predicted audio data and the sample audio data, network parameters of the initial audio generation network and the initial timbre characterization network are adjusted.

7. The method according to claim 6, It is characterized in that The initial timbre characterization network is used to extract sample timbre features based on timbre information obtained by identifying the sample audio data, including: The initial timbre characterization network is used to obtain a corresponding sample spectrogram based on the sample audio data, and the spectrum content in the sample spectrogram is reorganized with reference to the time dimension, and the sample timbre features are extracted from the processed sample spectrogram.

8. The method according to claim 5, It is characterized in that After obtaining the trained target timbre conversion model, the method further includes: Acquire a fine-tuning sample set, wherein a fine-tuning sample includes: a fine-tuning audio data recorded based on the target language and using the sample timbre, or a fine-tuning audio data recorded based on another language and using the target timbre; The fine-tuning sample set is used to perform multiple rounds of fine-tuning training on the target timbre representation network and the target audio generation network in the target timbre conversion model to obtain the fine-tuned final timbre conversion model and the timbre features extracted for the target timbre in the final timbre conversion model.

9. The method according to claim 1, It is characterized in that The obtaining of reference audio data based on the target language and recorded with a reference timbre includes: Any audio data recorded using the target timbre is obtained, and for the audio data, a target average pitch corresponding to the target timbre is obtained; In a preset audio database, select candidate audio data based on the target language, and obtain corresponding candidate average pitches for each candidate audio data; Among the candidate average pitches, a reference average pitch whose similarity with the target average pitch meets the set conditions is screened out, and the candidate audio data corresponding to the reference average pitch is determined as the obtained reference audio data recorded using the reference timbre.

10. The method according to any one of claims 1 to 9, It is characterized in that After the target audio data of the target timbre is generated inversely, the method further includes: Acquire each target audio data in the target language generated for the target timbre, and acquire each historical audio data based on other languages ​​and corresponding to the target timbre; Based on the target audio data and the historical audio data, generate speech synthesis samples respectively, wherein a speech synthesis sample includes: the timbre feature of the target timbre, a target audio data and a corresponding sample text, or the timbre feature of the target timbre, a historical audio data and a corresponding sample text; The speech synthesis samples are used to perform multiple rounds of iterative training on the constructed initial speech synthesis model to obtain a trained target speech synthesis model.

11. The method according to claim 10, It is characterized in that After obtaining the trained target speech synthesis model, the method further includes: In response to a request for reading aloud the target text data, synthesizing corresponding audio to be played using the target speech synthesis model based on text features extracted from the target text data and timbre features of the target timbre; The audio to be played is played to provide the target text data read aloud by the target tone.

12. A device for changing the tone of an audio frequency, It is characterized in that include: A determination unit, used to determine a target language and a target timbre, and obtain a timbre feature of the target timbre; An acquisition unit, configured to acquire reference audio data based on the target language and recorded using a reference timbre; An extraction unit, configured to extract content features based on the text content identified from the reference audio data, and to extract corresponding fundamental frequency features and phonological features based on the fundamental frequency information and phonological identification information identified from the reference audio data; A generating unit is used to reversely generate target audio data of the target timbre based on the timbre feature of the target timbre, the content feature, the fundamental frequency feature, and the phonological feature.

13. The device according to claim 12, It is characterized in that When extracting the corresponding fundamental frequency features and phonological features based on the fundamental frequency information and phonological identification information obtained by identifying the reference audio data, the extraction unit is used to: Identifying fundamental frequency information in the reference audio data, and based on the fundamental frequency information, obtaining phonological identification information for indicating unvoiced sounds and voiced sounds in the reference audio data, and extracting phonological features of the parameter audio data based on the phonological identification information; Determine a comprehensive value range corresponding to each information content in the baseband information, and re-assign a value to the information content associated with the unvoiced sound in the baseband information based on the principle of information content continuity, and reset the content of each information content in the re-assigned baseband information according to the comprehensive value range and the total number of preset value categories; According to resetting the baseband information of each information content, the corresponding baseband feature is extracted.

14. The device according to claim 12, It is characterized in that The timbre feature of the target timbre is extracted by the training unit in the device in the following manner: Acquire a sample data set; a sample data includes: a sample audio data based on any language and recorded using a sample timbre; each sample timbre covered by the sample data set includes a target timbre; The sample data set is used to perform multiple rounds of iterative training on the initial timbre characterization network and the initial audio generation network in the initial timbre conversion model to obtain the trained target timbre conversion model and the timbre features extracted for each sample timbre within the target timbre conversion model.

15. The device according to claim 14, It is characterized in that After obtaining the trained target timbre conversion model, the training unit is further used for: Acquire a fine-tuning sample set, wherein a fine-tuning sample includes: a fine-tuning audio data recorded based on the target language and using the sample timbre, or a fine-tuning audio data recorded based on another language and using the target timbre; The fine-tuning sample set is used to perform multiple rounds of fine-tuning training on the target timbre representation network and the target audio generation network in the target timbre conversion model to obtain the fine-tuned final timbre conversion model and the timbre features extracted for the target timbre in the final timbre conversion model.

16. The device according to any one of claims 12 to 15, It is characterized in that After the target audio data of the target timbre is inversely generated, the training unit in the device is further used for: Acquire each target audio data in the target language generated for the target timbre, and acquire each historical audio data based on other languages ​​and corresponding to the target timbre; Based on the target audio data and the historical audio data, generate speech synthesis samples respectively, wherein a speech synthesis sample includes: the timbre feature of the target timbre, a target audio data and a corresponding sample text, or the timbre feature of the target timbre, a historical audio data and a corresponding sample text; The speech synthesis samples are used to perform multiple rounds of iterative training on the constructed initial speech synthesis model to obtain a trained target speech synthesis model.

17. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, It is characterized in that When the processor executes the computer program, the method according to any one of claims 1 to 11 is implemented.

18. A computer-readable storage medium having a computer program stored thereon, Features: When the computer program is executed by a processor, the method according to any one of claims 1 to 11 is implemented.

19. A computer program product comprising a computer program, It is characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 11 is implemented.