Audio generation method, device, equipment and storage medium thereof

Through pre-trained audio generators and adversarial training networks, the language barrier problem in the international customer service of financial institutions is solved, multi-scale audio generation is achieved, and high-quality, automated and international audio services are provided.

CN119864015BActive Publication Date: 2025-09-30PING AN TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510028146.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-07
Publication Date
2025-09-30
Estimated Expiration
2045-01-07

AI Technical Summary

Technical Problem

As existing financial institutions go international, there are language barriers in generating customer service audio, which cannot meet multi-scale service needs.

Method used

A pre-trained audio generator with multi-scale acoustic feature capabilities is used. The generator is trained through an adversarial training network to identify scale selection instructions for human-computer interaction and generate audio data at multiple time scales, multiple frequency scales, multiple language scales, and multiple emotional scales.

Benefits of technology

It achieves high-quality, automated, and intelligent audio generation in multilingual broadcasting and intelligent voice customer service scenarios, meeting the service needs of international customers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119864015B_ABST
    Figure CN119864015B_ABST
Patent Text Reader

Abstract

The embodiments of the present application belong to the field of R&D design and audio processing technology, and are applied to audio generation scenarios. They relate to an audio generation method, apparatus, device, and storage medium thereof, which obtain target text data for audio generation; input the target text data into a pre-trained audio generator; identify the scale selection instruction obtained through human-computer interaction; and obtain the expected audio data output by the audio generator based on the scale selection instruction. The audio generation method described in this application is applied to multi-scale audio generation scenarios, especially in multi-language broadcasting or intelligent voice customer service follow-up scenarios. It can generate more detailed and high-quality multi-language transliterations based on the differences in audio language, audio time, and audio frequency range, which is more automated and intelligent, and provides more international broadcasting or inquiry services for customers of different languages.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the fields of R&D, design, and audio processing technology, and is applied to audio generation scenarios, and in particular to an audio generation method, apparatus, device, and storage medium thereof. Background Art

[0002] In recent years, large-scale language models have achieved remarkable success in generative tasks such as speech synthesis, music generation, and audio generation. Furthermore, the application of speech synthesis or audio generation to voice service scenarios is becoming increasingly widespread.

[0003] Because traditional voice customer service requires a high number of human agents, it not only increases a company's human resource costs but also faces language limitations. While most companies in the industry have made improvements to their customer service departments, increasing the proportion of intelligent voice customer service, this only addresses the issue of labor resource costs. Recently, an increasing number of banks and financial institutions are seeking to internationalize and provide customer service to an international audience. However, current audio generation often faces language barriers and remains unable to meet the diverse service needs. Summary of the Invention

[0004] The purpose of the embodiments of the present application is to propose an audio generation method, apparatus, device and storage medium thereof to solve the problem that existing financial institutions, when going international, have language barriers in customer service audio generation and are still unable to meet multi-scale service needs.

[0005] In order to solve the above technical problems, the embodiments of the present application provide an audio generation method, which adopts the following technical solutions:

[0006] An audio generation method comprises the following steps:

[0007] Obtain target text data for audio generation;

[0008] Inputting the target text data into a pre-trained audio generator, wherein the audio generator has a multi-scale acoustic feature generation function, wherein the multi-scale acoustic features include acoustic features of multiple time scales, acoustic features of multiple frequency scales, acoustic features of multiple language scales, and acoustic features of multiple emotion scales;

[0009] Identifying a scale selection instruction obtained through human-computer interaction, wherein the scale selection instruction includes a target time length, a target frequency range, target language information, and a target emotion type of the desired audio data;

[0010] Based on the scale selection instruction, desired audio data output by the audio generator is obtained.

[0011] Furthermore, before executing the step of inputting the target text data into the pre-trained audio generator, the method further includes:

[0012] Obtain training samples at multiple time scales, multiple frequency scales, multiple language scales, and multiple emotion scales;

[0013] Dividing the training sample according to multiple time scales to obtain a first subsample data set consisting of audio training data of different time lengths;

[0014] Dividing the training samples according to multiple frequency scales to obtain a second subsample data set consisting of audio training data in different frequency ranges;

[0015] Dividing the training samples according to a multilingual scale to obtain a third subsample data set consisting of audio training data of different language information;

[0016] Dividing the training samples according to multiple emotion scales to obtain a fourth subsample data set consisting of audio training data of different emotion types;

[0017] Inputting the first sub-sample data set, the second sub-sample data set, the third sub-sample data set and the fourth sub-sample data set into an audio generator to be trained;

[0018] The audio generator to be trained is trained using an adversarial training network to obtain the pre-trained audio generator.

[0019] Furthermore, the step of using an adversarial training network to train the audio generator to be trained to obtain the pre-trained audio generator specifically includes:

[0020] Generate simulated audio for the training sample using the audio generator to be trained, to obtain a simulated audio set corresponding to the training sample;

[0021] Using the discriminator in the adversarial training network to determine the difference between the training sample and the simulated audio set;

[0022] If the difference does not meet the preset difference threshold, a cyclic optimization method is adopted to adjust the adversarial training parameters of the audio generator to be trained and the discriminator respectively, and re-perform adversarial learning training until the difference meets the preset difference threshold;

[0023] If the differences all meet the preset difference threshold, the pre-training of the audio generator to be trained is completed.

[0024] Furthermore, the step of generating simulated audio for the training sample by the audio generator to be trained to obtain a simulated audio set corresponding to the training sample specifically includes:

[0025] Extracting temporal features of all audio training data in the first subsample dataset through a multi-scale feature extraction layer in the audio generator, wherein the multi-scale feature extraction layer is composed of a plurality of parallel convolutional layers;

[0026] Extracting frequency features of all audio training data in the second sub-sample dataset through the multi-scale feature extraction layer;

[0027] Extracting language features of all audio training data in the third subsample dataset through the multi-scale feature extraction layer;

[0028] Extracting emotional features of all audio training data in the fourth subsample dataset through the multi-scale feature extraction layer;

[0029] Based on the time features, frequency features, language features and emotional features corresponding to each audio training data, simulated audio is generated for all audio training data to obtain simulated audio corresponding to all audio training data.

[0030] Furthermore, the step of generating simulated audio for all audio training data based on the time feature, frequency feature, language feature, and emotion feature corresponding to each audio training data to obtain simulated audio corresponding to all audio training data specifically includes:

[0031] Adopting a cross-layer connection mechanism in the audio generator to transfer time features, frequency features, language features, and emotion features corresponding to the same audio training data, wherein the cross-layer connection mechanism is used to transfer multi-scale features between different convolutional layers;

[0032] For the time features, frequency features, language features and emotional features corresponding to each audio training data, a cascade generation method is used to generate the corresponding simulated audio, and the simulated audio corresponding to all audio training data is obtained.

[0033] Furthermore, the discriminator includes a local discriminator and a global discriminator, and the step of using the discriminator in the adversarial training network to discriminate the difference between the training sample and the simulated audio set specifically includes:

[0034] Extracting time features, frequency features, language features, and emotion features from all the analog audios in the analog audio set to obtain time features, frequency features, language features, and emotion features corresponding to all the analog audios;

[0035] Using the local discriminator, the difference between the training sample and the simulated audio set in terms of time characteristics, frequency characteristics, language characteristics and emotional characteristics is determined;

[0036] Based on the differences between the training samples and the simulated audio set in time features, frequency features, language features and emotional features, combined with a preset multi-scale loss function, the global discriminator is used to determine the global differences between the training samples and the simulated audio set.

[0037] Furthermore, before executing the step of identifying the scale selection instruction obtained through human-computer interaction, the method further includes:

[0038] Start an optional operation interface provided in advance for audio generation;

[0039] Based on the operator's click selection instruction, determining the audio time length, audio frequency range, audio language information and audio emotion type of the expected audio data when generating audio;

[0040] The audio time length, audio frequency range, audio language information and audio emotion type are used as scale selection fields to generate corresponding scale selection instructions;

[0041] Sending the scale selection instruction to a preset audio generation processing end through human-computer interaction;

[0042] The step of identifying the scale selection instruction obtained through human-computer interaction specifically includes:

[0043] The audio generation processing end identifies the audio time length, audio frequency range, audio language information and audio emotion type by parsing the scale selection instruction.

[0044] In order to solve the above technical problems, the embodiment of the present application also provides an audio generation device, which adopts the following technical solution:

[0045] An audio generating device, comprising:

[0046] A target text data acquisition module is used to acquire target text data for audio generation;

[0047] A target text data input module is used to input the target text data into a pre-trained audio generator, wherein the audio generator has a multi-scale acoustic feature generation function, and the multi-scale acoustic features include acoustic features of multiple time scales, acoustic features of multiple frequency scales, acoustic features of multiple language scales, and acoustic features of multiple emotion scales;

[0048] a scale selection instruction recognition module, configured to recognize a scale selection instruction obtained through human-computer interaction, wherein the scale selection instruction includes a target time length, a target frequency range, target language information, and a target emotion type of the desired audio data;

[0049] The expected audio data generating module is used to obtain the expected audio data output by the audio generator based on the scale selection instruction.

[0050] In order to solve the above technical problems, the embodiment of the present application further provides a computer device, which adopts the following technical solution:

[0051] A computer device includes a memory and a processor, wherein the memory stores computer-readable instructions, and the processor implements the steps of the above-mentioned audio generation method when executing the computer-readable instructions.

[0052] In order to solve the above technical problems, the embodiment of the present application further provides a computer-readable storage medium, which adopts the following technical solution:

[0053] A computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the audio generation method described above.

[0054] Compared with the prior art, the embodiments of the present application have the following beneficial effects:

[0055] The audio generation method described in the embodiment of the present application obtains target text data for audio generation; inputs the target text data into a pre-trained audio generator; identifies a scale selection instruction obtained through human-computer interaction; and obtains the expected audio data output by the audio generator based on the scale selection instruction. The audio generation method described in the present application is applied to multi-scale audio generation scenarios, especially in multi-language broadcasting or intelligent voice customer service follow-up scenarios. It can generate more detailed and high-quality multi-language transliterations based on differences in audio language, audio time, and audio frequency range, which is more automated and intelligent, and provides more international broadcasting or inquiry services for customers of different languages. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] In order to more clearly illustrate the solutions in this application, a brief introduction will be given below to the drawings required for use in the description of the embodiments of this application. Obviously, the drawings described below are some embodiments of this application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0057] Figure 1is an exemplary system architecture diagram to which the present application may be applied;

[0058] Figure 2 is a flowchart of an embodiment of an audio generation method according to the present application;

[0059] Figure 3 This is a flowchart of a specific embodiment of pre-training an audio generator in the audio generation method described in this application;

[0060] Figure 4 yes Figure 3 A flowchart of a specific embodiment of step 307 is shown;

[0061] Figure 5 yes Figure 4 A flowchart of a specific embodiment of step 401 is shown;

[0062] Figure 6 yes Figure 5 A flowchart of a specific embodiment of step 505 is shown;

[0063] Figure 7 yes Figure 4 A flowchart of a specific embodiment of step 402 is shown;

[0064] Figure 8 yes Figure 7 A flowchart of a specific embodiment of step 703 is shown;

[0065] Figure 9 is a flowchart of a specific embodiment of generating and processing a scale selection instruction in the audio generation method described in this application;

[0066] Figure 10 is a structural diagram of an embodiment of an audio generating device according to the present application;

[0067] Figure 11 It is a structural diagram of an embodiment of a computer device according to the present application. DETAILED DESCRIPTION

[0068] Unless otherwise defined, all technical and scientific terms used herein have the same meanings as commonly understood by those skilled in the art to which this application belongs. The terms used in the specification of the application are for the purpose of describing specific embodiments only and are not intended to limit this application. The terms "including" and "having" and any variations thereof in the specification and claims of this application and the above-mentioned drawings are intended to cover non-exclusive inclusions. The terms "first", "second", etc. in the specification and claims of this application or the above-mentioned drawings are used to distinguish different objects, not to describe a specific order.

[0069] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.

[0070] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings.

[0071] like Figure 1 As shown, system architecture 100 may include a terminal device 101, a network 102, and a server 103. Terminal device 101 may be a laptop computer 1011, a tablet computer 1012, or a mobile phone 1013. Network 102 is a medium for providing a communication link between terminal device 101 and server 103. Network 102 may include various connection types, such as wired or wireless communication links or fiber optic cables.

[0072] The user can use the terminal device 101 to interact with the server 103 via the network 102 to receive or send messages, etc. Various communication client applications can be installed on the terminal device 101, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.

[0073] The terminal device 101 can be various electronic devices with a display screen and supporting web browsing. In addition to the laptop computer 1011, tablet computer 1012 or mobile phone 1013, the terminal device 101 can also be an e-book reader, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 (Moving Picture Experts Group Audio Layer IV) player, a laptop computer and a desktop computer, etc.

[0074] The server 103 may be a server that provides various services, such as a background server that provides support for web pages displayed on the terminal device 101 .

[0075] It should be noted that the audio generation method provided in the embodiment of the present application is generally executed by a server, and accordingly, the audio generation device is generally set in the server.

[0076] It should be understood that Figure 1The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.

[0077] Continue to refer Figure 2 , shows a flow chart of an embodiment of an audio generation method according to the present application. The audio generation method comprises the following steps:

[0078] Step 201: Acquire target text data for audio generation.

[0079] In this embodiment, the target text data includes speech text data in a broadcast scenario, or speech text data when an intelligent voice customer service conducts a telemarketing return visit;

[0080] The audio generation method provided in this application generates audio from speech text data in broadcast scenarios. The generated audio can be applied to actual scenarios such as bus broadcasts, subway broadcasts, train carriage broadcasts, airport broadcasts, and bank service broadcasts, facilitating more detailed and high-quality broadcast services for customers.

[0081] Of course, through the audio generation method provided in this application, audio can be generated for the text data of the speech during the telephone sales follow-up of the intelligent voice customer service, and the generated audio can be applied to the telephone sales follow-up of the intelligent voice customer service, ensuring that in the field of financial technology services, more detailed and high-quality broadcast services are provided to financial customers.

[0082] Step 202: Input the target text data into a pre-trained audio generator, wherein the audio generator has a multi-scale acoustic feature generation function, and the multi-scale acoustic features include acoustic features of multiple time scales, acoustic features of multiple frequency scales, acoustic features of multiple language scales, and acoustic features of multiple emotion scales.

[0083] Specifically, the audio generator has the function of generating multi-scale acoustic features, and can generate audio with different time scales, frequency scales, language scales and emotional scales for the target text data according to the differences in time scales, frequency scales, language scales and emotional scales. It can generate different service audios according to different customer types, perform audio generation in a more targeted manner, and improve the quality of voice services.

[0084] Step 203 : Identify a scale selection instruction obtained through human-computer interaction, wherein the scale selection instruction includes target time length, target frequency range, target language information, and target emotion type of the desired audio data.

[0085] Specifically, due to the business development needs of enterprises or banks, more and more domestic banks or financial institutions need international services. Therefore, this application introduces scale selection during audio generation, specifically, for example: target time length, target frequency range, target language information and target emotion type, so that financial enterprises and institutions can provide international audio generation services to ensure synchronization with international business development.

[0086] Step 204: Obtain the expected audio data output by the audio generator based on the scale selection instruction.

[0087] In this embodiment, in order to facilitate the generation of the audio generator, when executing step 201, the scale selection instruction can be sent to the audio generator as a generation control parameter.

[0088] In this embodiment, target text data for audio generation is obtained; the target text data is input into a pre-trained audio generator; a scale selection instruction obtained through human-computer interaction is identified; and based on the scale selection instruction, the expected audio data output by the audio generator is obtained. The audio generation method described in this application is applied to multi-scale audio generation scenarios, especially in multilingual broadcasting or intelligent voice customer service follow-up scenarios. It can generate more detailed and high-quality multilingual transliterations based on differences in audio language, audio time, and audio frequency range, which is more automated and intelligent, and provides more international broadcasting or inquiry services for customers of different languages.

[0089] Continue to refer Figure 3 In some optional implementations, a step of pre-training the audio generator is included before step 202. Figure 3 This is a flowchart of a specific embodiment of performing pre-training of an audio generator in the audio generation method described in this application, comprising the following steps:

[0090] Step 301, obtaining training samples of multiple time scales, multiple frequency scales, multiple language scales, and multiple emotion scales;

[0091] Step 302: Divide the training sample according to multiple time scales to obtain a first subsample data set consisting of audio training data of different time lengths;

[0092] Step 303: Divide the training sample according to multiple frequency scales to obtain a second subsample data set consisting of audio training data in different frequency ranges;

[0093] Step 304: Divide the training samples according to a multilingual scale to obtain a third subsample data set consisting of audio training data of different language information;

[0094] Step 305: Divide the training sample according to multiple emotion scales to obtain a fourth subsample data set consisting of audio training data of different emotion types;

[0095] Step 306: input the first sub-sample data set, the second sub-sample data set, the third sub-sample data set, and the fourth sub-sample data set into an audio generator to be trained;

[0096] Step 307: Use an adversarial training network to train the audio generator to be trained to obtain the pre-trained audio generator.

[0097] In this embodiment, the training samples are first divided into multiple time scales, multiple frequency scales, multiple language scales and multiple emotion scales to obtain a first sub-sample data set, a second sub-sample data set, a third sub-sample data set and a fourth sub-sample data set. Afterwards, the first sub-sample data set, the second sub-sample data set, the third sub-sample data set and the fourth sub-sample data set are input into the audio generator to be trained and learned to perform pre-training of the audio generator, so as to ensure that the subsequently pre-trained audio generator can generate audio that meets specific time length requirements, frequency requirements, language requirements and emotion requirements.

[0098] Continue to refer Figure 4 , Figure 4 yes Figure 3 The flowchart of a specific embodiment of step 307 shown includes the following steps:

[0099] Step 401: Generate simulated audio for the training sample using the audio generator to be trained, to obtain a simulated audio set corresponding to the training sample;

[0100] Step 402: using a discriminator in the adversarial training network to discriminate the difference between the training sample and the simulated audio set;

[0101] Step 403: If the difference does not meet the preset difference threshold, a cyclic optimization method is used to adjust the adversarial training parameters of the audio generator to be trained and the discriminator respectively, and the adversarial learning training is repeated until the difference meets the preset difference threshold.

[0102] Specifically, it should be understood that until the differences meet the preset difference thresholds, different difference thresholds can be set according to different scale characteristics. For example, a first difference threshold is set for the difference based on the time feature, a second difference threshold is set for the frequency feature, a third difference threshold is set for the language feature, and a fourth difference threshold is set for the emotional feature. The specific settings are determined by actual needs and will not be elaborated here.

[0103] Step 404: If the differences all meet the preset difference threshold, the pre-training of the audio generator to be trained is completed.

[0104] Specifically, the audio generator is subjected to cyclic optimization adversarial training by adopting an adversarial training network, and during the cyclic optimization process, the adversarial training parameters of the audio generator to be trained and the discriminator are continuously adjusted to ensure the high availability of the audio generator that is finally pre-trained.

[0105] Continue to refer Figure 5 , Figure 5 yes Figure 4 The flowchart of a specific embodiment of step 401 includes the following steps:

[0106] Step 501: extracting temporal features of all audio training data in the first subsample dataset through a multi-scale feature extraction layer in the audio generator, wherein the multi-scale feature extraction layer is composed of a plurality of parallel convolutional layers;

[0107] Step 502: extracting frequency features of all audio training data in the second subsample dataset through the multi-scale feature extraction layer;

[0108] Step 503: extracting language features of all audio training data in the third subsample dataset through the multi-scale feature extraction layer;

[0109] Step 504: extracting emotional features of all audio training data in the fourth subsample dataset through the multi-scale feature extraction layer;

[0110] Step 505 : Based on the time feature, frequency feature, language feature, and emotion feature corresponding to each audio training data, simulated audio is generated for all audio training data to obtain simulated audio corresponding to all audio training data.

[0111] In this embodiment, several parallel convolutional layers are used in the audio generator to form a multi-scale feature extraction layer. During the pre-training stage of the audio generator, adversarial learning training is performed using a sub-sample data set that has been divided according to different scales. This enables the audio generator to generate audio according to the selection requirements of different scales during actual audio generation, thereby ensuring that the multi-scale feature extraction layer is fully trained.

[0112] Furthermore, the system can be flexibly expanded by adding or reducing parallel convolutional layers based on the number of features at different scales. For example, if there are five languages, the number of parallel convolutional layers can be set to five. Later, if the number of languages ​​increases to 20, additional parallel convolutional layers can be added for supplementary adversarial learning training without destroying the original pre-training results. This allows for flexible expansion of the audio generator.

[0113] Continue to refer Figure 6 , Figure 6 yes Figure 5 The flowchart of a specific embodiment of step 505 shown includes the following steps:

[0114] Step 601: Using a cross-layer connection mechanism in the audio generator to transfer time features, frequency features, language features, and emotion features corresponding to the same audio training data, wherein the cross-layer connection mechanism is used to transfer multi-scale features between different convolutional layers;

[0115] Step 602: For each audio training data corresponding to the time feature, frequency feature, language feature and emotion feature, a cascade generation method is used to generate the corresponding simulated audio, thereby obtaining the simulated audio corresponding to all the audio training data.

[0116] Specifically, the audio generator also introduces a cross-layer connection mechanism and a cascade generation mechanism for transferring multi-scale features between different convolutional layers, and performing cascade generation based on the multi-scale features to generate corresponding simulated audio. The cascade generation mechanism is implemented by introducing several cascade generation units.

[0117] In this embodiment, the discriminator includes a local discriminator and a global discriminator. The local discriminator includes a first local discriminator, a second local discriminator, a third local discriminator, and a fourth local discriminator. It should be understood that different local discriminators discriminate the difference between the simulated audio and the actual training samples based on features at different scales. By performing difference discrimination based on features at different scales and globally, the generated audio data is guaranteed to be more realistic, detailed, high-quality, and close to real speech in subsequent practical applications.

[0118] Continue to refer Figure 7 , Figure 7 yes Figure 4 The flowchart of a specific embodiment of step 402 includes the following steps:

[0119] Step 701: extracting time features, frequency features, language features, and emotion features from all analog audios in the analog audio set to obtain time features, frequency features, language features, and emotion features corresponding to all analog audios.

[0120] Step 702: using the local discriminator, determining the differences between the training sample and the simulated audio set in terms of time features, frequency features, language features, and emotional features;

[0121] Step 703: Based on the differences between the training samples and the simulated audio set in terms of time features, frequency features, language features, and emotional features, combined with a preset multi-scale loss function, the global discriminator is used to determine the global differences between the training samples and the simulated audio set.

[0122] Specifically, the preset multi-scale loss function can be a weighted sum function designed according to different scale features. The multi-scale loss function is used to calculate the total difference, that is, the global difference, based on the differences corresponding to different scale features.

[0123] Continue to refer Figure 8 , Figure 8 yes Figure 7 The flowchart of a specific embodiment of step 702 includes the following steps:

[0124] Step 801: Using the first local discriminator, based on the time features corresponding to all simulated audios and the time features of all audio training data extracted by the multi-scale feature extraction layer in the first subsample dataset, discriminate the difference in time features between the training samples and the simulated audio set.

[0125] Step 802: Using the second local discriminator, based on the frequency features corresponding to all the simulated audios and the frequency features of all the audio training data extracted by the multi-scale feature extraction layer in the second sub-sample dataset, discriminate the difference in frequency features between the training samples and the simulated audio set.

[0126] Step 803: Using the third local discriminator, based on the language features corresponding to all simulated audios and the language features of all audio training data extracted by the multi-scale feature extraction layer in the third subsample dataset, discriminate the difference in language features between the training samples and the simulated audio set.

[0127] Step 804: Based on the emotional features corresponding to all simulated audios and the emotional features of all audio training data extracted by the multi-scale feature extraction layer in the fourth sub-sample data set, the fourth local discriminator is used to discriminate the difference in emotional features between the training sample and the simulated audio set.

[0128] By performing difference judgment on features at different scales and globally, the generated audio data is ensured to be more realistic, detailed, high-quality and close to real speech in subsequent practical applications.

[0129] Continue to refer Figure 9 In some optional implementations, before step 203, a step of generating and processing a scale selection instruction is also included. Figure 9 This is a flowchart of a specific embodiment of generating and processing a scale selection instruction in the audio generation method described in this application, including the following steps:

[0130] Step 901, starting an optional operation interface provided in advance for audio generation;

[0131] Step 902: Determine the audio time length, audio frequency range, audio language information, and audio emotion type of the desired audio data when generating audio based on the operator's click selection instruction.

[0132] Step 903: Using the audio time length, audio frequency range, audio language information, and audio emotion type as scale selection fields, a corresponding scale selection instruction is generated;

[0133] Step 904: Send the scale selection instruction to a preset audio generation processing terminal through human-computer interaction;

[0134] In this embodiment, the step of identifying the scale selection instruction obtained through human-computer interaction specifically includes:

[0135] In step 905 , the audio generation processing end identifies the audio time length, audio frequency range, audio language information, and audio emotion type by parsing the scale selection instruction.

[0136] In this embodiment, the step of obtaining the expected audio data output by the audio generator based on the scale selection instruction specifically includes: using the identified audio time length, audio frequency range, audio language information and audio emotion type as generation control parameters of the audio generator to control the audio generator to generate the expected audio data.

[0137] By providing an optional audio generation operating interface, it is possible to generate corresponding audio based on customer needs to meet the service needs of customers in different regions and languages, making financial services more and more international.

[0138] This application obtains target text data for audio generation; inputs the target text data into a pre-trained audio generator; identifies the scale selection instruction obtained through human-computer interaction; and obtains the expected audio data output by the audio generator based on the scale selection instruction. The audio generation method described in this application is applied to multi-scale audio generation scenarios, especially in multi-language broadcasting or intelligent voice customer service follow-up scenarios. It can generate more detailed and high-quality multi-language transliterations based on the differences in audio language, audio time, and audio frequency range, which is more automated and intelligent, and provides more international broadcasting or inquiry services for customers of different languages.

[0139] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to achieve optimal results.

[0140] Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.

[0141] In an embodiment of the present application, target text data for audio generation is obtained; the target text data is input into a pre-trained audio generator; a scale selection instruction obtained through human-computer interaction is identified; and based on the scale selection instruction, the expected audio data output by the audio generator is obtained. The audio generation method described in this application is applied to multi-scale audio generation scenarios, especially in multi-language broadcasting or intelligent voice customer service follow-up scenarios. It can generate more detailed and high-quality multi-language transliterations based on differences in audio languages, audio time, and audio frequency ranges, which is more automated and intelligent, and provides more international broadcasting or inquiry services for customers of different languages.

[0142] Further references Figure 10 , as a response to the above Figure 2 The present application provides an embodiment of an audio generation device, which is similar to the embodiment of the present invention. Figure 2Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.

[0143] like Figure 10 As shown, the audio generating device 10 of this embodiment includes: a target text data acquisition module 10a, a target text data input module 10b, a scale selection instruction recognition module 10c and an expected audio data generation module 10d.

[0144] A target text data acquisition module 10a is used to acquire target text data for audio generation;

[0145] A target text data input module 10b is configured to input the target text data into a pre-trained audio generator, wherein the audio generator has a function of generating multi-scale acoustic features, including acoustic features at multiple time scales, acoustic features at multiple frequency scales, acoustic features at multiple language scales, and acoustic features at multiple emotion scales;

[0146] The scale selection instruction recognition module 10c is used to recognize the scale selection instruction obtained through human-computer interaction, wherein the scale selection instruction includes the target time length, target frequency range, target language information and target emotion type of the desired audio data;

[0147] The expected audio data generating module 10d is configured to obtain the expected audio data output by the audio generator based on the scale selection instruction.

[0148] This application obtains target text data for audio generation; inputs the target text data into a pre-trained audio generator; identifies the scale selection instruction obtained through human-computer interaction; and obtains the expected audio data output by the audio generator based on the scale selection instruction. The audio generation method described in this application is applied to multi-scale audio generation scenarios, especially in multi-language broadcasting or intelligent voice customer service follow-up scenarios. It can generate more detailed and high-quality multi-language transliterations based on the differences in audio language, audio time, and audio frequency range, which is more automated and intelligent, and provides more international broadcasting or inquiry services for customers of different languages.

[0149] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing related hardware via computer-readable instructions. The computer-readable instructions can be stored in a computer-readable storage medium, and when the program is executed, it can include the processes in the above-described method embodiments. The aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).

[0150] It should be understood that although the steps in the flowcharts of the accompanying drawings are shown in sequence as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some of the steps in the flowcharts of the accompanying drawings may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed in turn or alternately with other steps or at least a portion of the sub-steps or stages of other steps.

[0151] To solve the above technical problems, the present application also provides a computer device. Figure 11 , Figure 11 This is a basic structural block diagram of the computer device in this embodiment.

[0152] The computer device 11 includes a memory 11a, a processor 11b, and a network interface 11c that are interconnected via a system bus. Figure 11 Only a computer device 11 having components such as a memory 11a, a processor 11b, and a network interface 11c is shown. However, it should be understood that implementation of all illustrated components is not required, and more or fewer components may be implemented instead. Those skilled in the art will understand that a computer device herein is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to a microprocessor, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a digital signal processor (DSP), an embedded device, and the like.

[0153] The computer device may be a desktop computer, notebook computer, PDA, cloud server, etc. The computer device may interact with the user via a keyboard, mouse, remote control, touchpad, or voice control device.

[0154] The memory 11a includes at least one type of readable storage medium, including flash memory, a hard disk, a multimedia card, a card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic storage, a magnetic disk, an optical disk, etc. In some embodiments, the memory 11a may be an internal storage unit of the computer device 11, such as the hard disk or memory of the computer device 11. In other embodiments, the memory 11a may also be an external storage device of the computer device 11, such as a plug-in hard disk equipped on the computer device 11, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. Of course, the memory 11a may also include both the internal storage unit of the computer device 11 and its external storage device. In this embodiment, the memory 11a is generally used to store the operating system and various application software installed on the computer device 11, such as computer-readable instructions for an audio generation method. In addition, the memory 11a can also be used to temporarily store various types of data that have been output or are to be output.

[0155] In some embodiments, the processor 11b may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip. The processor 11b is generally used to control the overall operation of the computer device 11. In this embodiment, the processor 11b is used to execute computer-readable instructions or process data stored in the memory 11a, such as computer-readable instructions for executing the audio generation method.

[0156] The network interface 11 c may include a wireless network interface or a wired network interface. The network interface 11 c is generally used to establish a communication connection between the computer device 11 and other electronic devices.

[0157] The computer device proposed in this embodiment belongs to the field of R&D design and audio processing technology, and is used in audio generation scenarios. This application obtains target text data for audio generation; inputs the target text data into a pre-trained audio generator; identifies the scale selection instruction obtained through human-computer interaction; and obtains the expected audio data output by the audio generator based on the scale selection instruction. The audio generation method described in this application is applied to multi-scale audio generation scenarios, especially in multi-language broadcasting or intelligent voice customer service follow-up scenarios. It can generate more detailed and high-quality multi-language transliterations based on the differences in audio language, audio time, and audio frequency range, which is more automated and intelligent, and provides more international broadcasting or inquiry services for customers of different languages.

[0158] The present application also provides another embodiment, namely, providing a computer-readable storage medium, wherein the computer-readable storage medium stores computer-readable instructions, and the computer-readable instructions can be executed by a processor to enable the processor to perform the steps of the audio generation method as described above.

[0159] The computer-readable storage medium proposed in this embodiment belongs to the field of R&D design and audio processing technology, and is applied to audio generation scenarios. This application obtains target text data for audio generation; inputs the target text data into a pre-trained audio generator; identifies the scale selection instruction obtained through human-computer interaction; and obtains the expected audio data output by the audio generator based on the scale selection instruction. The audio generation method described in this application is applied to multi-scale audio generation scenarios, especially in multi-language broadcasting or intelligent voice customer service follow-up scenarios. It can generate more detailed and high-quality multi-language transliterations based on the differences in audio language, audio time, and audio frequency range, which is more automated and intelligent, and provides more international broadcasting or inquiry services for customers of different languages.

[0160] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present application.

[0161] Obviously, the embodiments described above are only some of the embodiments of the present application, rather than all of the embodiments. The preferred embodiments of the present application are given in the accompanying drawings, but they do not limit the patent scope of the present application. The present application can be implemented in many different forms. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosure of the present application more thorough and comprehensive. Although the present application has been described in detail with reference to the aforementioned embodiments, for those skilled in the art, it is still possible to modify the technical solutions described in the aforementioned specific embodiments, or to make equivalent replacements for some of the technical features therein. Any equivalent structure made using the contents of the present application specification and the accompanying drawings, directly or indirectly used in other related technical fields, is also within the scope of patent protection of the present application.

Claims

1. An audio generation method, characterized in that: The steps include: Obtain target text data for audio generation; Inputting the target text data into a pre-trained audio generator, wherein the audio generator has a multi-scale acoustic feature generation function, wherein the multi-scale acoustic features include acoustic features of multiple time scales, acoustic features of multiple frequency scales, acoustic features of multiple language scales, and acoustic features of multiple emotion scales; Identifying a scale selection instruction obtained through human-computer interaction, wherein the scale selection instruction includes a target time length, a target frequency range, target language information, and a target emotion type of the desired audio data; Based on the scale selection instruction, obtaining the desired audio data output by the audio generator, The audio generator uses several parallel convolutional layers to form a multi-scale feature extraction layer, and introduces a cross-layer connection mechanism and a cascade generation mechanism. During the pre-training phase of the audio generator, a cross-layer connection mechanism in the audio generator is used to transfer time features, frequency features, language features, and emotion features corresponding to the same audio training data, wherein the cross-layer connection mechanism is used to transfer multi-scale features between different convolutional layers; For the time features, frequency features, language features and emotional features corresponding to each audio training data, a cascade generation method is used to generate the corresponding simulated audio, and the simulated audio corresponding to all audio training data is obtained.

2. The audio generation method according to claim 1, wherein Before executing the step of inputting the target text data into the pre-trained audio generator, the method further includes: Obtain training samples at multiple time scales, multiple frequency scales, multiple language scales, and multiple emotion scales; Dividing the training sample according to multiple time scales to obtain a first subsample data set consisting of audio training data of different time lengths; Dividing the training samples according to multiple frequency scales to obtain a second subsample data set consisting of audio training data in different frequency ranges; Dividing the training samples according to a multilingual scale to obtain a third subsample data set consisting of audio training data of different language information; Dividing the training samples according to multiple emotion scales to obtain a fourth subsample data set consisting of audio training data of different emotion types; Inputting the first sub-sample data set, the second sub-sample data set, the third sub-sample data set and the fourth sub-sample data set into an audio generator to be trained; The audio generator to be trained is trained using an adversarial training network to obtain the pre-trained audio generator.

3. The audio generation method according to claim 2, wherein: The step of using an adversarial training network to train the audio generator to be trained to obtain the pre-trained audio generator specifically includes: Generate simulated audio for the training sample using the audio generator to be trained, to obtain a simulated audio set corresponding to the training sample; Using the discriminator in the adversarial training network to determine the difference between the training sample and the simulated audio set; If the difference does not meet the preset difference threshold, a cyclic optimization method is adopted to adjust the adversarial training parameters of the audio generator to be trained and the discriminator respectively, and re-perform adversarial learning training until the difference meets the preset difference threshold; If the differences all meet the preset difference threshold, the pre-training of the audio generator to be trained is completed.

4. The audio generation method according to claim 3, wherein: The step of generating simulated audio for the training sample by the audio generator to be trained to obtain a simulated audio set corresponding to the training sample specifically includes: Extracting temporal features of all audio training data in the first subsample dataset through a multi-scale feature extraction layer in the audio generator, wherein the multi-scale feature extraction layer is composed of a plurality of parallel convolutional layers; Extracting frequency features of all audio training data in the second sub-sample dataset through the multi-scale feature extraction layer; Extracting language features of all audio training data in the third subsample dataset through the multi-scale feature extraction layer; Extracting emotional features of all audio training data in the fourth subsample dataset through the multi-scale feature extraction layer; Based on the time features, frequency features, language features and emotional features corresponding to each audio training data, simulated audio is generated for all audio training data to obtain simulated audio corresponding to all audio training data.

5. The audio generation method according to claim 3, wherein: The discriminator includes a local discriminator and a global discriminator. The step of using the discriminator in the adversarial training network to discriminate the difference between the training sample and the simulated audio set specifically includes: Extracting time features, frequency features, language features, and emotion features from all the analog audios in the analog audio set to obtain time features, frequency features, language features, and emotion features corresponding to all the analog audios; Using the local discriminator, the difference between the training sample and the simulated audio set in terms of time characteristics, frequency characteristics, language characteristics and emotional characteristics is determined; Based on the differences between the training samples and the simulated audio set in time features, frequency features, language features and emotional features, combined with a preset multi-scale loss function, the global discriminator is used to determine the global differences between the training samples and the simulated audio set.

6. The audio generation method according to claim 1, wherein: Before executing the step of identifying the scale selection instruction obtained through human-computer interaction, the method further includes: Start an optional operation interface provided in advance for audio generation; Based on the operator's click selection instruction, determining the audio time length, audio frequency range, audio language information and audio emotion type of the expected audio data when generating audio; The audio time length, audio frequency range, audio language information and audio emotion type are used as scale selection fields to generate corresponding scale selection instructions; Sending the scale selection instruction to a preset audio generation processing end through human-computer interaction; The step of identifying the scale selection instruction obtained through human-computer interaction specifically includes: The audio generation processing end identifies the audio time length, audio frequency range, audio language information and audio emotion type by parsing the scale selection instruction.

7. An audio generating device, characterized in that: include: A target text data acquisition module is used to acquire target text data for audio generation; A target text data input module is used to input the target text data into a pre-trained audio generator, wherein the audio generator has a multi-scale acoustic feature generation function, and the multi-scale acoustic features include acoustic features of multiple time scales, acoustic features of multiple frequency scales, acoustic features of multiple language scales, and acoustic features of multiple emotion scales; a scale selection instruction recognition module, configured to recognize a scale selection instruction obtained through human-computer interaction, wherein the scale selection instruction includes a target time length, a target frequency range, target language information, and a target emotion type of the desired audio data; an expected audio data generating module, configured to obtain the expected audio data output by the audio generator based on the scale selection instruction, The audio generator uses several parallel convolutional layers to form a multi-scale feature extraction layer, and introduces a cross-layer connection mechanism and a cascade generation mechanism. During the pre-training phase of the audio generator, a cross-layer connection mechanism in the audio generator is used to transfer time features, frequency features, language features, and emotion features corresponding to the same audio training data, wherein the cross-layer connection mechanism is used to transfer multi-scale features between different convolutional layers; For the time features, frequency features, language features and emotional features corresponding to each audio training data, a cascade generation method is used to generate the corresponding simulated audio, and the simulated audio corresponding to all audio training data is obtained.

8. A computer device, characterized in that: The method comprises a memory and a processor, wherein the memory stores computer-readable instructions, and the processor implements the steps of the audio generation method according to any one of claims 1 to 6 when executing the computer-readable instructions.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the audio generation method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Image segmentation method and device and terminal equipment

    CN111047602A

  • Speech synthesis model training method and device, electronic equipment and storage medium

    CN114512112A