Voice generation method and device, equipment and storage medium

By introducing an auxiliary speech insertion planner and a speech decoder, the problem of lack of emotion and realism in speech generation in existing technologies is solved, and the generated speech is more realistic.

CN121789685APending Publication Date: 2026-04-03PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-08
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing speech generation methods cannot intelligently insert auxiliary speech, resulting in unrealistic and emotionless output.

Method used

An auxiliary speech insertion planner, trained through learning, combines word-level speech text with auxiliary speech, and then decodes and outputs the speech through a speech decoder to generate realistic speech.

Benefits of technology

It enables the intelligent insertion of auxiliary voice during speech generation, making the generated speech closer to the pronunciation of real people and enhancing the emotion and realism.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121789685A_ABST
    Figure CN121789685A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of artificial intelligence, and relates to a voice generation method and device, equipment and a storage medium. Inputting the word-level voice text into an auxiliary voice insertion planner after learning training is completed; acquiring a voice generation coding sequence output by the auxiliary voice insertion planner; and performing decoding output processing on the voice generation coding sequence by adopting a preset voice decoder to obtain a target generation voice. By introducing the auxiliary voice insertion planner, the voice which cannot be directly converted into the natural language is inserted into the word-level voice text to be subjected to voice generation, so that the fidelity of the finally generated voice is ensured, and the voice is closer to the real character pronunciation. The method is applied to a consultation reply scene in a financial science and technology service or a health medical service, and can assist an intelligent AI customer service in replacing manual response to generate a pronunciation effect closer to a real person.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology and is applied to the scenario of voice generation for AI question-answering agents. It relates to a voice generation method, device, equipment and storage medium. Background Technology

[0002] Currently, voice interaction systems, such as audio content generation and virtual human dialogue, have made significant progress in understanding and generating human speech. However, their core focus remains on vocabulary and grammar, severely neglecting the indispensable non-verbal components of communication, such as laughter, sighs, and coughs. These auxiliary voices carry rich emotions, intentions, and social signals, and their absence makes human-computer interaction appear stiff, lacking in emotion and realism.

[0003] In recent years, voice interaction systems have been gradually applied to consultation or marketing scenarios in fintech or healthcare. However, existing voice generation methods are mechanical and lack intelligence, failing to intelligently insert auxiliary voice based on the dialogue content, resulting in voice generation results that are not realistic enough. Summary of the Invention

[0004] The purpose of this application is to provide a speech generation method, apparatus, device, and storage medium to solve the technical problem that existing speech generation methods cannot intelligently insert auxiliary speech according to the dialogue content, resulting in speech generation results that are not realistic enough.

[0005] Firstly, embodiments of this application provide a speech generation method, which employs the following technical solution: A speech generation method includes the following steps: Obtain the word-level speech text to be generated; The word-level speech text is input into the auxiliary speech insertion planner that has been trained, wherein the auxiliary speech includes speech that cannot be directly converted into natural language; Obtain the speech generation coding sequence output by the auxiliary speech insertion planner; The speech generation encoding sequence is decoded and output using a preset speech decoder to obtain the target generated speech.

[0006] Secondly, embodiments of this application also provide a speech generation device, which adopts the technical solution described below: A speech generation device, comprising: The word-level speech-text acquisition module is used to acquire the word-level speech-text to be generated. A word-level speech-text input module is used to input the word-level speech-text into an auxiliary speech insertion planner that has been trained, wherein the auxiliary speech includes speech that cannot be directly converted into natural language; The speech generation coding sequence acquisition module is used to acquire the speech generation coding sequence output by the auxiliary speech insertion planner; The target generated speech acquisition module is used to decode and output the speech generation encoding sequence using a preset speech decoder to obtain the target generated speech.

[0007] Thirdly, embodiments of this application also provide a computer device that adopts the technical solution described below: A computer device includes a memory and a processor, wherein the memory stores computer-readable instructions, and the processor executes the computer-readable instructions to implement the steps of the speech generation method described above.

[0008] Fourthly, embodiments of this application also provide a computer-readable storage medium, which adopts the technical solutions described below: A computer-readable storage medium storing computer-readable instructions that, when executed by a processor, implement the steps of the speech generation method described above.

[0009] Compared with the prior art, the embodiments of this application have the following main advantages: The speech generation method described in this application involves: acquiring word-level speech text to be generated; inputting the word-level speech text into a trained auxiliary speech insertion planner, wherein the auxiliary speech includes speech that cannot be directly converted into natural language; acquiring the speech generation encoding sequence output by the auxiliary speech insertion planner; and decoding the speech generation encoding sequence using a preset speech decoder to obtain the target generated speech. By introducing an auxiliary speech insertion planner, speech that cannot be directly converted into natural language is inserted into the word-level speech text to be generated, ensuring the realism of the final generated speech and making it closer to the pronunciation of a real person. Applying this method to consultation and response scenarios in fintech or healthcare businesses can assist intelligent AI customer service in replacing human responses, generating pronunciation effects that are closer to those of a real person. Attached Figure Description

[0010] To more clearly illustrate the solutions in this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0011] Figure 1 This is an exemplary system architecture diagram to which this application can be applied; Figure 2 This is a flowchart of an embodiment of a speech generation method according to this application; Figure 3 This is a flowchart of a specific embodiment of the speech generation method described in this application, which involves learning and training an auxiliary speech insertion planner. Figure 4 yes Figure 3 A flowchart of a specific embodiment of step 302 shown; Figure 5 yes Figure 3 A flowchart of a specific embodiment of step 303 shown; Figure 6 yes Figure 3 A flowchart of a specific embodiment of step 304 shown; Figure 7 yes Figure 3 A flowchart of a specific embodiment of step 305 shown; Figure 8 yes Figure 7 A flowchart of a specific embodiment of step 705 shown; Figure 9 This is a schematic diagram of one embodiment of a speech generation device according to this application; Figure 10 This is a schematic diagram of the structure of one embodiment of the computer device according to this application. Detailed Implementation

[0012] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein in the specification of the application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application; the terms "comprising" and "having," and any variations thereof, in the specification, claims, and foregoing drawings of this application, are intended to cover non-exclusive inclusion. The terms "first," "second," etc., in the specification, claims, or foregoing drawings of this application are used to distinguish different objects, not to describe a particular order.

[0013] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0014] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.

[0015] like Figure 1 As shown, system architecture 100 may include terminal device 101, network 102, and server 103. Terminal device 101 may be a laptop 1011, tablet 1012, or mobile phone 1013. Network 102 is used as a medium to provide a communication link between terminal device 101 and server 103. Network 102 may include various connection types, such as wired, wireless communication links, or fiber optic cables.

[0016] Users can use terminal device 101 to interact with server 103 via network 102 to receive or send messages, etc. Various communication client applications can be installed on terminal device 101, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social media platform software, etc.

[0017] Terminal device 101 can be various electronic devices with a display screen and support web browsing. In addition to laptops 1011, tablets 1012, or mobile phones 1013, terminal device 101 can also be an e-book reader, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 player (Moving Picture Experts Group Audio Layer IV), a laptop computer, and a desktop computer, etc.

[0018] Server 103 can be a server that provides various services, such as a backend server that provides support for the pages displayed on terminal device 101.

[0019] It should be noted that the speech generation method provided in this application embodiment is generally executed by a server, and correspondingly, a speech generation device is generally set in the server.

[0020] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.

[0021] Continue to refer to Figure 2 The diagram illustrates a flowchart of an embodiment of a speech generation method according to this application. The speech generation method includes the following steps: Step 201: Obtain the word-level speech text to be generated.

[0022] In this embodiment, the word-level speech text includes: first obtaining the sentence-level speech text to be generated, and then performing word-level segmentation on the sentence-level speech text to obtain the word-level speech text. Alternatively, it may be possible to first obtain the paragraph-level speech text to be generated, and then perform sentence-level and word-level segmentation on the paragraph-level speech text sequentially to obtain the word-level speech text.

[0023] Specifically, the word-level speech text includes response scripts in fintech business application scenarios, such as the word-level speech text generated when intelligent AI customer service replaces human voice responses during telemarketing; it also includes response scripts in healthcare business scenarios, such as the word-level speech text generated when intelligent doctor avatars provide voice responses during online health consultations.

[0024] In this embodiment, the speech generation method is applied to consultation and response scenarios in fintech or healthcare businesses, which can assist intelligent AI customer service in speech generation, replacing human responses and making the system more intelligent.

[0025] Step 202: Input the word-level speech text into the auxiliary speech insertion planner that has completed learning and training, wherein the auxiliary speech includes speech that cannot be directly converted into natural language.

[0026] In this embodiment, the auxiliary speech, such as laughter, surprise, coughing, etc. during voice dialogue, specifically refers to a series of speech sounds that cannot be directly converted into natural language; natural language refers to the languages ​​used by humans in daily communication, such as English, Chinese, French, etc.; the auxiliary speech can also be understood as speech that cannot be directly represented by corresponding language text.

[0027] Specifically, the auxiliary speech insertion planner that has been trained can insert speech that cannot be directly represented by the corresponding language text into the word-level speech text, ensuring that the final generated speech is more realistic and closer to the real human pronunciation.

[0028] In this embodiment, the word-level speech text is input into the auxiliary speech insertion planner that has completed learning and training, so that the auxiliary speech insertion planner can insert auxiliary speech into the word-level speech text, ensuring that the final generated speech is more realistic and closer to the real human pronunciation.

[0029] Step 203: Obtain the speech generation coding sequence output by the auxiliary speech insertion planner.

[0030] Specifically, the auxiliary speech insertion planner outputs a speech generation encoding sequence, which includes the encoding sequence corresponding to the word-level speech text, as well as the encoding sequences of all auxiliary speech inserted by the auxiliary speech insertion planner.

[0031] In this embodiment, the auxiliary speech insertion planner not only completes the insertion of all auxiliary speech into the word-level speech text, but also performs acoustic encoding processing on "word-level speech text + all auxiliary speech" before output, ensuring that after output, the target generated speech can be obtained by directly decoding using a speech decoder.

[0032] Step 204: The speech generation encoding sequence is decoded and output using a preset speech decoder to obtain the target generated speech.

[0033] Specifically, the preset speech decoder includes the speech decoder in the preset TTS model, which decodes the speech to generate an encoded sequence and outputs the target generated speech.

[0034] In this embodiment, the process involves: acquiring word-level speech text to be generated; inputting the word-level speech text into a trained auxiliary speech insertion planner, where auxiliary speech includes speech that cannot be directly converted into natural language; acquiring the speech generation encoding sequence output by the auxiliary speech insertion planner; and using a preset speech decoder to decode and output the speech generation encoding sequence to obtain the target generated speech. By introducing the auxiliary speech insertion planner, speech that cannot be directly converted into natural language is inserted into the word-level speech text to be generated, ensuring the realism of the final generated speech and making it closer to the pronunciation of a real person. Applying this method to consultation and response scenarios in fintech or healthcare businesses can assist intelligent AI customer service in replacing human responses, generating pronunciation effects that are closer to those of a real person.

[0035] Continue to refer to Figure 3 In some embodiments, prior to step 202, the speech generation method further includes a step of learning and training an auxiliary speech insertion planner. Figure 3 This is a flowchart of a specific embodiment of the speech generation method described in this application, which involves learning and training an auxiliary speech insertion planner, including: Step 301: Obtain authentic audio from multiple sources; Specifically, original audio can be collected from diverse real media sources such as radio dramas, animations, and talk shows as the multi-source real audio.

[0036] Step 302: Preprocess the real audio to obtain the target speech sequence; Specifically, the target speech sequence can be a speech sequence of a specific person collected specifically, such as a speech sequence of a celebrity collected specifically as the target speech sequence; or it can be a speech sequence collected specifically when introducing a specific type of business knowledge, such as a speech sequence of a financial management interview program selected specifically as the target speech sequence.

[0037] Step 303: Perform speech-to-text conversion on the target speech sequence to obtain word-level speech text with timestamp annotation; Specifically, the target speech sequence can be directly converted into text using an AST model, and then timestamps can be added to obtain word-level speech text with timestamps.

[0038] Step 304: Perform auxiliary speech detection on the target speech sequence to obtain auxiliary speech segments labeled with timestamps; Step 305: Input the timestamped auxiliary speech segments and the timestamped word-level speech texts into the auxiliary speech insertion planner to be trained, and perform auxiliary speech insertion learning training to obtain the auxiliary speech insertion planner after training.

[0039] Continue to refer to Figure 4 , Figure 4 yes Figure 3 A flowchart of a specific embodiment of step 302 shown includes: Step 401: Use a preset voice activity detection tool to perform voice activity detection on the real audio; Specifically, since the actual audio may be conversational audio, meaning there may be multiple speakers, speech activity detection tools such as Emilia-Pipeline can be used to first identify the number of speakers. Then, the actual audio can be categorized and extracted according to the different speakers. Emilia-Pipeline is a speech activity detection tool that can detect the number of pronunciation targets in speech audio.

[0040] Step 402: Based on the voice activity detection results, perform voice log and voice source separation processing on the real audio, and obtain a clean single-speaker voice sequence through the voice source separation processing results; Specifically, the process of separating the voice log and voice source of the real audio includes: extracting voiceprint features from the real audio; and performing clustering and evaluation processing on the real audio based on the different results of voiceprint feature extraction, thereby obtaining the voice text corresponding to each pronunciation object.

[0041] More specifically, the speech source separation can be achieved by directly using independent component analysis to extract the speech waveform of each speaker from the mixed speech signal.

[0042] Step 403: Based on the speech log separation processing results, the pronunciation time information of the clean single-speaker speech sequence is annotated to obtain the speech sequence annotated with the pronunciation time information as the target speech sequence.

[0043] More specifically, the speech log separation, for example, involves assigning different speaker labels to each time point in the audio based on the different results of voiceprint feature clustering, thereby achieving speech log separation.

[0044] In this embodiment, in order to facilitate subsequent recognition processing and alignment between speech segments, after obtaining a clean single-speaker speech sequence, the speech time information of the clean single-speaker speech sequence can be marked according to the actual real audio playback time information, and the speech sequence marked with the speech time information is used as the target speech sequence.

[0045] Continue to refer to Figure 5 , Figure 5 yes Figure 3 A flowchart of a specific embodiment of step 303 shown includes: Step 501: Use a preset ASR model to extract acoustic features from the target speech sequence; Specifically, the preset ASR model is, for example, the Whisper-large-v3 speech-to-text model.

[0046] Step 502: The extracted acoustic features are transcribed into text using the text decoder in the ASR model to obtain the transcribed text; Step 503: Based on the pronunciation time information of the target speech sequence, add word-level timestamps to the transcribed text to obtain word-level speech text annotated with timestamps.

[0047] Specifically, the transcribed text is first divided into word-level segments. Then, combining the pronunciation time information of the target speech sequence, the start and end times of each word's pronunciation are labeled, serving as the corresponding word-level timestamps. Finally, the start and end times of each word's pronunciation are obtained, resulting in the word-level speech text labeled with timestamps.

[0048] In this embodiment, the target speech sequence is processed into speech-to-text to obtain word-level speech text with timestamps, so that the timestamps can be used for subsequent speech alignment processing.

[0049] Continue to refer to Figure 6 , Figure 6 yes Figure 3 A flowchart of a specific embodiment of step 304 shown includes: Step 601: Input the target speech sequence into a preset frame-level auxiliary speech detection model; In this embodiment, the preset frame-level auxiliary speech detection model, for example, is a frame-level auxiliary speech detection model with wav2vec 2.0 as the backbone network. The frame level refers to the speech frame detection level.

[0050] Step 602: Perform multi-layer audio feature extraction on the target speech sequence using the frame-level auxiliary speech detection model; Specifically, the frame-level auxiliary speech detection model with wav2vec 2.0 as the backbone network is used to extract multi-layer audio features from the target speech sequence, resulting in multi-layer audio features. ,in, Indicates the number of audio feature layers.

[0051] Step 603: The extracted multi-layer audio features are fused using a learnable weighted summation method to obtain audio fusion features; Specifically, using an audio feature fusion function: The audio fusion features are obtained, wherein, This indicates the layer number of the current audio feature. , A positive integer greater than 1, representing the number of extracted multi-layer audio feature data. This represents the automatically learned weights of the audio features in the current layer. This represents the audio features of the current layer, which can be understood as an N×1 sequence of audio feature code values. This refers to the audio fusion feature.

[0052] Step 604: Encode the audio fusion features using a preset Transformer encoder to obtain the encoded output result; Specifically, the audio fusion features are encoded using the Transformer encoder to obtain the corresponding encoded output matrix.

[0053] Step 605: Combine a set of preset decodeable vectors to decode the encoded output result and obtain the decoded output result; Specifically, a set of learnable query vectors Q can be used to decode the encoded output matrix to obtain the corresponding decoded output matrix.

[0054] Step 606: Calculate the similarity matrix between the decoding output and the encoding output, and activate them through a preset activation function to obtain frame-level auxiliary speech detection results; Specifically, by comparing the encoded output matrix and the decoded output matrix, a similarity matrix is ​​obtained between the two. The similarity matrix is ​​then activated using the Sigmoid activation function to obtain the frame-level auxiliary speech detection results, that is, to obtain the auxiliary speech corresponding to each speech frame.

[0055] Step 607: Based on the frame-level auxiliary speech detection results, identify all auxiliary speech segments; Specifically, if all the auxiliary speech segments from the first to the tenth speech frame have speech results, but there is no speech in the eleventh speech frame, then there is an auxiliary speech segment from the first to the tenth speech frame. And so on, all auxiliary speech segments are identified. Here, a frame-level auxiliary speech detection model is used to detect all auxiliary speech segments more accurately and avoid classifying different auxiliary speech segments together.

[0056] Step 608: Based on the pronunciation time information of the target speech sequence, add timestamps to all auxiliary speech segments to obtain auxiliary speech segments with timestamps.

[0057] Similarly, based on the frame value of the speech frame and the pronunciation time information of the target speech sequence, the pronunciation start time and pronunciation end time are added as corresponding timestamps to all auxiliary speech segments to facilitate the identification of the sequence insertion relationship between all auxiliary speech segments and all word-level speech texts.

[0058] Continue to refer to Figure 7 , Figure 7 yes Figure 3 A flowchart of a specific embodiment of step 305 shown includes: Step 701: Perform acoustic feature extraction on all auxiliary speech segments to obtain the acoustic feature extraction results; Specifically, due to the diversity of auxiliary speech segments, acoustic features are extracted for all auxiliary speech segments to facilitate subsequent classification of auxiliary speech segments.

[0059] Step 702: Perform cluster analysis on the acoustic feature extraction results to identify the pronunciation type corresponding to each of the auxiliary speech segments; Specifically, cluster analysis is performed on the acoustic feature extraction results to identify the pronunciation type corresponding to each auxiliary speech segment. For example, the pronunciation type of auxiliary speech segment A is laughter, the pronunciation type of auxiliary speech segment B is coughing, the pronunciation type of auxiliary speech segment C is sighing, and the pronunciation type of auxiliary speech segment D is surprise.

[0060] Step 703: Calculate the pronunciation duration of each auxiliary speech segment based on the timestamp marked on each auxiliary speech segment, and determine the pronunciation intensity of each auxiliary speech segment based on the pronunciation duration and the acoustic feature extraction result, wherein the timestamp includes a start timestamp and an end timestamp; Specifically, for example, the longer the pronunciation duration, the stronger the corresponding pronunciation intensity. That is, when the pronunciation type of the current auxiliary speech segment is laughter, the longer the laughter duration, the stronger the laughter intensity. The acoustic feature extraction results can also incorporate sound decibels, intonation, etc., to jointly determine the pronunciation intensity of each auxiliary speech segment.

[0061] Step 704: Determine the insertion probability of each auxiliary speech segment after different text content based on the timestamps marked on each auxiliary speech segment and the timestamps marked on the word-level speech text. Specifically, this process requires aligning the timestamps of each auxiliary speech segment and the word-level speech text according to their temporal sequence relationships. For example, assuming the pronunciation of word-level speech text A ends at 34 seconds, the pronunciation of auxiliary speech segment A begins at 35 seconds and ends at 40 seconds, and the pronunciation of word-level speech text B begins at 41 seconds, this indicates that auxiliary speech segment A exists between word-level speech text A and word-level speech text B in the target speech sequence. Through alignment, we can identify the text content frequently following different types of auxiliary speech segments. Then, statistical analysis methods are used to determine the insertion probability of each auxiliary speech segment after different text content.

[0062] Step 705: Combine pronunciation type, pronunciation intensity, and insertion probability to learn the structured embedding signal representations corresponding to all auxiliary speech segments; Specifically, the system utilizes all auxiliary speech segments and all word-level speech text contained in the multi-source real audio for learning and training. It combines pronunciation type, pronunciation intensity, and insertion probability as a behavior prediction combination to predict the combination representation of different pronunciation types, pronunciation intensities, and insertion probabilities corresponding to auxiliary speech segments appearing after different text content. The predicted combination representation is then used as the structured embedded signal representation corresponding to the corresponding auxiliary speech segment.

[0063] For example, for the text "I won the lottery", the auxiliary voice insertion planner may plan to insert the structured embedded signal representation, i.e., {pronunciation type: laugh, pronunciation intensity: 0.9, pronunciation probability: 0.95}, after the character "le" in the text. And this embedded signal representation is essentially the content learned through learning and training based on all the auxiliary voice segments and all the word-level voice texts included in the multi-source real audio.

[0064] Step 706, use the auxiliary voice insertion planner with the learned embedded signal representation as the auxiliary voice insertion planner after learning and training.

[0065] Specifically, using the auxiliary voice insertion planner with the learned embedded signal representation as the auxiliary voice insertion planner after learning and training is convenient for directly screening out the corresponding embedded signal representation according to the input word-level voice text during actual voice generation later, so as to implement applying the corresponding auxiliary voice segment to voice generation and generating a more realistic voice.

[0066] Continue to refer to Figure 8 , Figure 8 is Figure 7 a flowchart of a specific embodiment of step 705 shown, including: Step 801, obtain the structured embedded signal representation corresponding to the current auxiliary voice segment, and input the embedded signal representation into the voice decoder of a preset TTS model to obtain the decoded pronunciation audio corresponding to the current auxiliary voice segment; Specifically, the voice decoder of the preset TTS model includes the voice decoder of a non-autoregressive TTS model based on F5-TTS.

[0067] It should be understood that although the auxiliary voice insertion planner can comprehensively represent the pronunciation type, pronunciation intensity, and insertion probability for different auxiliary voice segments, that is, the embedded signal representation, for voice decoding when the corresponding embedded signal representation is recognized, whether the corresponding auxiliary voice segment can be obtained still requires voice decoding training to train the corresponding voice decoder so that the corresponding auxiliary voice segment can be inserted at the target insertion position when the corresponding embedded signal representation is recognized.

[0068] Specifically, inputting the embedded signal representation into the voice decoder of the preset TTS model can input the pronunciation type, pronunciation intensity, and insertion probability as different quantities into the voice decoder respectively. For example, input the pronunciation type as an embedded vector, the pronunciation intensity as a scalar, and the insertion probability as an attention mask.

[0069] Step 802: Obtain the actual pronunciation audio of the current auxiliary speech segment obtained when performing auxiliary speech detection on the target speech sequence; Step 803: Using the actual pronunciation audio as the discriminant speech in the discriminator, and the decoded pronunciation audio as the current speech to be discriminated, a generative adversarial training method is adopted to gradually optimize the pronunciation result of the auxiliary speech segment corresponding to the current embedded signal representation. Specifically, the gradual optimization of the pronunciation result of the auxiliary speech segment corresponding to the current embedded signal representation includes: optimizing the decoding parameters of the decoder of the TTS model so that the obtained decoded pronunciation audio becomes closer and closer to the actual pronunciation audio under each combination of pronunciation type, pronunciation intensity and insertion probability. Step 804: Once the decoded pronunciation audio corresponding to the currently optimized auxiliary speech segment is determined by the discriminator, the learning of the structured embedded signal representation corresponding to the current auxiliary speech segment is complete.

[0070] Specifically, this involves selecting the actual pronunciation audio of the current auxiliary speech segment from all the auxiliary speech segments obtained in step 607. Using the actual pronunciation audio as the discrimination audio, and employing generative adversarial training, the current auxiliary speech segment is progressively generated and optimized to ensure that the final speech decoder can generate more realistic auxiliary speech. Essentially, the auxiliary speech insertion planner incorporates the speech decoder based on the F5-TTS non-autoregressive TTS model. This F5-TTS non-autoregressive TTS model speech decoder focuses on the generation and insertion of auxiliary speech based on its pronunciation type, pronunciation intensity, and insertion probability, ultimately inserting the target auxiliary speech between corresponding word-level speech text. This ensures that more realistic synthesized speech is generated.

[0071] In this embodiment, the process involves: acquiring word-level speech text to be generated; inputting the word-level speech text into a trained auxiliary speech insertion planner, where auxiliary speech includes speech that cannot be directly converted into natural language; acquiring the speech generation encoding sequence output by the auxiliary speech insertion planner; and using a preset speech decoder to decode and output the speech generation encoding sequence to obtain the target generated speech. By introducing the auxiliary speech insertion planner, speech that cannot be directly converted into natural language is inserted into the word-level speech text to be generated, ensuring the realism of the final generated speech and making it closer to the pronunciation of a real person. Applying this method to consultation and response scenarios in fintech or healthcare businesses can assist intelligent AI customer service in replacing human responses, generating pronunciation effects that are closer to those of a real person.

[0072] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.

[0073] Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.

[0074] In this embodiment, the process involves: acquiring word-level speech text to be generated; inputting the word-level speech text into a trained auxiliary speech insertion planner, where auxiliary speech includes speech that cannot be directly converted into natural language; acquiring the speech generation encoding sequence output by the auxiliary speech insertion planner; and using a preset speech decoder to decode and output the speech generation encoding sequence to obtain the target generated speech. By introducing the auxiliary speech insertion planner, speech that cannot be directly converted into natural language is inserted into the word-level speech text to be generated, ensuring the realism of the final generated speech and making it closer to the pronunciation of a real person. Applying this method to consultation and response scenarios in fintech or healthcare businesses can assist intelligent AI customer service in replacing human responses, generating pronunciation effects that are closer to those of a real person.

[0075] Further reference Figure 9 As a response to the above Figure 2 To implement the method shown, this application provides an embodiment of a speech generation device, which is similar to... Figure 2 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.

[0076] like Figure 9 As shown, the speech generation device 900 described in this embodiment includes: a word-level speech-text acquisition module 901, a word-level speech-text input module 902, a speech generation encoding sequence acquisition module 903, and a target generated speech acquisition module 904. Wherein: The word-level speech-text acquisition module 901 is used to acquire the word-level speech-text to be generated. The word-level speech-text input module 902 is used to input the word-level speech-text into the auxiliary speech insertion planner that has been trained, wherein the auxiliary speech includes speech that cannot be directly converted into natural language; The speech generation coding sequence acquisition module 903 is used to acquire the speech generation coding sequence output by the auxiliary speech insertion planner; The target generated speech acquisition module 904 is used to decode and output the speech generation encoding sequence using a preset speech decoder to obtain the target generated speech.

[0077] This application involves: acquiring word-level speech text to be generated; inputting the word-level speech text into a trained auxiliary speech insertion planner, wherein the auxiliary speech includes speech that cannot be directly converted into natural language; acquiring the speech generation encoding sequence output by the auxiliary speech insertion planner; and using a preset speech decoder to decode and output the speech generation encoding sequence to obtain the target generated speech. By introducing an auxiliary speech insertion planner, speech that cannot be directly converted into natural language is inserted into the word-level speech text to be generated, ensuring the realism of the final generated speech and making it closer to the pronunciation of a real person. Applying this method to consultation and response scenarios in fintech or healthcare businesses can assist intelligent AI customer service in replacing human responses, generating pronunciation effects that are closer to those of a real person.

[0078] In this embodiment, the speech generation device 900 further includes a training audio acquisition module, an audio preprocessing module, a speech-to-text module, an auxiliary speech detection module, and an auxiliary speech insertion planner training module. Wherein: Train the audio acquisition module to acquire real audio from multiple sources; An audio preprocessing module is used to preprocess the real audio to obtain a target speech sequence; The speech-to-text module is used to perform speech-to-text processing on the target speech sequence to obtain word-level speech text with timestamp annotation; An auxiliary speech detection module is used to perform auxiliary speech detection on the target speech sequence to obtain auxiliary speech segments labeled with timestamps; The auxiliary speech insertion planner training module is used to input the timestamped auxiliary speech segments and the timestamped word-level speech texts into the auxiliary speech insertion planner to be trained, and to perform auxiliary speech insertion learning training to obtain the auxiliary speech insertion planner after learning and training.

[0079] In this embodiment, the audio preprocessing module includes a speech activity detection unit, a speech source separation unit, and a speech log separation unit. Wherein: A voice activity detection unit is used to perform voice activity detection on the real audio using a preset voice activity detection tool; The speech source separation unit is used to perform speech log and speech source separation processing on the real audio based on the speech activity detection results, and obtain a clean single-speaker speech sequence through the speech source separation processing results; The speech log separation unit is used to annotate the pronunciation time information of a clean single-speaker speech sequence based on the speech log separation processing result, and obtain the speech sequence annotated with the pronunciation time information as the target speech sequence.

[0080] In this embodiment, the auxiliary speech detection module includes an auxiliary speech detection input unit, a multi-layer audio feature extraction unit, an audio feature fusion processing unit, an encoding output unit, a decoding output unit, a contrast activation unit, an auxiliary speech determination unit, and an auxiliary speech annotation unit. Wherein: An auxiliary speech detection input unit is used to input the target speech sequence into a preset frame-level auxiliary speech detection model; A multi-layer audio feature extraction unit is used to extract multi-layer audio features from the target speech sequence using the frame-level auxiliary speech detection model. The audio feature fusion processing unit is used to fuse the extracted multi-layer audio features using a learnable weighted summation method to obtain audio fusion features; The encoding output unit is used to encode the audio fusion features using a preset Transformer encoder to obtain the encoded output result; The decoding output unit is used to decode the encoded output result by combining a set of preset decodeable vectors to obtain the decoded output result; The contrast activation unit is used to calculate the similarity matrix between the decoded output and the encoded output, and then activates it through a preset activation function to obtain frame-level auxiliary speech detection results. The auxiliary speech determination unit is used to determine all auxiliary speech segments based on the frame-level auxiliary speech detection results; An auxiliary speech annotation unit is used to add timestamps to all auxiliary speech segments based on the pronunciation time information of the target speech sequence, so as to obtain auxiliary speech segments annotated with timestamps.

[0081] In this embodiment, the semantic fingerprint generation module 901 is further used to perform multimodal feature extraction on all compliant cases to generate a semantic fingerprint corresponding to each compliant case; and to perform multimodal feature extraction on all non-compliant cases to generate a semantic fingerprint corresponding to each non-compliant case.

[0082] In this embodiment, the auxiliary speech insertion planner training module includes an acoustic feature extraction unit, a pronunciation type recognition unit, a pronunciation intensity determination unit, an insertion probability determination unit, an embedding signal representation learning unit, and a learning and training completion unit. Wherein: The acoustic feature extraction unit is used to extract acoustic features from all auxiliary speech segments and obtain the acoustic feature extraction results. The pronunciation type recognition unit is used to perform cluster analysis on the acoustic feature extraction results to identify the pronunciation type corresponding to each of the auxiliary speech segments; The pronunciation intensity determination unit is used to calculate the pronunciation duration of each auxiliary speech segment based on the timestamp marked on each auxiliary speech segment, and to determine the pronunciation intensity of each auxiliary speech segment based on the pronunciation duration and the acoustic feature extraction result, wherein the timestamp includes a start timestamp and an end timestamp; The insertion probability determination unit is used to determine the insertion probability of each auxiliary speech segment after different text contents based on the timestamp marked on each auxiliary speech segment and the timestamp marked on the word-level speech text. The embedded signal representation learning unit is used to learn the structured embedded signal representations corresponding to all auxiliary speech segments by combining pronunciation type, pronunciation intensity and insertion probability; A learning and training completion unit is used to represent the learning completion auxiliary speech insertion planner as the learning and training completion auxiliary speech insertion planner by embedding signals.

[0083] In this embodiment, the auxiliary speech insertion planner training module further includes a decoded pronunciation audio acquisition unit, an actual pronunciation audio acquisition unit, a generated adversarial training unit, and an adversarial training termination unit. Wherein: The decoded pronunciation audio acquisition unit is used to acquire the structured embedded signal representation corresponding to the current auxiliary speech segment, and input the embedded signal representation into the speech decoder of the preset TTS model to obtain the decoded pronunciation audio corresponding to the current auxiliary speech segment; The actual pronunciation audio acquisition unit is used to acquire the actual pronunciation audio of the current auxiliary speech segment obtained when performing auxiliary speech detection on the target speech sequence; A generative adversarial training unit is used to take the actual pronunciation audio as the discriminant speech in the discriminator and the decoded pronunciation audio as the current speech to be discriminated. The generative adversarial training method is used to gradually optimize the pronunciation result of the auxiliary speech segment corresponding to the current embedded signal representation. The adversarial training termination unit is used to ensure that the structured embedded signal representation learning of the current auxiliary speech segment is completed when the decoded pronunciation audio corresponding to the currently optimized auxiliary speech segment is discriminated by the discriminator.

[0084] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by instructing related hardware through computer-readable instructions. These computer-readable instructions can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the methods described above. The aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, optical disk, or read-only memory (ROM), or random access memory (RAM).

[0085] It should be understood that although the steps in the flowcharts of the accompanying figures are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the accompanying figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.

[0086] To address the aforementioned technical problems, embodiments of this application also provide a computer device. Please refer to [link / reference needed] for details. Figure 10 , Figure 10 This is a basic structural block diagram of the computer device in this embodiment.

[0087] The computer device 10 includes a memory 10a, a processor 10b, and a network interface 10c, which are interconnected via a system bus. It should be noted that... Figure 10 Only a computer device 10 with component memory 10a, processor 10b, and network interface 10c is shown. However, it should be understood that it is not required to implement all the components shown, and more or fewer components may be implemented instead. Those skilled in the art will understand that the computer device described herein is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions. Its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.

[0088] The computer device can be a desktop computer, laptop, handheld computer, or cloud server, etc. The computer device can interact with the user via a keyboard, mouse, remote control, touchpad, or voice control.

[0089] The memory 10a includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 10a may be an internal storage unit of the computer device 10, such as the hard disk or memory of the computer device 10. In other embodiments, the memory 10a may also be an external storage device of the computer device 10, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc. Of course, the memory 10a may include both internal storage units and external storage devices of the computer device 10. In this embodiment, the memory 10a is typically used to store the operating system and various application software installed on the computer device 10, such as computer-readable instructions for a speech generation method. In addition, the memory 10a can also be used to temporarily store various types of data that have been output or will be output.

[0090] In some embodiments, the processor 10b may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip. The processor 10b is typically used to control the overall operation of the computer device 10. In this embodiment, the processor 10b is used to execute computer-readable instructions stored in the memory 10a or to process data, for example, to execute computer-readable instructions for the aforementioned speech generation method.

[0091] The network interface 10c may include a wireless network interface or a wired network interface, which is typically used to establish communication connections between the computer device 10 and other electronic devices.

[0092] The computer device proposed in this embodiment belongs to the field of artificial intelligence technology and is applied to the scenario of speech generation for AI question-answering agents. This application obtains word-level speech text to be generated; inputs the word-level speech text into an auxiliary speech insertion planner that has completed learning and training, wherein the auxiliary speech includes speech that cannot be directly converted into natural language; obtains the speech generation encoding sequence output by the auxiliary speech insertion planner; and uses a preset speech decoder to decode and output the speech generation encoding sequence to obtain the target generated speech. By introducing an auxiliary speech insertion planner, speech that cannot be directly converted into natural language is inserted into the word-level speech text to be generated, ensuring the realism of the final generated speech and making it closer to the pronunciation of a real person. Applying this method to consultation and response scenarios in fintech or healthcare businesses can assist intelligent AI customer service in replacing human responses, generating pronunciation effects that are closer to those of a real person.

[0093] This application also provides another embodiment, namely, a computer-readable storage medium storing computer-readable instructions that can be executed by a processor to cause the processor to perform the steps of the speech generation method described above.

[0094] The computer-readable storage medium proposed in this embodiment belongs to the field of artificial intelligence technology and is applied to the scenario of speech generation for AI question-answering agents. This application obtains word-level speech text to be generated; inputs the word-level speech text into a trained auxiliary speech insertion planner, wherein the auxiliary speech includes speech that cannot be directly converted into natural language; obtains the speech generation encoding sequence output by the auxiliary speech insertion planner; and decodes the speech generation encoding sequence using a preset speech decoder to obtain the target generated speech. By introducing an auxiliary speech insertion planner, speech that cannot be directly converted into natural language is inserted into the word-level speech text to be generated, ensuring the realism of the final generated speech and making it closer to the pronunciation of a real person. Applying this method to consultation and response scenarios in fintech or healthcare businesses can assist intelligent AI customer service in replacing human responses, generating pronunciation effects that are closer to those of a real person.

[0095] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0096] Obviously, the embodiments described above are only some embodiments of this application, not all embodiments. The accompanying drawings show preferred embodiments of this application, but do not limit the patent scope of this application. This application can be implemented in many different forms; rather, the purpose of providing these embodiments is to make the disclosure of this application more thorough and comprehensive. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing specific embodiments, or make equivalent substitutions for some of the technical features. Any equivalent structures made using the content of this application's specification and drawings, directly or indirectly applied to other related technical fields, are similarly within the scope of patent protection of this application. Software tools or components not belonging to this company that appear in the embodiments of this application are merely illustrative examples and do not represent actual use.

Claims

1. A speech generation method, characterized in that, Includes the following steps: Obtain the word-level speech text to be generated; The word-level speech text is input into the auxiliary speech insertion planner that has been trained, wherein the auxiliary speech includes speech that cannot be directly converted into natural language; Obtain the speech generation coding sequence output by the auxiliary speech insertion planner; The speech generation encoding sequence is decoded and output using a preset speech decoder to obtain the target generated speech.

2. The speech generation method according to claim 1, characterized in that, Before performing the step of inputting the word-level speech text into the learned and trained auxiliary speech insertion planner, the method further includes: Obtain authentic audio from multiple sources; The real audio is preprocessed to obtain the target speech sequence; The target speech sequence is processed into speech-to-text to obtain word-level speech-to-text with timestamp annotations; Auxiliary speech detection is performed on the target speech sequence to obtain auxiliary speech segments labeled with timestamps; The timestamped auxiliary speech segments and the timestamped word-level speech texts are input together into the auxiliary speech insertion planner to be trained, and auxiliary speech insertion learning training is performed to obtain the auxiliary speech insertion planner after training.

3. The speech generation method according to claim 2, characterized in that, The step of preprocessing the real audio to obtain the target speech sequence includes: A preset voice activity detection tool is used to detect voice activity in the real audio. Based on the voice activity detection results, the real audio is processed by voice log and voice source separation. Through the voice source separation processing results, a clean single-speaker voice sequence is obtained. Based on the speech log separation processing results, the pronunciation time information of the clean single-speaker speech sequence is annotated to obtain the speech sequence annotated with the pronunciation time information as the target speech sequence.

4. The speech generation method according to claim 2, characterized in that, The step of performing speech-to-text conversion on the target speech sequence to obtain word-level speech text annotated with timestamps specifically includes: Acoustic features are extracted from the target speech sequence using a pre-defined ASR model; The extracted acoustic features are transcribed into text using the text decoder in the ASR model to obtain the transcribed text. Based on the pronunciation time information of the target speech sequence, word-level timestamps are added to the transcribed text to obtain word-level speech text annotated with timestamps.

5. The speech generation method according to claim 2, characterized in that, The step of performing auxiliary speech detection on the target speech sequence to obtain auxiliary speech segments labeled with timestamps specifically includes: The target speech sequence is input into a preset frame-level auxiliary speech detection model; The frame-level auxiliary speech detection model is used to extract multi-layer audio features from the target speech sequence; A learnable weighted summation method is used to fuse the extracted multi-layer audio features to obtain audio fusion features; The audio fusion features are encoded and output using a preset Transformer encoder to obtain the encoded output result; The encoded output result is decoded by combining a set of preset decodeable vectors to obtain the decoded output result; The similarity matrix between the decoded output and the encoded output is calculated and activated by a preset activation function to obtain frame-level auxiliary speech detection results. Based on the frame-level auxiliary speech detection results, all auxiliary speech segments are identified; Based on the pronunciation time information of the target speech sequence, timestamps are added to all auxiliary speech segments to obtain auxiliary speech segments with timestamps.

6. The speech generation method according to claim 2, characterized in that, The step of inputting the timestamped auxiliary speech segments and the timestamped word-level speech text into the auxiliary speech insertion planner to be trained, and performing auxiliary speech insertion learning training to obtain the trained auxiliary speech insertion planner includes: Acoustic features were extracted from all auxiliary speech segments to obtain the acoustic feature extraction results; Cluster analysis is performed on the acoustic feature extraction results to identify the pronunciation type corresponding to each of the auxiliary speech segments; Based on the timestamp marked on each auxiliary speech segment, the pronunciation duration of each auxiliary speech segment is calculated, and based on the pronunciation duration and the acoustic feature extraction results, the pronunciation intensity of each auxiliary speech segment is determined. Based on the timestamps marked on each auxiliary speech segment and the timestamps marked on the word-level speech text, determine the insertion probability of each auxiliary speech segment after different text content; By combining pronunciation type, pronunciation intensity, and insertion probability, a structured embedding signal representation corresponding to all auxiliary speech segments is learned; The auxiliary speech insertion planner that has completed learning is represented by the embedded signal as the auxiliary speech insertion planner that has completed learning and training.

7. The speech generation method according to claim 6, characterized in that, The step of learning structured embedded signal representations for all auxiliary speech segments by combining pronunciation type, pronunciation intensity, and insertion probability includes: Obtain the structured embedded signal representation corresponding to the current auxiliary speech segment, and input the embedded signal representation into the speech decoder of the preset TTS model to obtain the decoded pronunciation audio corresponding to the current auxiliary speech segment; When performing auxiliary speech detection on the target speech sequence, obtain the actual pronunciation audio of the current auxiliary speech segment; The actual pronunciation audio is used as the discriminant speech in the discriminator, and the decoded pronunciation audio is used as the current speech to be discriminated. A generative adversarial training method is adopted to gradually optimize the pronunciation results of the auxiliary speech segments corresponding to the current embedded signal representation. The structured embedded signal representation learning for the current auxiliary speech segment is completed once the decoded pronunciation audio corresponding to the currently optimized auxiliary speech segment is discriminated by the discriminator.

8. A speech generation device, characterized in that, include: The word-level speech-text acquisition module is used to acquire the word-level speech-text to be generated. A word-level speech-text input module is used to input the word-level speech-text into an auxiliary speech insertion planner that has been trained, wherein the auxiliary speech includes speech that cannot be directly converted into natural language; The speech generation coding sequence acquisition module is used to acquire the speech generation coding sequence output by the auxiliary speech insertion planner; The target generated speech acquisition module is used to decode and output the speech generation encoding sequence using a preset speech decoder to obtain the target generated speech.

9. A computer device, characterized in that, The method includes a memory and a processor, wherein the memory stores computer-readable instructions, and the processor executes the computer-readable instructions to implement the steps of the speech generation method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the speech generation method as described in any one of claims 1 to 7.