Generation method and generation device for anthropomorphic audio and electronic equipment

Through the combination of sharding processing and TTS model, semantic and emotional feature vectors are identified and audio files are synthesized, and the problem of insufficient emotional expression of digital human audio is solved, and the degree of anthropomorphism and emotional appeal are improved.

CN119964546APending Publication Date: 2025-05-09GUOKE INTELLIGENT MFG (BINZHOU) TECH DEV CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510103024.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

The audio generated by digital life has shortcomings in emotional expression, the output is relatively ‘machine’, lacks fluency and emotional appeal, especially in scenarios involving rich emotional expression and interesting interactions.

Method used

By obtaining the streaming reply information output from the large language model, the semantic vector extraction model and emotion vector generation model of the target TTS model are used after sharding processing, and the semantic feature vector of the text slice and the emotional feature vector of the target character are identified, and the audio file corresponding to each text slice is synthesized.

Benefits of technology

It improves the anthropomorphic level of digital human output audio, making it more natural and smooth in emotional expression and tone adjustment, and enhances emotional appeal.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119964546A_ABST
    Figure CN119964546A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of voice broadcast, and discloses a generation method and a generation device for anthropomorphic audio, and electronic equipment. The generation method comprises the steps of obtaining streaming reply information output by a large language model, and performing fragmentation processing on the streaming reply information to determine a plurality of text slices; identifying a semantic feature vector of each text slice by adopting a semantic vector extraction model of the target TTS model; adopting an emotion vector generation model of the target TTS model, processing the audio file and the emotion parameters of the target person, and determining an emotion feature vector of the target person; and synthesizing an audio file corresponding to each text slice according to the semantic feature vector of each text slice and the emotional feature vector of the target person. According to the invention, the personification degree of the digital human output audio can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of voice broadcasting technology, for example, to a method and device for generating anthropomorphic audio, and an electronic device. Background Art

[0002] At present, Digital Human Technology, as a virtual image technology, integrates multiple technologies such as computer graphics, artificial intelligence, natural language processing, deep learning, etc. to create virtual people with human image, voice, behavior and even emotional expression.

[0003] In related technologies, the voice broadcast function of digital humans is widely used in entertainment, education, medical care, customer service, social networking and other fields to answer users' questions in digital humans and provide anthropomorphic voice broadcast services.

[0004] In the process of implementing the embodiments of the present disclosure, it is found that there are at least the following problems in the related art:

[0005] The audio generated by digital humans in related technologies is insufficient in terms of emotional expression. In terms of multi-emotion and natural tone adjustment, the generated audio output is relatively "mechanical". This results in the audio effects output by digital humans lacking fluency and emotional appeal compared to the voices of real people in scenes involving rich emotional expressions and interesting interactions. Therefore, how to improve the degree of anthropomorphism of the audio output by digital humans has become a technical problem that needs to be solved urgently. Summary of the invention

[0006] In order to provide a basic understanding of some aspects of the disclosed embodiments, a brief summary is given below. The summary is not an extensive review, nor is it intended to identify key / critical components or delineate the scope of protection of these embodiments, but rather serves as a prelude to the detailed description that follows.

[0007] The embodiments of the present disclosure provide a method and device for generating anthropomorphic audio, and an electronic device, which can improve the degree of anthropomorphism of audio output by a digital human.

[0008] In some embodiments, a method for generating anthropomorphic audio includes: obtaining streaming response information output by a large language model, and processing the streaming response information in segments to determine multiple text slices; using a semantic vector extraction model of a target TTS model to identify a semantic feature vector of each text slice; using an emotion vector generation model of a target TTS model to process an audio file and emotion parameters of a target character to determine an emotion feature vector of the target character; and synthesizing an audio file corresponding to each text slice based on the semantic feature vector of each text slice and the emotion feature vector of the target character.

[0009] Optionally, the streaming response information output by the large language model is sliced ​​to determine multiple text slices, including: slicing the streaming response information according to punctuation marks to determine multiple text slices; confirming the text length of each text slice; and using a word segmenter to re-slice the text slices whose text length is greater than or equal to a set threshold.

[0010] Optionally, a semantic vector extraction model of a target TTS model is used to identify the semantic feature vector of each text slice, including: saving multiple text slices to a text queue; reading text slices from the text queue in sequence, and performing colloquial processing on the current text slice according to the professional field to which the streaming reply information belongs, to obtain the colloquial text corresponding to the current text slice; extracting the semantic feature vector of the colloquial text as the semantic feature vector of the current text slice.

[0011] Optionally, an emotion vector generation model of a target TTS model is used to process the audio file and emotion parameters of a target person to determine the emotion feature vector of the target person, including: determining the audio feature vector of the target person based on the audio file of the target person; encoding the emotion parameters of the target person into a vector form to obtain an emotion parameter vector; and concatenating and fusing the audio feature vector and the emotion parameter vector to determine the emotion feature vector of the target person.

[0012] Optionally, an audio file corresponding to each text slice is synthesized based on the semantic feature vector of each text slice and the emotional feature vector of the target person, including: concatenating the semantic feature vector of each text slice with the emotional feature vector of the target person to determine the Mel spectrum of each text slice; converting the Mel spectrum of each text slice into an audio waveform to determine the audio file corresponding to each text slice.

[0013] Optionally, before using the semantic vector extraction model of the target TTS model to identify the semantic feature vector of each text slice, the generation method also includes: inputting the current text slice into a model pool, and determining a target TTS model in the model pool for processing the current text slice.

[0014] Optionally, after synthesizing the audio file corresponding to each text slice, the generation method further includes: saving the audio file corresponding to each text slice to an audio queue, so that the audio player sequentially reads the audio files from the audio queue for playing.

[0015] In some embodiments, a device for generating anthropomorphic audio includes: a slicing module, configured to obtain streaming response information output by a large language model, and to process the streaming response information in slices to determine multiple text slices; a semantic recognition module, configured to adopt a semantic vector extraction model of a target TTS model to identify a semantic feature vector of each text slice; an emotion recognition module, configured to adopt an emotion vector generation model of a target TTS model to process an audio file and emotion parameters of a target character, and determine the emotion feature vector of the target character; and a synthesis module, configured to synthesize an audio file corresponding to each text slice based on the semantic feature vector of each text slice and the emotion feature vector of the target character.

[0016] In some embodiments, a device for generating anthropomorphic audio includes a processor and a memory storing program instructions, and the processor is configured to execute the method for generating anthropomorphic audio as described above.

[0017] In some implementations, an electronic device includes: a device body; and a device for generating anthropomorphic audio as described above, which is disposed in the device body.

[0018] The prediction method, prediction device, and electronic device for time series provided by the embodiments of the present disclosure can achieve the following technical effects:

[0019] In the disclosed embodiment, first, the streaming reply information output by the virtual person based on the large language model is processed in slices to obtain multiple text slices, and the semantic vector extraction model of the target TTS model is used to identify the semantic feature vector of each text slice. Then, the emotional vector generation model of the target TTS model is used to process the audio file and emotional parameters of the target person (for example, the section chief of electrolytic aluminum) to determine the emotional feature vector of the target person. Finally, the audio file corresponding to each text slice is synthesized based on the semantic feature vector of each text slice and the emotional feature vector of the target person. It can be seen that in the disclosed embodiment, the emotional feature vector including the audio features (such as intonation, timbre, etc.) of the target person and the emotional parameters (such as happiness, sadness, anger, etc.) of the target person is introduced during audio synthesis, and each time only the audio of a text slice in the synthesized streaming reply information is synthesized, which helps to achieve more detailed emotions and tone. Therefore, the disclosed embodiment can improve the degree of anthropomorphism of the digital person output audio.

[0020] The above general description and the following description are exemplary and explanatory only and are not intended to limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] One or more embodiments are exemplarily described by corresponding drawings, which do not limit the embodiments. Elements with the same reference numerals in the drawings are shown as similar elements, and the drawings do not constitute a scale limitation, and wherein:

[0022] Figure 1 is a schematic diagram of an electronic device provided by an embodiment of the present disclosure;

[0023] Figure 2 is a schematic diagram of a method for generating anthropomorphic audio provided by an embodiment of the present disclosure;

[0024] Figure 3 is a schematic diagram of a vocoder architecture provided by an embodiment of the present disclosure;

[0025] Figure 4 is a schematic diagram of a method for generating anthropomorphic audio provided by an embodiment of the present disclosure;

[0026] Figure 5 is a schematic diagram of a device for generating anthropomorphic audio provided by an embodiment of the present disclosure;

[0027] Figure 6 It is a schematic diagram of another device for generating anthropomorphic audio provided by an embodiment of the present disclosure. DETAILED DESCRIPTION

[0028] In order to be able to understand the features and technical contents of the embodiments of the present disclosure in more detail, the implementation of the embodiments of the present disclosure is described in detail below in conjunction with the accompanying drawings. The attached drawings are for reference only and are not used to limit the embodiments of the present disclosure. In the following technical description, for the convenience of explanation, a full understanding of the disclosed embodiments is provided through multiple details. However, one or more embodiments can still be implemented without these details. In other cases, to simplify the drawings, well-known structures and devices can be simplified for display.

[0029] The terms "first", "second", etc. in the specification and claims of the embodiments of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the terms used in this way can be interchanged where appropriate, so that the embodiments of the embodiments of the present disclosure described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions.

[0030] Unless otherwise stated, the term "plurality" of a feature means two or more.

[0031] In the embodiment of the present disclosure, the character " / " feature indicates that the preceding and following objects are in an "or" relationship. For example, the A / B feature indicates: A or B.

[0032] The term "and / or" is a description of the association relationship between objects, and the characteristic indicates that there can be three relationships. For example, A and / or B, the characteristic indicates: A or B, or, A and B.

[0033] The term "correspondence" may refer to an association relationship or a binding relationship. The correspondence between A and B means that there is an association relationship or a binding relationship between A and B.

[0034] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present disclosure may be combined with each other.

[0035] like Figure 1 As shown, the electronic device 100 provided by the embodiment of the present disclosure includes a device body 110 and a device 500 (600) for generating anthropomorphic audio. The device 500 (600) for generating anthropomorphic audio is disposed in the device body 110.

[0036] Specifically, a voice broadcasting system based on digital human technology is deployed in the electronic device, and the device 500 (600) for generating anthropomorphic audio is used to process the streaming reply information generated by the large voice model in the voice playback system, generate a corresponding audio file and play it.

[0037] Optionally, the device 600 for generating anthropomorphic audio includes a processor, which can obtain the streaming reply information output by the large language model, and perform segmentation processing on the streaming reply information to determine multiple text slices. The semantic vector extraction model of the target TTS model can be used to identify the semantic feature vector of each text slice. The emotional vector generation model of the target TTS model can be used to process the audio file and emotional parameters of the target person to determine the emotional feature vector of the target person. The audio file corresponding to each text slice can be synthesized based on the semantic feature vector of each text slice and the emotional feature vector of the target person.

[0038] In combination with the above electronic device, the present disclosure provides a method for generating anthropomorphic audio, such as Figure 2 As shown, the generation method includes:

[0039] S201, the processor obtains streaming response information output by the large language model, and processes the streaming response information in segments to determine multiple text segments.

[0040] Specifically, the large language model refers to a large model used for dialogue interaction with users in a voice broadcast system based on digital human technology.

[0041] Specifically, streaming reply information means that when a large language model generates a reply message, it does not give a complete answer at once, but returns the answer step by step or in blocks. This method can significantly reduce the user's waiting time, simulate the rhythm of natural conversation, and make the entire interaction process more in line with the experience of daily communication.

[0042] It is understandable that the text length of the streaming reply information output by the large language model may be relatively long, and directly processing the entire streaming reply information may consume a lot of computing resources and time. However, by processing the streaming reply information in slices, dividing the streaming reply information into multiple text slices with smaller text lengths, the burden of a single processor processing can be significantly reduced, and processing efficiency can be improved.

[0043] Optionally, the streaming response information output by the large language model is sliced ​​to determine multiple text slices, including: slicing the streaming response information according to punctuation marks to determine multiple text slices; confirming the text length of each text slice; and using a word segmenter to re-slice the text slices whose text length is greater than or equal to a set threshold.

[0044] Specifically, punctuation marks can express the pause and tone of a text and separate components in the text. Therefore, after obtaining the streaming reply information, the streaming reply information can be sliced ​​according to the punctuation marks to determine multiple text slices.

[0045] Specifically, if the text length of the determined text slice is still relatively long, the subsequent audio conversion of the text slice with a relatively long text length will still consume a lot of computing resources and take a long time. Therefore, for the multiple text slices initially obtained, the word segmenter will be used to re-slice the text slices whose text length is greater than or equal to the set threshold.

[0046] Specifically, the basis for the word segmenter to perform slicing processing again on the text switch is the connecting words in the text slices.

[0047] For example, the streaming reply information of the large language model is: "The basic pressure of electrolytic aluminum usually refers to the cell voltage of the electrolytic cell. The cell voltage refers to the minimum voltage to maintain the normal operation of the electrolytic cell, which is composed of the following parts: Anode voltage drop: related to the anode current density, the resistivity of the anode carbon block, the assembly quality of the anode carbon block steel claw-guide rod, and the contact between the guide rod and the busbar."

[0048] After slicing this streaming reply information according to punctuation marks, we can obtain 7 text slices: "The basic pressure of electrolytic aluminum usually refers to the cell voltage of the electrolytic cell", "The cell voltage refers to the minimum voltage to maintain the normal operation of the electrolytic cell", "It consists of the following parts", "Anode voltage drop", "Related to anode current density", "Resistivity of anode carbon block" and "Assembly quality of anode carbon block steel claw-guide rod and contact between guide rod and busbar".

[0049] The length of the text "Assembly quality of anode carbon block steel claw-guide rod and contact condition between guide rod and busbar" in the above 7 text slices reaches the set threshold. Therefore, the word segmenter is also used to divide it into two text slices, "Assembly quality of anode carbon block steel claw-guide rod" and "Assembly quality of contact condition between guide rod and busbar" based on the connecting word "and" in the text slice.

[0050] S202: The processor uses a semantic vector extraction model of a target TTS model to identify a semantic feature vector of each text slice.

[0051] Specifically, the TTS (Text-to-Speech) model is used to convert text into audio. The TTS model includes multiple functional models, and the semantic vector extraction model is a model in the TTS model used to identify the semantics of the text. Therefore, after determining multiple text slices, the semantic vector extraction model of the target TTS model can be used to identify the semantic feature vector of each text slice, so as to subsequently synthesize the audio corresponding to the text slice based on the semantic feature vector.

[0052] S203, the processor uses the emotion vector generation model of the target TTS model to process the audio file and emotion parameters of the target person to determine the emotion feature vector of the target person.

[0053] Specifically, the emotion vector generation model is a model used to identify the emotional characteristics of a character in the TTS model.

[0054] Specifically, the target person is the person that the digital human is to simulate (for example, when the voice broadcast system based on digital human technology is applied to an aluminum electrolytic workshop, the target person may be the section chief of the aluminum electrolytic workshop).

[0055] Specifically, the target person's audio features such as timbre and intonation can be identified based on the target person's audio file, and the target person's emotions (such as happiness, sadness, anger, etc.) can be identified based on the target person's emotional parameters. Therefore, by processing the target person's audio file and emotional parameters, the target person's emotional feature vector can be determined.

[0056] Optionally, the emotional parameters of the target person include: spoken language parameters, pause parameters, laughter parameters, etc.

[0057] For example, various emotion parameters and their related meanings are shown in Table 1 below:

[0058] Table 1

[0059]

[0060] S204, the processor synthesizes the audio file corresponding to each text slice according to the semantic feature vector of each text slice and the emotional feature vector of the target person.

[0061] Specifically, by concatenating the semantic feature vector of each text slice and the emotional feature vector of the target person, each text slice can be converted into an audio file with the audio features and emotions of the target person.

[0062] Specifically, the TTS model includes a synthesizer and a vocoder, which concatenate the semantic feature vector of each text slice and the emotional feature vector of the target person to determine the audio file corresponding to each text slice.

[0063] In the disclosed embodiment, first, the streaming reply information output by the virtual person based on the large language model is processed in slices to obtain multiple text slices, and the semantic vector extraction model of the target TTS model is used to identify the semantic feature vector of each text slice. Then, the emotional vector generation model of the target TTS model is used to process the audio file and emotional parameters of the target person (for example, the section chief of electrolytic aluminum) to determine the emotional feature vector of the target person. Finally, the audio file corresponding to each text slice is synthesized based on the semantic feature vector of each text slice and the emotional feature vector of the target person. It can be seen that in the disclosed embodiment, the emotional feature vector including the audio features (such as intonation, timbre, etc.) of the target person and the emotional parameters (such as happiness, sadness, anger, etc.) of the target person is introduced during audio synthesis, and each time only the audio of a text slice in the synthesized streaming reply information is synthesized, which helps to achieve more detailed emotions and tone. Therefore, the disclosed embodiment can improve the degree of anthropomorphism of the digital person output audio.

[0064] In some embodiments, a semantic vector extraction model of a target TTS model is used to identify a semantic feature vector of each text slice, including: saving multiple text slices to a text queue; reading text slices from the text queue in sequence, and performing colloquial processing on the current text slice according to the professional field to which the streaming reply information belongs, to obtain a colloquial text corresponding to the current text slice; and extracting the semantic feature vector of the colloquial text as the semantic feature vector of the current text slice.

[0065] Specifically, after the streaming reply information output by the large language model is processed in pieces to determine multiple text slices, the multiple text slices are saved to the text queue, and then the text slices are read sequentially from the text queue for subsequent processing. This can improve the processing efficiency of the streaming reply information and avoid the situation where text slices are lost, resulting in missing audio content generated later.

[0066] It is understandable that the current text slices may contain a lot of professional terms and specific expressions, which are difficult for non-professionals to understand. Through colloquial processing, these professional terms and expressions can be converted into more understandable language, making it easier to extract the semantics.

[0067] It should be noted that the professional terms in the text slices should not be too colloquial, so as to prevent the colloquial text from failing to express the original meaning. Taking the streaming reply information in the field of electrolytic aluminum as an example, when processing the text slices in colloquial language, the colloquial degree can be reduced for the professional knowledge text of electrolytic aluminum, and the colloquial degree can be increased for other introductory texts.

[0068] In this embodiment, when identifying the semantic feature vector of the current text slice, first, the current text slice is colloquially processed according to the professional field to which the reply information belongs, and then the semantic feature vector of the colloquial text is extracted as the semantic feature vector of the current text slice. This is conducive to improving the accuracy of the identified semantic features.

[0069] In addition, by processing the text slices in a colloquial way, the subsequently generated audio files can be more consistent with the characteristics of the person's speech.

[0070] In some embodiments, an emotion vector generation model of a target TTS model is used to process the audio file and emotion parameters of a target person to determine the emotion feature vector of the target person, including: determining the audio feature vector of the target person based on the audio file of the target person; encoding the emotion parameters of the target person into a vector form to obtain an emotion parameter vector; and concatenating and fusing the audio feature vector and the emotion parameter vector to determine the emotion feature vector of the target person.

[0071] Specifically, an encoding layer can be used to convert the audio file into a vector form to determine the audio feature vector of the target person. The dimension of the audio feature vector is [1, num_wav_token, d_emb], where num_wav_token is the number of tokens the audio file is segmented into, and d_emb is the embedding dimension of each token.

[0072] Specifically, by encoding the target person's emotional parameters (e.g., onehot encoding), the emotional parameters can be converted into text form. The emotional parameters in text form are: Text = [oral_x1], [speed_x2], [laugh_x3], [break_x4]. Among them, x1, x2, x3, x4 are specific emotional parameter values. By segmenting the emotional parameters in text form and vectorizing the segmented results through the embedding layer, the emotional parameter vector can be determined.

[0073] Specifically, after determining the audio feature vector and the emotion parameter vector, the two are concatenated and fused to obtain the final emotion feature vector of the target person. The concatenation and fusion of the audio feature vector and the emotion parameter vector is performed by inputting the audio feature vector and the emotion parameter vector into an emotion vector generation model (the emotion vector generation model here can adopt the LLAMA model).

[0074] Specifically, when the audio feature vector and the emotion parameter vector are spliced ​​and fused through the emotion vector generation model, a sampling and penalty mechanism can be used to make the emotion vector generation model cyclically execute the processing and sampling steps. In each cycle, further generation and refinement are performed based on the output of the previous cycle, and the generated emotion feature vector is gradually optimized. Among them, when the sampling and penalty mechanism is used, the last output of the emotion vector generation model is the final emotion feature vector.

[0075] In some embodiments, an audio file corresponding to each text slice is synthesized based on the semantic feature vector of each text slice and the emotional feature vector of the target person, including: concatenating the semantic feature vector of each text slice with the emotional feature vector of the target person to determine the Mel spectrum of each text slice; converting the Mel spectrum of each text slice into an audio waveform to determine the audio file corresponding to each text slice.

[0076] Specifically, the synthesizer in the disclosed embodiment adopts the Tacotron2 architecture without the WaveNet network, including an encoder, an attention mechanism (position-sensitive attention mechanism) and a decoder. The encoder is provided with a two-layer LSTM memory network (Long Short-Term Memory) to avoid the problem of subsequence duplication or missing during the decoding process.

[0077] Specifically, in order to make the final generated audio file have both the timbre and emotional characteristics of the target person, the Tacontron2 architecture is modified as follows in the disclosed embodiment: the emotional feature vector of the target person is input into the synthesizer, and is spliced ​​with the semantic feature vector generated by the encoder, and the generation process of the Mel spectrum is adjusted. The decoder predicts the corresponding Mel spectrum based on the semantic feature vector in the encoder input sequence, and uses the emotional feature vector of the target person to change the relevant features of the Mel spectrum so that it can express the emotional characteristics of the target person, and finally outputs the Mel spectrum with the audio characteristics and emotions of the target person.

[0078] Specifically, the TTS model includes a vocoder, and the vocoder of the target TTS model can convert the Mel spectrum of each text slice into an audio waveform to determine the audio file corresponding to each text slice.

[0079] Optionally, the vocoder of the target TTS model can be an improved WaveRNN, whose architecture is as follows: Figure 3 As shown. The vocoder can process the input Mel spectrum through an upsampling network and correct the length of the Mel spectrum to be consistent with the input target waveform. The residual network can extract the information of the emotional feature vector in the Mel spectrum and use this information to adjust the generation of the audio waveform so that the generated audio waveform has the information of the emotional feature vector of the target person. GRU is a recurrent neural network that can solve the long-term memory problem of the vocoder. Using GRU to model waveform information in the vocoder can increase the speed at which the vocoder converts the Mel spectrum into an audio waveform.

[0080] The present disclosure provides another method for generating anthropomorphic audio. Figure 4 As shown, the generation method includes:

[0081] S401, the processor obtains streaming response information output by the large language model, and processes the streaming response information in segments to determine multiple text segments.

[0082] S402, the processor inputs the current text slice into the model pool, and determines a target TTS model in the model pool for processing the current text slice.

[0083] Specifically, a model pool with multiple TTS models deployed is pre-built, and each TTS model is set with a different socket address. When processing the current text slice, the current text slice is input into the model pool to determine the target TTS model in the model pool that can be used except the current text slice.

[0084] Specifically, if the TTS model is accessed by n clients or threads T1, T2, ...Tn simultaneously, the number of possible conflicts is calculated by the following expression:

[0085]

[0086] Where C is the number of conflicts, P(T i ∩T j ) is thread T i and T j The probability of accessing the TTS model simultaneously.

[0087] Specifically, it can be seen from the above expression that when multiple clients or threads access the TTS model at the same time, high concurrency problems will occur. Therefore, in the embodiment of the present disclosure, a model pool including multiple TTS models is constructed. When processing the current text slice, the current text slice is input into the model pool to determine the target TTS model in the model pool that can be used to process the current text slice. In this way, only one client or thread can access a TTS model resource at any time.

[0088] S403: The processor uses the semantic vector extraction model of the target TTS model to identify the semantic feature vector of each text slice.

[0089] S404, the processor uses the emotion vector generation model of the target TTS model to process the audio file and emotion parameters of the target person to determine the emotion feature vector of the target person.

[0090] S405, the processor synthesizes the audio file corresponding to each text slice according to the semantic feature vector of each text slice and the emotional feature vector of the target person.

[0091] In the disclosed embodiment, when processing the current text slice, the current text slice will first be input into the model pool, and the target TTS model in the model pool for processing the current text slice will be determined. In this way, it is possible to avoid handing over the current text slice to the TTS model being accessed for processing. This is helpful in reducing the blocking phenomenon caused by insufficient model resources when multiple clients or multiple threads access the TTS model at the same time, and solving the high concurrency problem caused by multiple clients and multiple threads generating anthropomorphic audio at the same time.

[0092] In some embodiments, after synthesizing the audio file corresponding to each text slice, the generation method further includes: saving the audio file corresponding to each text slice to an audio queue, so that the audio player sequentially reads the audio files from the audio queue for playing.

[0093] Specifically, by saving the audio file corresponding to each text slice to the audio queue, when the voice broadcast system based on digital human technology needs to broadcast the streaming reply information of the large voice model, the audio player of the voice broadcast system can read the audio files from the audio queue in sequence for playback. In this way, low-latency and fast voice broadcast can be achieved.

[0094] Specifically, when a text queue and an audio queue are established at the same time, a double-layer queue algorithm can be implemented. That is, multiple text slices determined by the streaming reply information output by the large language model are saved in the text queue. The text queue pops up the text slices in order, and the target TTS model processes them, determines the audio files corresponding to the text slices, and saves the audio files in the audio queue. Finally, the audio player plays the audio files in the audio queue in order.

[0095] In this embodiment, a double-layer queue algorithm including a text queue and an audio queue is adopted, which can improve the speed of converting the streaming reply information of the large language model into an audio file, avoid the jamming when the audio player plays the audio, and realize the rapid connection between the streaming reply information of the large language model and the audio file broadcast.

[0096] Combination Figure 5 As shown, the embodiment of the present disclosure provides a device 500 for generating anthropomorphic audio, including: a slicing module 501, a semantic recognition module 502, an emotion recognition module 503 and a synthesis module 504. The slicing module 501 is configured to obtain the streaming reply information output by the large language model, and process the streaming reply information in segments to determine multiple text slices. The semantic recognition module 502 is configured to adopt the semantic vector extraction model of the target TTS model to identify the semantic feature vector of each text slice. The emotion recognition module 503 is configured to adopt the emotion vector generation model of the target TTS model, process the audio file and emotion parameters of the target character, and determine the emotion feature vector of the target character. The synthesis module 504 is configured to synthesize the audio file corresponding to each text slice according to the semantic feature vector of each text slice and the emotion feature vector of the target character.

[0097] Combination Figure 6 As shown, the embodiment of the present disclosure provides a device 600 for generating anthropomorphic audio, including: a processor (processor) 601 and a memory (memory) 602. Optionally, the device may also include a communication interface (Communication Interface) 603 and a bus 604. Among them, the processor 601, the communication interface 603, and the memory 602 can communicate with each other through the bus 604. The communication interface 603 can be used for information transmission. The processor 601 can call the logic instructions in the memory 602 to execute the method for generating anthropomorphic audio of the above embodiment.

[0098] In addition, the logic instructions in the memory 602 described above may be implemented in the form of software functional units and when sold or used as independent products, may be stored in a computer-readable storage medium.

[0099] The memory 602 is a computer-readable storage medium that can be used to store software programs and computer executable programs, such as program instructions / modules corresponding to the method in the embodiment of the present disclosure. The processor 601 executes the function application and data processing by running the program instructions / modules stored in the memory 602, that is, the method for generating anthropomorphic audio in the above embodiment is implemented.

[0100] The memory 602 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and an application required for at least one function; the data storage area may store data created according to the use of the terminal device, etc. In addition, the memory 602 may include a high-speed random access memory and may also include a non-volatile memory.

[0101] An embodiment of the present disclosure provides a computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions are configured to execute the above-mentioned method for generating anthropomorphic audio.

[0102] The technical solution of the embodiment of the present disclosure can be embodied in the form of a software product, which is stored in a storage medium and includes one or more instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in the embodiment of the present disclosure. The aforementioned storage medium may be a non-transient storage medium, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a disk or an optical disk, and other media that can store program codes.

[0103] The above description and the accompanying drawings fully illustrate the embodiments of the present disclosure so that those skilled in the art can practice them. Other embodiments may include structural, logical, electrical, process and other changes. The embodiments represent only possible changes. Unless explicitly required, separate components and functions are optional, and the order of operation may vary. The parts and features of some embodiments may be included in or replace the parts and features of other embodiments. Moreover, the words used in this application are only used to describe the embodiments and are not used to limit the claims. As used in the description of the embodiments and the claims, unless the context clearly indicates, the singular forms of "a", "an" and "the" are intended to include plural forms as well. Similarly, the term "and / or" as used in this application refers to any and all possible combinations of listings containing one or more associated ones. In addition, when used in the present application, the term "comprise" and its variants "comprises" and / or comprising refer to the presence of stated features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or groups thereof. In the absence of further restrictions, the elements defined by the sentence "comprising a ..." do not exclude the presence of other identical elements in the process, method or device comprising the elements. In this article, each embodiment may focus on the differences from other embodiments, and the same and similar parts between the various embodiments may refer to each other. For the methods, products, etc. disclosed in the embodiments, if they correspond to the method part disclosed in the embodiments, then the relevant parts can refer to the description of the method part.

[0104] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software may depend on the specific application and design constraints of the technical solution. The technicians may use different methods for each specific application to implement the described functions, but such implementations should not be considered to exceed the scope of the embodiments of the present disclosure. The technicians may clearly understand that, for the convenience and simplicity of description, the specific working processes of the systems, devices and units described above may refer to the corresponding processes in the aforementioned method embodiments, and will not be repeated here.

[0105] In the embodiments disclosed herein, the disclosed methods and products (including but not limited to devices, equipment, etc.) can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units can be only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between each other shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms. The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the units may be selected according to actual needs to implement this embodiment. In addition, each functional unit in the embodiment of the present disclosure may be integrated in a processing unit, or each unit may exist physically alone, or two or more units may be integrated in one unit.

[0106] The flowchart and block diagram in the accompanying drawings show the possible architecture, function and operation of the system, method and computer program product according to the embodiment of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of the code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. In some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, which can depend on the functions involved. In the description corresponding to the flowchart and the block diagram in the accompanying drawings, the operations or steps corresponding to different boxes can also occur in a different order from the order disclosed in the description, and sometimes there is no specific order between different operations or steps. For example, two consecutive operations or steps can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, which can depend on the functions involved. Each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented by a dedicated hardware-based system that performs the specified functions or actions, or may be implemented by a combination of dedicated hardware and computer instructions.

Claims

1. A method for generating anthropomorphic audio, characterized in that: include: Obtain the streaming response information output by the large language model, and process the streaming response information in segments to determine multiple text slices; Using the semantic vector extraction model of the target TTS model, the semantic feature vector of each text slice is identified; The target TTS model's emotion vector generation model is used to process the target person's audio file and emotion parameters to determine the target person's emotion feature vector; According to the semantic feature vector of each text slice and the emotional feature vector of the target person, the audio file corresponding to each text slice is synthesized.

2. The generation method according to claim 1, characterized in that: The sharding process streams the reply information to determine multiple text slices, including: Slicing the streaming reply information according to punctuation marks to determine a plurality of text slices; Confirm the text length of each text slice; The word segmenter is used to re-slice text slices whose length is greater than or equal to the set threshold.

3. The generation method according to claim 1, characterized in that: The semantic vector extraction model of the target TTS model is used to identify the semantic feature vector of each text slice, including: Save multiple text slices to the text queue; Read text slices from the text queue in sequence, and perform colloquial processing on the current text slice according to the professional field to which the streaming reply information belongs, to obtain the colloquial text corresponding to the current text slice; The semantic feature vector of the spoken text is extracted as the semantic feature vector of the current text slice.

4. The generation method according to claim 1, characterized in that: The target TTS model's emotion vector generation model is used to process the target person's audio file and emotion parameters to determine the target person's emotion feature vector, including: Determine the audio feature vector of the target person according to the audio file of the target person; Encode the target person's emotional parameters into a vector form to obtain an emotional parameter vector; The audio feature vector and the emotion parameter vector are concatenated and fused to determine the emotion feature vector of the target person.

5. The generation method according to any one of claims 1 to 4, characterized in that: According to the semantic feature vector of each text slice and the emotional feature vector of the target person, the audio file corresponding to each text slice is synthesized, including: The semantic feature vector of each text slice is concatenated with the emotional feature vector of the target person to determine the Mel spectrum of each text slice; Convert the Mel spectrum of each text slice into an audio waveform and determine the audio file corresponding to each text slice.

6. The generation method according to any one of claims 1 to 4, characterized in that: Before using the semantic vector extraction model of the target TTS model to identify the semantic feature vector of each text slice, the generation method further includes: The current text slice is input into the model pool, and a target TTS model in the model pool for processing the current text slice is determined.

7. The generation method according to any one of claims 1 to 4, characterized in that: After synthesizing the audio file corresponding to each text slice, the generation method further includes: The audio file corresponding to each text slice is saved to the audio queue, so that the audio player reads the audio files from the audio queue in sequence for playback.

8. A device for generating anthropomorphic audio, characterized in that: include: A slicing module is configured to obtain streaming response information output by the large language model, and process the streaming response information in segments to determine a plurality of text slices; A semantic recognition module, configured to adopt a semantic vector extraction model of a target TTS model to recognize a semantic feature vector of each text slice; The emotion recognition module is configured to adopt the emotion vector generation model of the target TTS model, process the audio file and emotion parameters of the target person, and determine the emotion feature vector of the target person; The synthesis module is configured to synthesize the audio file corresponding to each text slice according to the semantic feature vector of each text slice and the emotional feature vector of the target person.

9. A device for generating anthropomorphic audio, comprising a processor and a memory storing program instructions, characterized in that: The processor is configured to be able to execute the method for generating human-like audio according to any one of claims 1 to 7.

10. An electronic device, characterized in that: include: Equipment body; The device for generating anthropomorphic audio as described in claim 8 or 9 is arranged on the device body.

Citation Information

Cited By

  • Data transmission method and device, equipment, storage medium and program product

    CN120528959A

  • Methods, apparatus, equipment, storage media and software products for data transmission

    CN120528959B