3D digital human system based on large model native streaming audio interaction
By adopting streaming audio interaction technology based on end-to-end voice model in the 3D digital human system, the problem of insufficient real-time response and voice comprehension in traditional systems is solved, and a higher sense of user experience and system intelligence is achieved, which is suitable for a variety of application scenarios.
Patent Information
- Application Number
- CN202510079485.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-18
- Publication Date
- 2025-05-06
AI Technical Summary
The traditional audio interactive 3D digital human system has difficulties in real-time response, insufficient voice comprehension, and the inability to accurately capture user intentions and generate context-compatible answers, resulting in low user experience of realism and immersion.
The 3D digital human-streaming audio interaction system based on the end-to-end voice model is adopted. The user's audio signal reception module is used to receive user voice in real time, and the voice dialogue interaction feedback module is used to perform voice recognition, embed coding and semantic enhancement processing, generate feedback audio signals, and stream drive through the 3D digital human-drive module to realize real-time adjustment of expressions, body movements and lip shape changes.
It improves the realism and immersion of interaction, enhances the intelligence level of the 3D digital human system, realizes instant response and efficient communication, and opens up new possibilities for various application scenarios.
Smart Images

Figure CN119943044A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of digital human technology, and more specifically, to a 3D digital human streaming audio interaction system based on an end-to-end speech large model. Background Art
[0002] With the development of artificial intelligence, computer vision and computer graphics, 3D digital human technology has gradually become an important part of virtual reality (VR), augmented reality (AR), games, education, customer service and other fields. 3D digital human refers to three-dimensional character models generated by computers, which can simulate human appearance, movements and even emotional expressions. Audio interaction is one of the key ways for 3D digital humans to communicate with users. It allows users to interact with digital humans through voice commands or conversations, and digital humans can respond to users in the form of voice feedback.
[0003] Although 3D digital human technology has made great progress, traditional audio interaction systems still have some limitations. For example, traditional audio interaction systems may not be able to achieve real-time response, causing users to feel that the waiting time is too long, affecting the realism and immersion of the user experience. This is mainly due to high latency in the data transmission process or the failure to complete complex computing tasks in a timely manner. In addition, traditional audio interactive 3D digital humans may have deficiencies in understanding and generating speech, and cannot accurately capture the user's intentions or produce responses that fit the context, causing the communication to appear mechanical rather than human. In addition, non-verbal cues such as the 3D digital human's body language and facial expressions are often ignored, reducing the authenticity and coherence of the interaction.
[0004] Therefore, an optimized audio interactive 3D digital human system is desired. Summary of the invention
[0005] The present application provides a 3D digital human streaming audio interaction system based on an end-to-end voice large model, which not only improves the realism and immersion of the interaction, but also enhances the intelligence level of the 3D digital human system, opening up new possibilities for efficient communication in various application scenarios.
[0006] In a first aspect, a 3D digital human streaming audio interaction system based on an end-to-end speech large model is provided, comprising: A user audio signal receiving module is used to receive the user's voice interaction signal in real time through a streaming audio input interface; A voice dialogue interaction feedback module, used for inputting the voice interaction signal into a voice dialogue engine based on a large model to obtain a feedback audio signal; The 3D digital human driving module is used to drive the 3D digital human model using audio that can be rendered, and to perform streaming driving of the 3D digital human's expression, body movements and lip shape changes based on the feedback audio signal.
[0007] In a possible implementation, the voice dialogue interaction feedback module includes: A speech interaction signal recognition unit, configured to perform speech recognition on the speech interaction signal using an audio encoder to obtain a speech interaction signal encoding result; A speech interaction content embedding coding unit, used for performing word-granularity-based embedding coding on the speech interaction signal coding result to obtain a sequence of speech interaction content word-granularity semantic embedding coding features; The speech interaction content semantic enhancement processing subunit is used to perform semantic enhancement processing on the sequence of the speech interaction content word granularity semantic embedding coding features based on grammatical depth association constraints to obtain a sequence of medium-granularity speech interaction content semantic enhancement word features, It is represented as the aggregation of contextual speech elements after encoding the speech interaction content embedding coding unit: , Where k is the selected number of consecutive frames, and this step can complete the enhanced splicing of k consecutive frames along the feature dimension; A speech interaction content context semantic encoding unit is used to input the sequence of medium-granularity speech interaction content semantic reinforcement word features into the audio adapter to obtain speech interaction content context semantic encoding features, Represents the concatenated array of audio elements processed by the speech interaction content semantic enhancement processing sub-unit: ; The feedback audio signal generating unit is used to input the semantic coding features of the speech interaction content context into an audio decoder based on a large language model to obtain the feedback audio signal, and pass it through a 2-layer CNN perceptron composed of MBConv6 blocks. The activation generates the final speech representation: .
[0008] In one possible implementation, the speech interaction content embedding encoding unit is used to perform word segmentation processing on the speech interaction signal text recognition result and then pass it through a word embedding encoder based on the Bert model to obtain a sequence of speech interaction content word granularity semantic embedding encoding vectors as a sequence of speech interaction content word granularity semantic embedding encoding features.
[0009] In a possible implementation, the voice interaction content semantic enhancement processing unit includes: A speech interaction content mapping subunit is used to map each speech interaction content word granularity semantic embedding coding vector in the sequence of the speech interaction content word granularity semantic embedding coding vector to a hyperbolic space to obtain a sequence of speech interaction content hyperbolic space word semantic coding feature vectors, using a context-aware mapping method with variable weights as follows; Given a word-level semantic embedding encoding vector sequence ,in , perform adaptive weight adjustment: , in, is the vector sequence after adaptive adjustment, It is The adaptive weight matrix of words, based on Nonlinear transformation of the network, using The network encoder obtains the context vector To achieve a globally consistent understanding of features: ,in, represents the transformed sequence vector, Use the Poincare sphere model to map the transformed vector to hyperbolic space: in, is the vector mapped to the hyperbolic space; The grammatical depth implicit representation value calculation subunit is used to calculate the grammatical depth implicit representation value of each speech interaction content hyperbolic space word semantic encoding feature vector in the sequence of the speech interaction content hyperbolic space word semantic encoding feature vector to obtain the sequence distribution of the grammatical depth implicit representation value of the speech interaction content, including: calculating each word vector Global Grammar Center The hyperbolic distance of , as the basic measure of grammatical depth, the global grammatical center By sequence The Fréchet mean of is calculated as: in is the Poincare distance in hyperbolic space: The implicit representation value of the grammatical depth is calculated. The implicit representation value of the grammatical depth is Defined as word vector Global Grammar Center The hyperbolic distance is normalized by the Sigmoid function: in is the Sigmoid function; h, Representing the word vector in the hyperbolic space, we finally get the sequence distribution of the grammatical depth implicit representation value of the speech interaction content; A medium-granularity semantic association window determination subunit is used to determine the medium-granularity semantic association window of each speech interaction content word granularity semantic embedding coding vector in the sequence of the speech interaction content word granularity semantic embedding coding vector, that is, the value of the selected continuous frame number k, based on the sequence distribution of the speech interaction content grammatical depth implicit representation value; The medium-granularity contextual semantic reinforcement processing subunit is used to perform medium-granularity contextual semantic reinforcement on each speech interaction content word granularity semantic embedding coding vector in the sequence of speech interaction content word granularity semantic embedding coding vectors based on the medium-granularity semantic association window of each speech interaction content word granularity semantic embedding coding vector in the sequence of speech interaction content word granularity semantic embedding coding vectors to obtain a sequence of medium-granularity speech interaction content semantic reinforcement word feature vectors as a sequence of medium-granularity speech interaction content semantic reinforcement word features.
[0010] In a possible implementation, the medium-granularity semantic association window determination subunit includes: Extracting the grammatical depth implicit representation value of the first target speech interaction content word granularity semantic embedding coding vector in the sequence of the speech interaction content word granularity semantic embedding coding vectors as the starting position of the medium-granularity semantic association window; Along the sequence distribution of the grammatical depth implicit representation value of the speech interaction content, find the second target speech interaction content word granularity semantic embedding coding vector that is closest to the grammatical depth implicit representation value of the first target speech interaction content word granularity semantic embedding coding vector as the termination position of the medium granularity semantic association window; Based on the starting position of the medium-granularity semantic association window and the ending position of the medium-granularity semantic association window, the medium-granularity semantic association window corresponding to the first target speech interaction content word granularity semantic embedding coding vector is determined from the sequence of the speech interaction content word granularity semantic embedding coding vectors.
[0011] In a possible implementation, the speech interaction content context semantic encoding unit is used to: input the sequence of medium-granularity speech interaction content semantic reinforcement word feature vectors into the audio adapter to obtain a speech interaction content context semantic encoding feature vector as the speech interaction content context semantic encoding feature.
[0012] In a possible implementation, the audio decoder is composed of two layers of Transformer layers of the Llama architecture.
[0013] In one possible implementation, the feedback audio signal generating unit includes: inputting the semantic encoding feature vector of the speech interaction content context into a reply feature extractor based on a large language model to obtain a sequence of reply feature embedding encoding vectors; performing a two-fold upsampling process on the sequence of reply feature embedding encoding vectors to obtain a sequence of upsampled reply feature embedding encoding vectors; passing the sequence of upsampled reply feature embedding encoding vectors through a Transformer network based on a two-layer Llama architecture to obtain a sequence of reply feature embedding semantic vectors; using CTC to align the sequence of reply feature embedding semantic vectors with discrete units to obtain a sequence of aligned reply embedding semantic vectors; and inputting the sequence of aligned reply embedding semantic vectors into a vocoder to synthesize audio to obtain the feedback audio signal.
[0014] In a possible implementation, the 3D digital human driving module is used to convert the feedback audio signal into PCM code stream data and then input the PCM code stream data into the audio-driven 3D digital human model capable of rendering through streaming transmission.
[0015] In a possible implementation, the 3D digital human driving module is used to input the feedback audio signal into a streaming audio decoder based on the EmoTalk model for processing to generate a driving signal to perform streaming driving on the expression, body movements and lip shape changes of the 3D digital human.
[0016] The present application provides a 3D digital human streaming audio interaction system based on an end-to-end voice large model, which uses a streaming audio input interface to receive the user's voice commands and interactive content in real time, ensuring the continuity and low latency of data transmission, thereby providing an instant response user experience. This allows the 3D digital human to respond to the user's questions or commands as quickly as a real person. In addition, an advanced large model is used to extract voice features and perform semantic analysis, which helps to understand the semantics of the user's voice interaction content more timely and accurately, generate accurate reply features, and convert the reply features into feedback audio signals to achieve streaming drive of the 3D digital human. In this way, not only the realism and immersion of the interaction are improved, but also the intelligence level of the 3D digital human system is enhanced, opening up new possibilities for efficient communication in various application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings of the embodiments of the present application are briefly introduced below. Obviously, the drawings described below only relate to some embodiments of the present application, and are not intended to limit the present application.
[0018] Figure 1 It is a schematic block diagram of a 3D digital human streaming audio interaction system based on an end-to-end speech large model according to an embodiment of the present application.
[0019] Figure 2 This is a schematic flow chart of a voice dialogue interaction feedback module in a 3D digital human streaming audio interaction system based on an end-to-end voice large model according to an embodiment of the present application.
[0020] Figure 3 This is a data flow diagram of a 3D digital human streaming audio interaction system based on an end-to-end speech large model according to an embodiment of the present application.
[0021] Figure 4 This is a schematic block diagram of a speech interaction content semantic enhancement processing unit in a 3D digital human streaming audio interaction system based on an end-to-end speech large model according to an embodiment of the present application. DETAILED DESCRIPTION
[0022] The technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, not all embodiments. Based on the embodiments of the present application, all other embodiments obtained by ordinary technicians in this field without creative work also fall within the scope of protection of the present application.
[0023] In recent years, the advancement of deep learning and natural language processing (NLP) technology has brought a qualitative leap in the audio interaction capabilities of 3D digital humans. Large models such as BERT, GPT series, and models optimized for specific tasks have greatly improved the ability of machines to understand human language, thereby promoting a more intelligent and natural conversation experience. At the same time, the application of real-time rendering technology and streaming processing frameworks has significantly improved the response speed and interaction fluency of 3D digital humans.
[0024] Based on this, in the technical solution of this application, a 3D digital human streaming audio interaction system based on an end-to-end voice large model is proposed, which combines the latest deep learning and artificial intelligence technologies, especially the application of large models, to provide a more natural and real-time user interaction experience. At the same time, in the interaction of 3D digital humans, in order to achieve smooth interaction with users, native streaming technology plays a vital role. It optimizes the utilization of computing resources and processes data streams in parallel through distributed computing, achieving more efficient hardware resource utilization and improving the speed and efficiency of data processing. How to combine the large model to realize native streaming to build a 3D digital human real-time interaction system that needs to process large amounts of data and complex calculations is a problem that needs to be solved.
[0025] In view of the above technical problems, in the technical solution of this application, a 3D digital human streaming audio interaction system based on an end-to-end speech large model is proposed, such as Figure 1 As shown, the 3D digital human streaming audio interaction system based on the end-to-end voice big model includes: a user audio signal receiving module 10, used to execute S1: receiving the user's voice interaction signal in real time through the streaming audio input interface; a voice dialogue interaction feedback module 20, used to execute S2: inputting the voice interaction signal into the voice dialogue engine based on the big model to obtain a feedback audio signal; a 3D digital human driving module 30, used to execute S3: using audio that can be rendered to drive the 3D digital human model, and streaming driving the 3D digital human's expression, body movement and lip shape changes based on the feedback audio signal.
[0026] Exemplarily, in the user audio signal receiving module 10, S1 is executed: receiving the user's voice interaction signal in real time through the streaming audio input interface. It should be understood that when the user starts to communicate with the 3D digital person, their voice will be captured through a microphone or other recording device and converted into a digitized audio signal. These audio signals are then immediately sent to the streaming audio input interface of the system. In one embodiment, the streaming audio input interface is a large model native input interface for implementing streaming audio input based on a large model. Specifically, the streaming audio input interface is specially customized for a large model. It can be directly connected to the latest deep learning architecture. Such a configuration allows the system to perform efficient and accurate feature extraction on each incoming audio segment, maintaining good performance even in a noisy environment. At the same time, due to the use of a large model native interface, the entire process can seamlessly connect to the subsequent large-scale language model processing steps. In addition, the streaming audio input interface also supports multi-channel audio input, which means that it can receive sound information from multiple directions at the same time. This is very useful for simulating real face-to-face conversation scenarios, because people often have movements such as head rotation when speaking, causing the sound source to change. With multi-channel support, the system can locate the sound source more accurately, further enhancing the realism and immersion of the interaction.
[0027] Exemplarily, in the voice dialogue interaction feedback module 20, S2 is executed: the voice interaction signal is input into the voice dialogue engine based on the large model to obtain a feedback audio signal. In particular, the above-mentioned 3D digital human streaming audio interaction system based on the end-to-end voice large model adopts a streaming audio input interface to realize the real-time reception of the user's voice commands and interactive content, ensuring the continuity and low latency of data transmission, thereby providing an instant response user experience. This allows the 3D digital human to respond to the user's questions or commands as quickly as a real person. In addition, an advanced large model is used for voice feature extraction and semantic analysis, which helps to understand the semantics of the user's voice interaction content more timely and accurately, generate accurate reply features, and convert the reply features into feedback audio signals to realize the streaming drive of the 3D digital human.
[0028] Specifically, in the process of using a large model-based voice dialogue engine to process the voice interaction signal to generate a feedback audio signal, the timeliness and accuracy of the voice processing and semantic understanding of the large model can further affect the audio interaction experience of the 3D digital human. The technical concept of this application is to introduce a data processing and semantic understanding algorithm based on an artificial intelligence large model and natural language processing at the back end to analyze the text recognition result of the voice interaction signal after performing voice recognition on the voice interaction signal to generate a text recognition result, so as to capture the contextual semantic association features of the word granularity of the voice interaction content, and use the contextual semantics of the voice interaction content to generate a feedback audio signal through audio adaptation coding and large language model processing to input audio to drive the 3D digital human model to realize the subsequent 3D digital human expression, body movements and lip shape change streaming drive. In this way, not only the sense of reality and immersion of the interaction is improved, but also the intelligence level of the 3D digital human system is enhanced, opening up new possibilities for efficient communication in various application scenarios.
[0029] In one embodiment, Figure 2 and Figure 3 As shown, the voice dialogue interaction feedback module 20 includes: a voice interaction signal recognition unit 21, which is used to perform voice recognition on the voice interaction signal using an audio encoder to obtain a voice interaction signal text recognition result; a voice interaction content embedding encoding unit 22, which is used to perform word granularity-based embedding encoding on the voice interaction signal text recognition result to obtain a sequence of voice interaction content word granularity semantic embedding encoding features; a voice interaction content semantic enhancement processing unit 23, which is used to perform semantic enhancement processing based on grammatical deep association constraints on the sequence of voice interaction content word granularity semantic embedding encoding features to obtain a sequence of medium-granularity voice interaction content semantic enhancement word features, which can be more specifically expressed as an aggregation of contextual voice elements encoded by the voice interaction content embedding encoding unit: Wherein k is the selected number of consecutive frames, and this step can complete the enhanced splicing of k consecutive frames along the feature dimension; the speech interaction content context semantic encoding unit 24 is used to input the sequence of the medium-granularity speech interaction content semantic enhancement word features into the audio adapter to obtain the speech interaction content context semantic encoding features, which can be more specifically represented as a spliced array of audio elements processed by the speech interaction content semantic enhancement processing subunit: ; The feedback audio signal generating unit 25 is used to input the semantic coding features of the speech interaction content context into an audio decoder based on a large language model to obtain the feedback audio signal, more specifically, through a 2-layer CNN perceptron composed of similar MBConv6 blocks. The activation generates the final speech representation: .
[0030] Exemplarily, in the voice interaction signal recognition unit 21, an audio encoder is used to perform voice recognition on the voice interaction signal to obtain a voice interaction signal text recognition result. It should be understood that since the user's interactive voice is a continuous, unstructured audio signal, and computers are better at processing discrete, structured text data. By performing voice recognition on the voice interaction signal through an audio encoder, the user's voice instructions can be converted into a text form that the machine can understand and process, which is the basis for further analysis and response. That is, once the voice is transcribed into text, the system can use natural language processing (NLP) technology to parse these voice interaction signal text recognition results to better understand the user's request or question. In the technical solution of the present application, Whisper-large-v3 can be used as an audio encoder to perform voice recognition on the voice interaction signal, which has a high voice-to-text accuracy rate, and can maintain good performance and stability even in a noisy environment, different speaking speeds and different accents. This ensures the quality of the conversion from speech to text, reduces the possibility of misrecognition, and thus improves the effect of subsequent natural language processing tasks.
[0031] More specifically, the workflow of Whisper-large-v3 is as follows. First, when the user's voice is captured by the microphone, it will go through a series of preprocessing steps, including removing background noise, normalizing the volume, etc., to ensure clear and stable recording quality even in a noisy environment. This step is very important for improving the effect of subsequent feature extraction. Next, the encoder will perform feature extraction on each small audio clip to extract important information that can represent the content of the voice. For example, Mel-spectrogram is calculated, convolutional neural network (CNN) is used for acoustic modeling, and long short-term memory network (LSTM) is used to capture dependencies on time series. All of this is to extract the most semantic information from the original audio. Once the preliminary feature extraction is completed, the next step is the actual decoding process. At this stage, Whisper-large-v3 will construct the most likely text sequence based on the previously extracted features. Due to the diversity of natural language, the same sound clip may correspond to multiple different text interpretations, so the encoder needs to use probabilistic statistical methods to select the best solution from them. For example, when faced with some unclear or accented pronunciations, it may consider multiple possibilities and make the best choice based on contextual information. It is worth noting that in addition to single-language recognition, Whisper-large-v3 also supports multi-language recognition, which means it can process voice input from users in different countries and regions, further expanding the scope of applicability of the system.
[0032] Exemplarily, in the voice interaction content embedding coding unit 22, the voice interaction signal text recognition result is embedded based on word granularity to obtain a sequence of semantic embedding coding features of the voice interaction content word granularity. It should be understood that natural language is highly complex and diverse, and the same word may have different meanings in different contexts. In order to accurately understand the user's intention, the system needs to have fine-grained language parsing capabilities. By mapping each word in the text into a high-dimensional vector representation (i.e., embedding coding), its semantic information can be better retained. Traditional rule-based methods are difficult to handle this phenomenon of polysemy, while modern word embedding technology can automatically learn the specific meaning of each word in different contexts through a large amount of training data, thereby improving the system's understanding and response accuracy. At the same time, word-granular embedding coding helps to build a richer semantic representation. Compared with simple character or word level processing, word-level embedding can combine more levels of information, such as part of speech, grammatical role, etc., to form a more comprehensive sentence structure description. This is particularly important for capturing long-distance dependencies, because many important semantic clues are often scattered throughout a sentence or even a paragraph. In this way, the system can more accurately grasp the core idea expressed by the user and will not get lost even when faced with complex sentence structures.
[0033] In one embodiment, the speech interaction content embedding encoding unit is used to: after the speech interaction signal text recognition result is processed by word segmentation, a word embedding encoder based on the Bert model is used to obtain a sequence of the speech interaction content word granularity semantic embedding encoding vectors as a sequence of the speech interaction content word granularity semantic embedding encoding features. That is, after the speech interaction signal text recognition result is processed by word segmentation, it is encoded by a word embedding encoder based on the Bert model to extract the semantic association feature representation information based on word granularity in the speech interaction signal text recognition result, thereby obtaining a sequence of speech interaction content word granularity semantic embedding encoding vectors. It should be understood that BERT is a bidirectional encoder representation method that can simultaneously consider the context information on both sides of a word, thereby generating a vector representation containing rich context information for each word. This means that even if the same word has different meanings in different sentences, BERT can give an appropriate semantic interpretation according to the specific context. For complex natural languages, this fine-grained understanding can significantly improve the system's grasp of user intent. Specifically, the speech interaction signal text recognition result will first be processed by word segmentation to cut the continuous text into independent vocabulary units. These vocabulary units are then fed into a word embedding encoder based on the Bert model for encoding, and the semantic association feature representation information based on word granularity is extracted, thereby obtaining a sequence of word granular semantic embedding encoding vectors of the speech interaction content. In this process, the Bert model will make full use of the advantages of its internal multi-layer Transformer architecture, focus on important information within the entire sentence through the self-attention mechanism, and integrate it into the embedding representation of each word. In this way, not only the meaning of individual words is accurately portrayed, but also the potential connections between them are fully revealed.
[0034] Exemplarily, in the speech interaction content semantic enhancement processing unit 23, the sequence of the speech interaction content word granularity semantic embedding coding features is subjected to semantic enhancement processing based on grammatical deep association constraints to obtain a sequence of medium-granularity speech interaction content semantic enhancement word features. It should be understood that each speech interaction content word granularity semantic embedding coding vector in the sequence of speech interaction content word granularity semantic embedding coding vectors respectively contains word granularity embedded semantic coding features related to the user's speech interaction content. This implied coding feature cannot reflect long-distance dependencies and contextual semantic feature information, nor can it capture the grammatical role of vocabulary in a sentence, which will produce a phenomenon of polysemy, resulting in a decrease in the accuracy of understanding within the speech interaction. Based on this, in order to be able to understand and enhance the semantic information in the text at a deeper level, in the technical solution of the present application, the sequence of the speech interaction content word granularity semantic embedding coding features is further subjected to semantic enhancement processing based on grammatical deep association constraints to obtain a sequence of medium-granularity speech interaction content semantic enhancement word features.
[0035] In one embodiment, Figure 4 As shown, the speech interaction content semantic enhancement processing unit 23 includes: a speech interaction content mapping subunit 231, which is used to map each speech interaction content word granularity semantic embedding coding vector in the sequence of the speech interaction content word granularity semantic embedding coding vector to a hyperbolic space to obtain a sequence of speech interaction content hyperbolic space word semantic coding feature vectors. More specifically, here we adopt a context-aware mapping method with variable weights; Given a word-level semantic embedding encoding vector sequence ,in , first we perform adaptive weight adjustment: in It is We then perform an adaptive weight matrix based on Nonlinear transformation of the network, using The network encoder obtains the context vector To achieve a globally consistent understanding of features: Finally, the transformed vector is mapped to the hyperbolic space using the Poincare sphere model: ; The grammatical depth implicit representation value calculation subunit 232 is used to calculate the grammatical depth implicit representation value of each speech interaction content hyperbolic space word semantic encoding feature vector in the sequence of the speech interaction content hyperbolic space word semantic encoding feature vector to obtain the sequence distribution of the grammatical depth implicit representation value of the speech interaction content. More specifically, here we first calculate each word vector Global Grammar Center The hyperbolic distance of , as the basic measure of grammatical depth. Global grammar center Through the sequence The Fréchet mean (hyperbolic space mean) of is calculated as: in is the Poincare distance in hyperbolic space: Then the implicit representation value of the grammatical depth is calculated. Defined as word vector Global Grammar Center The hyperbolic distance is normalized by the Sigmoid function: in is the Sigmoid function; h, Representing the word vector in the hyperbolic space, finally obtaining the sequence distribution of the grammatical depth implicit representation value of the speech interaction content; the medium-granularity semantic association window determination subunit 233 is used to determine the medium-granularity semantic association window size of each speech interaction content word granularity semantic embedding coding vector in the sequence of the speech interaction content word granularity semantic embedding coding vector based on the sequence distribution of the grammatical depth implicit representation value of the speech interaction content, that is, the value of the above-selected number of continuous frames k; the medium-granularity contextual semantic enhancement processing subunit 234 is used to perform medium-granularity contextual semantic enhancement on each speech interaction content word granularity semantic embedding coding vector in the sequence of the speech interaction content word granularity semantic embedding coding vector based on the medium-granularity semantic association window of each speech interaction content word granularity semantic embedding coding vector in the sequence of the speech interaction content word granularity semantic embedding coding vector to obtain a sequence of medium-granularity speech interaction content semantic enhancement word feature vectors as the sequence of medium-granularity speech interaction content semantic enhancement word features.
[0036] Exemplarily, in the voice interaction content mapping subunit 231, each voice interaction content word granularity semantic embedding coding vector in the sequence of the voice interaction content word granularity semantic embedding coding vector is mapped to the hyperbolic space to obtain a sequence of voice interaction content hyperbolic space word semantic coding feature vectors. Specifically, the process can be expressed by the formula: in, is a sequence of word-granular semantic embedding encoding vectors of the speech interaction content, are the first, second, and third words in the sequence of the word-granular semantic embedding coding vectors of the speech interaction content. and The word-granular semantic embedding encoding vector of the speech interaction content, and are the first weight matrix and the second weight matrix respectively, represents a hyperbolic space mapping, for The corresponding hyperbolic space word semantic encoding feature vector of the speech interaction content.
[0037] It should be understood that words in natural language do not exist in isolation, and there is often a complex hierarchical structure between them. For example, when describing an event, words such as "car", "engine", and "tire" may all be associated with "transportation", which can be classified into a larger category such as "mobility". It is difficult for traditional Euclidean space to effectively represent this hierarchical semantic relationship because the distance between all points is linear and cannot intuitively reflect the conceptual inclusion or subordination of certain words. In contrast, hyperbolic space has a natural hierarchy, which can better simulate a tree or network structure, so that words close to the root node (such as "transportation") maintain an appropriate distance from other leaf nodes (such as "car" and "bicycle"), while retaining the relative position between internal nodes. In this way, the logical connection between words can be captured more accurately, thereby improving the quality of dialogue generation. That is, by mapping the granular semantic embedding encoding vector of each voice interaction content word to the hyperbolic space, the hierarchical structure characteristics of hyperbolic geometry can be used to better capture the hierarchical relationship and complex structure between words, so that the semantic relationship between words can be more naturally represented.
[0038] Exemplarily, in the grammatical depth implicit representation value calculation subunit 232, the grammatical depth implicit representation value of each speech interaction content hyperbolic space word semantic encoding feature vector in the sequence of the speech interaction content hyperbolic space word semantic encoding feature vector is calculated to obtain the sequence distribution of the grammatical depth implicit representation value of the speech interaction content. Specifically, the process can be expressed as follows: in, for The corresponding hyperbolic space word semantic encoding feature vector of the speech interaction content, It means calculating the square of the vector-norm. For calculation The syntax depth of implies the value, for The corresponding speech interaction content grammar depth implicitly represents the value.
[0039] It should be understood that when the user's voice is converted into text, and these text words are mapped to the hyperbolic space through embedding coding, each word obtains its unique semantic encoding feature vector. However, these vectors alone are not enough to fully understand the structure and meaning of the sentence. Because in natural language, the order of words, parts of speech, and the dependencies between them (i.e., grammatical structure) also determine the meaning of the entire sentence. To compensate for this, the application further calculates the grammatical depth implicit representation value of each word, a step designed to measure the importance and influence of the word in the entire sentence structure.
[0040] Specifically, the implicit representation value of grammatical depth reflects the grammatical status of a word in a sentence and the closeness of its relationship with other words. For example, in a simple sentence "Xiao Ming likes to eat apples", "like" as a verb is the key bridge connecting the subject "Xiao Ming" and the object "apple", so its implicit representation value of grammatical depth may be relatively high. On the contrary, although "apple" is also an important information carrier, because it is in the object position and relies on verbs to express the complete meaning, its implicit representation value of grammatical depth may be slightly lower. In this way, the present application can distinguish the importance of different words in a more detailed way, avoiding the limitations of traditional methods that treat all words equally.
[0041] More importantly, the implicit representation value of grammatical depth helps to reveal long-distance dependencies. In complex sentence structures, although some words are far apart, there may be causal or logical connections between them. For example, when describing a storyline, "At first he was scared, but with the help of his friends, he finally became brave." Although "fear" and "brave" are not in the same phrase, they constitute the key nodes of the transition before and after. By calculating the implicit representation value of grammatical depth, the present application is able to identify such associations in a wider range, ensuring that no important details are missed when the dialogue is generated. In addition, the implicit representation value of grammatical depth also supports context-awareness. Since the grammatical role of each word varies with the context, fixed semantic encoding is difficult to adapt to diverse expressions. By dynamically calculating the implicit representation value of grammatical depth, the present application can flexibly adjust the understanding of vocabulary according to the current dialogue context, and can respond appropriately even when faced with different uses of the same word. This context sensitivity greatly improves the coherence and accuracy of the dialogue.
[0042] That is, by calculating the implicit representation value of grammatical depth, it is possible to more accurately capture the grammatical role of each word in the voice interaction content in the sentence, thereby generating more precise dynamic representations for the words in different contexts. This dynamic representation method helps the model better understand the specific meaning of each interaction content word in a specific context, improves the ability to handle and understand the phenomenon of polysemy, thereby strengthening the semantic expression of voice interaction content and providing a basis for subsequent feedback audio signal generation tasks.
[0043] In one embodiment, the medium-granularity semantic association window determination subunit 233 is used to: extract the grammatical depth implicit representation value of the first target speech interaction content word granularity semantic embedding coding vector in the sequence of the speech interaction content word granularity semantic embedding coding vector as the starting position of the medium-granularity semantic association window; along the sequence distribution of the grammatical depth implicit representation value of the speech interaction content, find the second target speech interaction content word granularity semantic embedding coding vector that is closest to the grammatical depth implicit representation value of the first target speech interaction content word granularity semantic embedding coding vector as the ending position of the medium-granularity semantic association window; based on the starting position of the medium-granularity semantic association window and the ending position of the medium-granularity semantic association window, determine the medium-granularity semantic association window corresponding to the first target speech interaction content word granularity semantic embedding coding vector from the sequence of the speech interaction content word granularity semantic embedding coding vector. Specifically, the process can be expressed by the formula: in, for The corresponding speech interaction content grammar depth implicit representation value, yes The starting position of the corresponding medium-grained semantic association window, for The corresponding speech interaction content grammar depth implicit representation value, is the absolute value, Returns the position corresponding to the minimum value , for The corresponding end position of the medium-granularity semantic association window, for The corresponding medium-granularity semantic association window.
[0044] It should be understood that when the user's voice is converted into text and mapped to the hyperbolic space through embedding coding, each word obtains its unique semantic encoding feature vector. In order to have a deeper understanding of the relationship between sentence structure and vocabulary, this application further calculates the grammatical depth implicit representation value of each word, which is a step designed to measure the importance and influence of the word in the entire sentence. On this basis, by determining the medium-granularity semantic association window, the focus range of the local context can be dynamically adjusted to avoid the information loss or irrelevant information interference caused by the fixed window method.
[0045] First, a key word (i.e., the first target word) is selected as the starting point, and its grammatical depth implicit representation value is extracted. For example, when processing a user's question "I want to know what the weather is like today?", the word "understand" may be the key to understanding the intention of the entire question. Therefore, this application will select "understand" as the first target word and extract its grammatical depth implicit representation value as the starting position of the medium-grained semantic association window. This step ensures that subsequent analysis can focus on the core words and the contextual information around them, rather than blindly covering the entire sentence. Next, along the sequence distribution of the grammatical depth implicit representation value, this application will find another word (i.e., the second target word) closest to the first target word and define it as the end position of the medium-grained semantic association window. This process is similar to finding the shortest path between two points in a continuous space. Continuing with the above example, if "today" is the word that is semantically closest to "understand" and has a similar grammatical status, then "today" will become the end position of the medium-grained semantic association window. Doing so not only limits the boundaries of the analysis, but also ensures that there is a strong semantic connection between the selected words, avoiding the interference of irrelevant words. Finally, based on the start bit and the end bit, the present application can determine the corresponding medium-granularity semantic association window of the first target vocabulary from the original word-granularity semantic embedding encoding vector sequence. In this window, all words are considered to be closely related to the first target vocabulary in the current context. For example, in the sentence "I want to know what the weather is like today?", the determined medium-granularity semantic association window may include "want", "understand", "today", and "weather", which together constitute the core content of expressing the user's intention. In this way, the present application can integrate contextual information in a larger range while maintaining attention to details, so that the dialogue generation is both in line with the actual needs of the user and has good coherence. This dynamic window mechanism is particularly suitable for dealing with complex sentence structures or polysemous words. For example, when facing long sentences or complex sentences, traditional methods may miss important semantic clues due to fixed window size; and when dealing with polysemous phenomena, fixed windows may also lead to misunderstandings. In contrast, the present application can adaptively capture appropriate semantic information in different situations and improve the quality of dialogue generation by flexibly adjusting the scope of the medium-granularity semantic association window.
[0046] That is, the semantic enhancement processing based on grammatical deep association constraints can more flexibly capture the medium-scale semantic associations between words by determining the medium-granularity semantic association window and dynamically adjusting the size of the consideration range of the word-granularity semantics in the speech interaction content. This dynamic window mechanism can not only avoid the problem of semantic information loss caused by the fixed window method when considering the relationship between words, but also effectively reduce the interference of irrelevant information.
[0047] Exemplarily, in the medium-granularity contextual semantic enhancement processing subunit 234, based on the medium-granularity semantic association window of each speech interaction content word granularity semantic embedding coding vector in the sequence of the speech interaction content word granularity semantic embedding coding vector, each speech interaction content word granularity semantic embedding coding vector in the sequence of the speech interaction content word granularity semantic embedding coding vector is subjected to medium-granularity contextual semantic enhancement to obtain a sequence of medium-granularity speech interaction content semantic enhancement word feature vectors as a sequence of medium-granularity speech interaction content semantic enhancement word features. Specifically, the process can be expressed by the formula: in, is the first word in the sequence of the word granularity semantic embedding encoding vector of the speech interaction content. The word-granular semantic embedding encoding vector of the speech interaction content, and denote the linear modulation weight matrix and the linear modulation bias vector respectively, is matrix multiplication, for The position in the sequence, It means to calculate the value of the natural exponential function with the natural constant e as the base. is the first in the sequence of the feature vectors of the semantic reinforcement words in the medium-granularity speech interaction content. A feature vector of semantically enhanced words in medium-granularity speech interaction content.
[0048] That is, based on the determined medium-granularity semantic association window, the optimizer will perform contextual semantic enhancement on the word-granularity semantic embedding encoding vector of each voice interaction content. Through medium-granularity contextual semantic enhancement, the word-granularity context information of the voice interaction content can be integrated on a larger scale, which means that each word now not only represents its own meaning, but also incorporates the influence of surrounding words and long-distance dependency associations, thereby more accurately reflecting its actual meaning in the current context, so that the understood phrase can better express the user's specific intentions.
[0049] Exemplarily, in the speech interaction content context semantic encoding unit 24, the sequence of the medium-granularity speech interaction content semantic reinforcement word features is input into the audio adapter to obtain the speech interaction content context semantic encoding features. In one embodiment, the speech interaction content context semantic encoding unit is used to: input the sequence of the medium-granularity speech interaction content semantic reinforcement word feature vectors into the audio adapter to obtain the speech interaction content context semantic encoding feature vector as the speech interaction content context semantic encoding feature. It is worth mentioning that the role of the audio adapter here is to further refine and strengthen the obtained sequence of medium-granularity speech interaction content semantic reinforcement word feature vectors, and capture the contextual semantic association information between different medium-granularity speech interaction content semantic word feature representations, so as to more accurately understand the user's speech interaction intentions and needs. In addition, the role of the audio adapter also includes adjusting these semantic feature vectors to a form suitable for subsequent large language model processing. This is because different models may have different requirements for input data, such as dimensions, formats, etc. The adapter ensures that the semantic features can be seamlessly connected to the next step of processing.
[0050] In a specific embodiment, the sequence of the medium-granularity speech interaction content semantic reinforcement word feature vectors is input into the audio adapter to obtain the speech interaction content context semantic encoding feature vector as the speech interaction content context semantic encoding feature, including: first, mapping and reducing the dimension of each medium-granularity speech interaction content semantic reinforcement word feature. The goal of this step is to convert the high-dimensional feature vector into a form that is more suitable for subsequent processing. Specifically, the feature dimension can be reduced by linear transformation (such as PCA principal component analysis) or nonlinear transformation (such as t-SNE) while retaining as much important information as possible. In addition, regularization technology can be applied to prevent overfitting and ensure the generalization ability of the model. Through feature mapping and dimensionality reduction, not only the computational complexity is reduced, but also the processing efficiency is improved, making the subsequent context modeling more efficient. In order to capture the potential connection between different feature vectors and construct a more complete context representation, the audio adapter adopts a self-attention mechanism. In this way, the interaction between words can be considered simultaneously in multiple subspaces to enhance the understanding of complex sentence structures. For example, the multi-head self-attention mechanism allows the system to examine the same problem from different angles and improve overall performance. In addition, positional encoding can be introduced to enable the model to perceive the positional relationship of words, which is very important for processing long-distance dependencies. Through the application of the self-attention mechanism, the audio adapter can integrate contextual information in a larger range, avoid overall understanding bias caused by local optimization, and thus significantly improve the quality and coherence of dialogue generation. After the above processing, the audio adapter will fuse all feature vectors into a new sequence, namely the contextual semantic encoding features of the voice interaction content. This new sequence not only retains the important information in the original input, but also incorporates the influence from a wider range of contexts, making the meaning of each word clearer and more coherent. The final output feature vector sequence can be directly passed to the large language model to generate reply features or other forms of responses.
[0051] Exemplarily, in the feedback audio signal generating unit 25, the speech interaction content context semantic coding features are input into an audio decoder based on a large language model to obtain the feedback audio signal. It should be understood that the speech interaction content context semantic coding feature vector processed by the audio adapter contains rich speech interaction content context context information, which is essential for generating a coherent answer that is close to the conversation background. At this time, the speech interaction content context semantic coding feature vector is input into an audio decoder based on a large language model, and a feedback audio signal that is both in line with the current conversation situation and highly relevant can be generated according to the user's speech interaction content context semantics.
[0052] Specifically, the process of the feedback audio signal generating unit outputting feedback audio in real time includes: inputting the semantic coding feature vector of the speech interaction content context into the reply feature extractor based on the large language model to obtain a sequence of reply feature embedding coding vectors; performing a two-fold upsampling process on the sequence of reply feature embedding coding vectors to obtain a sequence of upsampled reply feature embedding coding vectors; passing the sequence of upsampled reply feature embedding coding vectors through a Transformer network based on a two-layer Llama architecture to obtain a sequence of reply feature embedding semantic vectors; using CTC to align the sequence of reply feature embedding semantic vectors with discrete units to obtain a sequence of aligned reply embedding semantic vectors. More specifically, the system achieves alignment between semantic vectors by outputting indefinite-length tokens with padding values of 0; ; Inputting the aligned sequence of restored embedded semantic vectors into a vocoder to synthesize audio to obtain the feedback audio signal.
[0053] More specifically, first, the semantic encoding feature vector of the speech interaction content context containing the semantic features of the speech interaction content context is used to generate and extract features of the reply corpus through a large language model, and the reply features obtained from the large language model (such as the fine-tuned Llama-3.1-8B-Instruct) are converted into a serialized representation, usually an embedding vector. This embedding vector represents the semantic information of the text. Then, in order to enable these embedding vectors to be more finely mapped to the audio features, they are upsampled twice in the time dimension. This means that the original time step of each step is now half of the original, thereby increasing the time resolution and facilitating the subsequent generation of smoother and more natural sound waveforms. Further, the upsampled structure is input into the Transformer layer of the two-layer Llama architecture. That is, the upsampled embedding vector sequence is sent to the Transformer network composed of the two-layer Llama architecture. This process further enhances the temporal dependency and contextual relevance of the text information, so that the embedding at each time point contains more information about the context. The Transformer layer processes the input data through a self-attention mechanism, ensuring that important grammatical and semantic relationships are captured even for longer sentences. The output after the Transformer layer is a new embedding sequence, which then needs to be aligned with the actual target audio features. The discrete units used here are features extracted by applying the HuBERT model to real audio samples and a series of discrete categories obtained by the k-means clustering algorithm. This alignment is achieved using the Connectionist Temporal Classification (CTC) loss function, which allows the prediction results to not have to exactly match the position of the label, but to find the optimal path, that is, a series of actions that are most likely to produce the correct output. This step ensures that the generated audio accurately reflects the content of the reply features in both time and content. Finally, the CTC-aligned embedding sequence is used as a conditional input to the vocoder, a tool specifically designed to reconstruct the original sound waveform from a spectrogram or similar intermediate representation. The vocoder uses its internally trained parameters to generate the corresponding time-domain audio signal according to the input conditions, which is the voice file that can finally be played.
[0054] Exemplarily, in the 3D digital human driving module 30, S3 is executed: using the audio that can be rendered to drive the 3D digital human model, and streaming driving the expression, body movements and lip shape changes of the 3D digital human based on the feedback audio signal. It should be understood that when users interact with 3D digital humans, they expect to get instant and contextual responses, rather than just mechanical text or voice output. With the help of audio driving technology, the 3D digital human can dynamically adjust its facial expressions, body movements and lip shape changes according to the feedback audio signal, making the entire conversation process look more smooth and humane. For example, nodding slightly to show affirmation when answering questions, or using gestures to assist expression when telling stories, these subtle movements can greatly enhance the user's sense of participation and trust. At the same time, this driving method helps to convey emotions and intentions. In addition to the content of the language itself, non-verbal clues (such as facial expressions, body postures, etc.) are also an important part of communication. By accurately controlling the expressions and movements of 3D digital humans, the system can better convey emotional color and tone information, helping users to understand the other party's meaning more accurately. For example, an expression with a smile may imply friendliness and support, while a serious expression may indicate seriousness or warning. In addition, appropriate body language can ease tension and establish better interactive relationships. Furthermore, streaming drive ensures real-time and continuity. Traditional methods often require pre-recording all possible animation clips and selecting them for playback according to specific conditions, which not only increases development costs but also limits the flexibility of the system. In contrast, streaming drive based on feedback audio signals can automatically generate corresponding animation effects during the conversation, without waiting for the entire sentence to end before starting to respond. This means that 3D digital humans can respond to every instruction or question from the user as quickly as real people, providing a more immediate and efficient interactive experience.
[0055] In one embodiment, the 3D digital human driving module is used to: convert the feedback audio signal into PCM code stream data and then input it into the audio-driven 3D digital human model that can be rendered through streaming transmission. It should be understood that in streaming audio interaction, the audio data will be divided into a series of small data packets and continuously transmitted through a pre-agreed transmission channel. The sender will continuously send data packets to the receiver, and the receiver will receive and process these data in a streaming manner. Therefore, by converting the feedback audio signal into PCM code stream data, the feedback audio signal can be divided into a series of small data packets. Such a processing method can ensure the real-time nature of audio interaction, reduce delays and improve transmission efficiency.
[0056] In one embodiment, the 3D digital human driving module is used to: input the feedback audio signal into a streaming audio decoder based on the EmoTalk model for processing to generate a driving signal to perform streaming driving on the expression, body movements and lip shape changes of the 3D digital human.
[0057] Furthermore, the 3D digital human streaming audio interaction system based on the end-to-end speech large model also includes: a model parameter dynamic optimization module, which is used to construct a human feedback learning module to realize the dynamic optimization of the digital human model parameters, and can adjust the response actions of the digital human in real time to ensure the fluency and accuracy of the interaction process. At the same time, as the number of interactions increases, the system gradually adapts to the personalized needs of users, thereby providing more accurate and user-friendly responses in subsequent interactions. In one embodiment, the human feedback learning module scores based on human feedback, and uses the human scores and facial expression driving parameters obtained by the scoring as inputs, and the optimal facial expression driving parameters as outputs, to construct an MLP network for parameter mapping, and to realize hot updates of control parameters based on human feedback.
[0058] In summary, according to the embodiment of the present application, a 3D digital human streaming audio interaction system based on an end-to-end voice big model is explained, which adopts a streaming audio input interface to realize real-time reception of user's voice commands and interactive content, ensuring the continuity and low latency of data transmission, thereby providing an instant response user experience. This allows the 3D digital human to respond to the user's questions or commands as quickly as a real person. In addition, an advanced big model is used for voice feature extraction and semantic analysis, which helps to understand the semantics of the user's voice interaction content more timely and accurately, generate accurate reply features, and convert the reply features into feedback audio signals to realize the streaming drive of the 3D digital human. In this way, not only the realism and immersion of the interaction are improved, but also the intelligence level of the 3D digital human system is enhanced, opening up new possibilities for efficient communication in various application scenarios.
[0059] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0060] It should be understood that the specific examples in this article are only intended to help those skilled in the art better understand the embodiments of the present application, rather than to limit the scope of the embodiments of the present application.
[0061] It should also be understood that in the various embodiments of the present application, the size of the serial number of each process does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0062] It should also be understood that the various implementation methods described in this specification can be implemented individually or in combination, and the embodiments of the present application are not limited to this.
[0063] Unless otherwise stated, all technical and scientific terms used in the embodiments of the present application are the same as the meanings generally understood by those skilled in the art of the technical field of the present application. The terms used in the present application are only for the purpose of describing specific embodiments and are not intended to limit the scope of the present application. The term "and / or" used in the present application includes any and all combinations of one or more related listed items. The singular forms of "a kind of", "above" and "the" used in the embodiments of the present application and the appended claims are also intended to include majority forms, unless the context clearly indicates other meanings. In addition, the terms "first", "second" etc. are only used for description purposes and cannot be understood as indicating or suggesting relative importance.
[0064] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0065] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0066] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0067] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0068] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
Claims
1. A 3D digital human streaming audio interaction system based on an end-to-end speech large model, characterized in that: include: A user audio signal receiving module is used to receive the user's voice interaction signal in real time through a streaming audio input interface; A voice dialogue interaction feedback module, used for inputting the voice interaction signal into a voice dialogue engine based on a large model to obtain a feedback audio signal; The 3D digital human driving module is used to drive the 3D digital human model using audio that can be rendered, and to perform streaming driving of the 3D digital human's expression, body movements and lip shape changes based on the feedback audio signal.
2. The 3D digital human streaming audio interaction system based on the end-to-end speech large model according to claim 1 is characterized in that: The voice dialogue interaction feedback module includes: A speech interaction signal recognition unit, configured to perform speech recognition on the speech interaction signal using an audio encoder to obtain a speech interaction signal encoding result; A speech interaction content embedding coding unit, used for performing word-granularity-based embedding coding on the speech interaction signal coding result to obtain a sequence of speech interaction content word-granularity semantic embedding coding features; The speech interaction content semantic enhancement processing subunit is used to perform semantic enhancement processing on the sequence of the speech interaction content word granularity semantic embedding coding features based on grammatical depth association constraints to obtain a sequence of medium-granularity speech interaction content semantic enhancement word features, It is represented as the aggregation of contextual speech elements after encoding the speech interaction content embedding coding unit: , Where k is the selected number of consecutive frames, and this step can complete the enhanced splicing of k consecutive frames along the feature dimension; A speech interaction content context semantic encoding unit is used to input the sequence of medium-granularity speech interaction content semantic reinforcement word features into the audio adapter to obtain speech interaction content context semantic encoding features, Represents the concatenated array of audio elements processed by the speech interaction content semantic enhancement processing sub-unit: ; The feedback audio signal generating unit is used to input the semantic coding features of the speech interaction content context into an audio decoder based on a large language model to obtain the feedback audio signal, and pass it through a 2-layer CNN perceptron composed of MBConv6 blocks. The activation generates the final speech representation: 。 3. The 3D digital human streaming audio interaction system based on the end-to-end speech large model according to claim 1 is characterized in that: The speech interaction content embedding encoding unit is used to perform word segmentation processing on the speech interaction signal text recognition result and then pass it through a word embedding encoder based on the Bert model to obtain a sequence of speech interaction content word granularity semantic embedding encoding vectors as a sequence of speech interaction content word granularity semantic embedding encoding features.
4. The 3D digital human streaming audio interaction system based on the end-to-end speech large model according to claim 1 is characterized in that: The speech interaction content semantic enhancement processing unit includes: A speech interaction content mapping subunit is used to map each speech interaction content word granularity semantic embedding coding vector in the sequence of the speech interaction content word granularity semantic embedding coding vector to a hyperbolic space to obtain a sequence of speech interaction content hyperbolic space word semantic coding feature vectors, using a context-aware mapping method with variable weights as follows; Given a word-level semantic embedding encoding vector sequence ,in , perform adaptive weight adjustment: , in, is the vector sequence after adaptive adjustment, It is The adaptive weight matrix of words, based on Nonlinear transformation of the network, using The network encoder obtains the context vector To achieve a globally consistent understanding of features: ,in, represents the transformed sequence vector, Use the Poincare sphere model to map the transformed vector to hyperbolic space: in, is the vector mapped to the hyperbolic space; The grammatical depth implicit representation value calculation subunit is used to calculate the grammatical depth implicit representation value of each speech interaction content hyperbolic space word semantic encoding feature vector in the sequence of the speech interaction content hyperbolic space word semantic encoding feature vector to obtain the sequence distribution of the grammatical depth implicit representation value of the speech interaction content, including: calculating each word vector Global Grammar Center The hyperbolic distance of , as the basic measure of grammatical depth, the global grammatical center By sequence The Fréchet mean of is calculated as: in is the Poincare distance in hyperbolic space: The implicit representation value of the grammatical depth is calculated. The implicit representation value of the grammatical depth is Defined as word vector Global Grammar Center The hyperbolic distance is normalized by the Sigmoid function: in is the Sigmoid function; h, Representing the word vector in the hyperbolic space, we finally get the sequence distribution of the grammatical depth implicit representation value of the speech interaction content; A medium-granularity semantic association window determination subunit is used to determine the medium-granularity semantic association window of each speech interaction content word granularity semantic embedding coding vector in the sequence of the speech interaction content word granularity semantic embedding coding vector, that is, the value of the selected continuous frame number k, based on the sequence distribution of the speech interaction content grammatical depth implicit representation value; The medium-granularity contextual semantic reinforcement processing subunit is used to perform medium-granularity contextual semantic reinforcement on each speech interaction content word granularity semantic embedding coding vector in the sequence of speech interaction content word granularity semantic embedding coding vectors based on the medium-granularity semantic association window of each speech interaction content word granularity semantic embedding coding vector in the sequence of speech interaction content word granularity semantic embedding coding vectors to obtain a sequence of medium-granularity speech interaction content semantic reinforcement word feature vectors as a sequence of medium-granularity speech interaction content semantic reinforcement word features.
5. The 3D digital human streaming audio interaction system based on the end-to-end speech large model according to claim 4 is characterized in that: The medium-granularity semantic association window determination subunit includes: Extracting the grammatical depth implicit representation value of the first target speech interaction content word granularity semantic embedding coding vector in the sequence of the speech interaction content word granularity semantic embedding coding vectors as the starting position of the medium-granularity semantic association window; Along the sequence distribution of the grammatical depth implicit representation value of the speech interaction content, find the second target speech interaction content word granularity semantic embedding coding vector that is closest to the grammatical depth implicit representation value of the first target speech interaction content word granularity semantic embedding coding vector as the termination position of the medium granularity semantic association window; Based on the starting position of the medium-granularity semantic association window and the ending position of the medium-granularity semantic association window, the medium-granularity semantic association window corresponding to the first target speech interaction content word granularity semantic embedding coding vector is determined from the sequence of the speech interaction content word granularity semantic embedding coding vectors.
6. The 3D digital human streaming audio interaction system based on end-to-end speech large model according to claim 1 is characterized in that: The speech interaction content context semantic encoding unit is used to input the sequence of medium-granularity speech interaction content semantic reinforcement word feature vectors into the audio adapter to obtain a speech interaction content context semantic encoding feature vector as the speech interaction content context semantic encoding feature.
7. The 3D digital human streaming audio interaction system based on the end-to-end speech large model according to claim 1 is characterized in that: The audio decoder consists of two layers of Transformer layers of the Llama architecture.
8. The 3D digital human streaming audio interaction system based on the end-to-end speech large model according to claim 1 is characterized in that: The feedback audio signal generating unit comprises: Inputting the semantic encoding feature vector of the speech interaction content context into a reply feature extractor based on a large language model to obtain a sequence of reply feature embedding encoding vectors; Performing a two-fold upsampling process on the sequence of recovery feature embedding coding vectors to obtain a sequence of upsampled recovery feature embedding coding vectors; Passing the sequence of upsampled recovery feature embedding encoding vectors through a Transformer network based on a two-layer Llama architecture to obtain a sequence of recovery feature embedding semantic vectors; Using CTC, aligning the sequence of the response feature embedded semantic vectors with the discrete units to obtain a sequence of aligned response embedded semantic vectors; The aligned sequence of restored embedded semantic vectors is input into a vocoder to synthesize audio to obtain the feedback audio signal.
9. The 3D digital human streaming audio interaction system based on end-to-end speech large model according to claim 1 is characterized in that: The 3D digital human driving module is used to convert the feedback audio signal into PCM code stream data and then input it into the audio-driven 3D digital human model that can be rendered through streaming transmission.
10. The 3D digital human streaming audio interaction system based on end-to-end speech large model according to claim 1, characterized in that: The 3D digital human driving module is used to input the feedback audio signal into a streaming audio decoder based on the EmoTalk model for processing to generate a driving signal to perform streaming driving on the expression, body movements and lip shape changes of the 3D digital human.
Citation Information
Patent Citations
Multi-modal driving algorithm for real-time generation of 3D digital human limb movements
CN118015157A
Digital people stream type multi-modal interaction method based on large language model
CN118227746A
Cited By
Digital human question and answer system
CN121658531A
A digital human question answering system
CN121658531B