Digital life individuation implementation method, device, equipment and medium

By acquiring, preprocessing and emotional analysis of user interaction information, combining long-term memory databases and knowledge graphs, personalized replies are generated and digital human rendering technology is used to solve the shortcomings of emotional interaction and cultural heritage in the personalized realization of digital life, deep emotional interaction and long-term memory storage are achieved, and vivid interactive experience is provided.

CN120353929AInactive Publication Date: 2025-07-22HANGZHOU LINGWA TECHNOLOGY CO LTD
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510411424.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-22
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing personalized implementation methods of digital life cannot deeply understand users' emotional needs in emotional interaction, lack systematic integration of user knowledge graphs and long-term memory, and cannot meet the needs of emotional companionship and cultural heritage.

Method used

By obtaining user interaction information, preprocessing and sentiment analysis, combining long-term memory database and user knowledge graph, personalized replies are generated, and interactive videos are generated using digital human rendering technology to achieve deep emotional interaction and long-term memory storage.

Benefits of technology

It provides a more intuitive and vivid interactive experience, can deeply understand user emotions, realize emotional companionship and cultural heritage, and enhance the authenticity and immersion of interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120353929A_ABST
    Figure CN120353929A_ABST
Patent Text Reader

Abstract

The invention relates to a digital life individuation implementation method and device, equipment and a medium. The method comprises the steps of obtaining user interaction information in response to an obtained user input instruction; preprocessing the user interaction information according to a set mode to obtain a user interaction text; inputting the user interaction text into the digital life entity to obtain a personalized reply; the personalized reply comprises a target person and audio data; and generating an interactive video corresponding to the personalized reply based on digital human rendering. By adopting the method, personalized digital cloned construction, cross-space-time emotional interaction and cultural heritage digital inheritance can be realized through emotional analysis, behavior pattern learning and the like, and the method is suitable for scenes of emotional science and technology, digital immortality, family archive management and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of artificial intelligence voice interaction, and particularly relates to a method, device, equipment and medium for realizing digital life personalization. Background Art

[0002] With the development of digital technology, technologies such as digital avatars and virtual assistants have emerged. Through certain intelligent interaction capabilities, functions such as simple voice command recognition and information query can be realized, forming a preliminary way to realize digital life personalization mainly based on virtual images and preset dialogue modes, which are widely used in simple customer service scenarios, entertainment interactions and other fields.

[0003] In traditional technologies, when constructing digital life personalization functions, dialogue interaction is mainly achieved by presetting a large number of fixed conversation templates. For example, in a customer service scenario, corresponding answer options are set for common questions. When a user asks a question, the system selects a suitable reply according to keyword matching. In terms of virtual image display, static or simple animated virtual character models are usually used, and their actions and expressions have limited changes.

[0004] However, there are many problems in the current traditional ways to realize digital life personalization. In terms of emotional interaction, it is impossible to deeply understand the emotional needs of users and it is difficult to provide real emotional companionship. In terms of information integration and long-term memory, there is a lack of systematic integration of user knowledge graphs, behavior patterns and long-term memories. In terms of cultural inheritance, traditional methods lack structured storage and interactive inheritance means for family memories, oral history, etc., and cannot meet people's needs for digital preservation and intergenerational inheritance of family culture. Summary of the Invention

[0005] Based on this, in view of the above technical problems, it is necessary to provide a method, device, equipment and medium for realizing digital life personalization that can achieve in-depth emotional interaction and long-term memory storage.

[0006] In a first aspect, the present application provides a method for realizing digital life personalization, including:

[0007] Responding to the acquired user input instruction, and acquiring user interaction information;

[0008] Preprocessing the user interaction information according to a set mode to obtain a user interaction text;

[0009] Inputting the user interaction text into a digital life body to obtain a personalized reply; the personalized reply includes a target person and audio data;

[0010] Generating an interaction video corresponding to the personalized reply based on digital human rendering.

[0011] In one embodiment, the user interaction information is preprocessed according to a set mode to obtain a user interaction text, including:

[0012] Identify the data type of the user interaction information to obtain a classification result; the classification result includes a voice data type and a text data type;

[0013] If the classification result is the voice data type, process the user interaction information according to the voice processing mode to obtain a user interaction text;

[0014] If the classification result is the text data type, perform natural language processing on the user interaction information to obtain a user interaction text;

[0015] Among them, the voice processing mode corresponds to the following operation steps:

[0016] Perform noise suppression and invalid audio filtering on the user interaction information according to the voice activity detection based on the memory network to obtain normalized voice data;

[0017] Use end-to-end automatic speech recognition to convert the normalized voice data into text to obtain text data;

[0018] Perform error correction on the text data based on the sequence processing method to obtain a user interaction text.

[0019] In one embodiment, the digital life form obtains a personalized response through the following method:

[0020] Perform sentiment analysis on the user interaction text to obtain a sentiment analysis result; the sentiment analysis result includes a sentiment label and a target person;

[0021] Index according to the time stamp in the memory repository based on the user interaction text and the sentiment analysis result to obtain a personalized memory; the memory repository includes a long-term memory repository and a user knowledge graph;

[0022] Perform sentiment enhancement on the user interaction text and the personalized memory based on the behavior pattern database to obtain a personalized text response;

[0023] Perform voice cloning on the personalized text response according to the pre-stored voiceprint data of the target person to obtain a personalized response.

[0024] In one embodiment, indexing according to the time stamp in the memory repository based on the user interaction text and the sentiment analysis result to obtain a personalized memory, including:

[0025] Calculate a semantic vector based on the user interaction text and the sentiment analysis result;

[0026] Calculate the similarity between the semantic vector and the sentiment feature vector in the long-term memory repository to obtain a semantic similarity;

[0027] Determine the historical memories in the long-term memory library corresponding to the semantic similarity exceeding the set threshold as historical conversation data based on the timestamp index;

[0028] Search for the knowledge graph corresponding to the target person in the user knowledge graph based on the timestamp index to obtain historical interest data;

[0029] Obtain personalized memories based on the historical conversation data and historical interest data;

[0030] Obtain the semantic similarity through the following formula:

[0031]

[0032] where Q is the semantic vector; D is the emotional feature vector; n is the number of sub-parts into which the semantic vector and the emotional feature vector are divided; Q i , D i respectively correspond to the i-th sub-part of the semantic vector and the emotional feature vector; α i is the weight corresponding to each sub-part.

[0033] In one embodiment, an interactive video corresponding to a personalized response is generated based on digital human rendering, including:

[0034] Synchronize the virtual image with the audio data based on the phoneme interpolation algorithm to obtain a video stream; the virtual image corresponds to the target person;

[0035] Combine and render the audio data and the video stream to obtain an interactive video.

[0036] In one embodiment, a digital life form is constructed by the following method:

[0037] Obtain user personalized data; the user personalized data includes voice interaction data, text interaction data, and behavior pattern data;

[0038] Preprocess the user personalized data to obtain user knowledge entries;

[0039] Perform emotional classification on the user knowledge entries based on an emotional classification model and construct a long-term memory library in the order of timestamp index; the long-term memory library contains emotional feature vectors;

[0040] Extract the entities and corresponding relationships in the user knowledge entries based on a graph convolutional network and an attention mechanism, and record them in the order of timestamp index to obtain a user knowledge graph;

[0041] Use a long short-term memory network to analyze the user interaction habits of the user knowledge entries to obtain a behavior pattern database; the behavior pattern database includes a personalized conversation style;

[0042] Use a generative adversarial network and a 256-dimensional voiceprint vector to perform individual timbre voice cloning on user knowledge entries to obtain a voiceprint data set corresponding to the person.

[0043] Obtain a digital life form based on the long-term memory library, the user knowledge graph, and the voiceprint data set corresponding to the person.

[0044] In one embodiment, the method further includes:

[0045] Construct new interaction data from the user interaction information and personalized responses corresponding to each user input instruction obtained in response.

[0046] Annotate the new interaction data with a timestamp and store it in the long-term memory library.

[0047] Use a graph convolutional network and an attention mechanism to extract new entities and corresponding relationships from the new interaction data, and update the user knowledge graph.

[0048] Based on a long short-term memory network, optimize the personalized conversation style according to the new interaction data.

[0049] Distributedly store the long-term memory library and the user knowledge graph.

[0050] In a second aspect, the present application also provides a digital life personalization implementation device, including:

[0051] A data acquisition module, configured to acquire user interaction information in response to a user input instruction obtained.

[0052] A data processing module, configured to preprocess the user interaction information in a set mode to obtain a user interaction text.

[0053] A personalization module, configured to input the user interaction text into the digital life form to obtain a personalized response; the personalized response includes a target person and audio data.

[0054] An interaction module, configured to generate an interaction video corresponding to the personalized response based on digital human rendering.

[0055] In a third aspect, the present application also provides a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the steps of any of the above digital life personalization implementation methods are implemented.

[0056] In a fourth aspect, the present application also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of any of the above digital life personalization implementation methods are implemented.

[0057] The above digital life personalization implementation method, device, computer device, and storage medium support various types of user interaction information input, whether it is voice, text, or other forms. After being converted into standard user interaction text through a unified preprocessing process, the versatility and adaptability of the method are enhanced. The user interaction text is input into the digital life body, and a personalized response containing the target person and audio data is obtained based on the integration of the user knowledge graph, behavior pattern, and long-term memory. An interactive video corresponding to the personalized response is generated based on the digital human rendering, providing a more intuitive and vivid interaction experience for users. Brief Description of the Drawings

[0058] To more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments or related technologies. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0059] Figure 1 It is a flowchart of the digital life personalization implementation method of the present invention;

[0060] Figure 2 It is a sub-step flowchart of step S103;

[0061] Figure 3 It is a component structure diagram of the digital life personalization implementation device of the present invention. Detailed Embodiments

[0062] In order to make the objectives, technical solutions, and advantages of the present application clearer, the following further details the present application in conjunction with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0063] In one embodiment, as Figure 1 shown, a digital life personalization implementation method is provided. In this embodiment, the method is exemplified by being applied to a terminal. It can be understood that the method can also be applied to a server, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. The decoupling (response time < 50ms) through RESTful API (Representational State Transfer Application Programming Interface) supports the hybrid deployment of the terminal and the server. In this embodiment, the method includes the following steps:

[0064] S101. In response to the obtained user input instruction, obtain user interaction information.

[0065] The user inputs user interaction information through means such as voice and text. By monitoring the user in real time and obtaining the user input instruction, the user interaction information is parsed. Illustratively, the user can input user interaction information through means such as the device microphone and manual input in the text input box. The user interaction information includes voice signals, text, etc. Optionally, in combination with historical interaction records, it is determined whether the user interaction information belongs to a continuous conversation, and the user interaction information is stored in the short-term conversation cache to ensure the context coherence of subsequent personalized responses.

[0066] S102. Preprocess the user interaction information in a set mode to obtain user interaction text.

[0067] Illustratively, perform processing such as cleaning and conversion on the obtained user interaction information. Exemplarily, for voice information, use speech recognition technology to convert it into text, while removing background noise and correcting mispronunciations; for text information, perform operations such as word segmentation, stop word removal, and standardization to organize the original interaction information into a standardized and concise user interaction text.

[0068] S103. Input the user interaction text into the digital life form to obtain a personalized response; the personalized response includes the target character and audio data.

[0069] Input the preprocessed user interaction text into the digital life form. Illustratively, the digital life form is a user personalization model constructed based on the user's historical inputs, and it integrates a variety of technologies internally, including sentiment analysis models, long-term memory libraries, user knowledge graphs, behavior pattern learning, etc. The digital life form generates a personalized response based on the content of the user's interaction text, combined with its understanding of the user's interests, emotional states, past interaction records, etc. Among them, the target character information determines the role style of the response, and the audio data is the voice presentation form of the final response.

[0070] S104. Generate an interaction video corresponding to the personalized response based on digital human rendering.

[0071] Illustratively, based on digital human rendering technology, combine the audio data in the personalized response with the image of the target character. According to the rhythm, intonation, and semantic content of the audio, drive the facial expressions, lip movements, etc. of the digital human to generate a smooth and natural interaction video, so as to intuitively display the process of the digital human responding in a specific role style and provide a more vivid and immersive interaction experience for the user.

[0072] In the above method for realizing digital life personalization, user personalization data is comprehensively collected and analyzed. By deeply understanding the user's identity, preferences, emotional characteristics, and interaction habits, the emotion classification model captures the user's emotional state in real time. The long-term memory library stores historical emotional data, and adjusts the response style according to the user's emotional changes, offering comfort when the user is sad and sharing joy when the user is happy, achieving emotional resonance, meeting the need for emotional companionship, making the interaction warmer and more considerate, shortening the distance between the user and the digital life form. Through the user knowledge graph, entities and relationships are accurately extracted, covering knowledge in multiple fields such as family relationships and hobbies, providing rich knowledge support for the digital life form, enabling it to accurately understand the user's intention in the conversation and provide professional and targeted answers. The intonation, speech rate, and pauses are adjusted in combination with the emotional characteristics to make the synthesized voice natural and fluent. In the cross-time and space conversation scenario, the user hears a familiar voice, enhancing the authenticity and immersion of the interaction and improving the emotional experience.

[0073] In one embodiment, the user interaction information is preprocessed according to a set mode to obtain a user interaction text, including:

[0074] S21. Identify the data type of the user interaction information to obtain a classification result; the classification result includes a voice data type and a text data type.

[0075] Voice data usually exists in the form of audio files or real-time voice streams, while text data is directly presented in text form. Schematically, it is judged according to the data format of the user interaction information. Generally, voice data is transmitted as audio data, while text input is directly transmitted as text data. Voice and text data usually carry different meta-information during transmission, including HTTP request headers, data formats, etc. The data type can be quickly classified by parsing the meta-information. Optionally, for user interaction information whose meta-information cannot be recognized through encryption, a data type recognition algorithm can be used to analyze the obtained user interaction information. The data type recognition algorithm is based on the multi-modal feature extraction technology of deep learning and comprehensively judges the input information from multiple aspects such as signal characteristics, frequency distribution, and syntactic structure. For voice data, the algorithm extracts the time-domain and frequency-domain characteristics of the audio signal, such as fundamental frequency, formants, etc.; for text data, it focuses on the analysis of vocabulary, grammar, and syntactic structure. In this way, the data type of the user interaction information is accurately identified, and it is determined as a voice data type or a text data type.

[0076] S22. If the classification result is a voice data type, process the user interaction information according to the voice processing mode to obtain a user interaction text.

[0077] Among them, the voice processing mode corresponds to the following operation steps:

[0078] S221. Perform noise suppression and invalid audio filtering on the user interaction information according to the voice activity detection based on the memory network to obtain standardized voice data.

[0079] Schematically, when the recognition result is of the voice data type, the voice activity detection based on the memory network (FSMN-VAD) technology is used to process the user interaction information. The memory network can store and utilize the historical voice data information. In a complex audio environment, by analyzing the short-term and long-term characteristics of the voice signal, it can distinguish the voice part and the noise part. Specifically, according to the preset noise and voice activity thresholds, the input voice signal is analyzed frame by frame to suppress background noises such as environmental noise and device noise, and filter out invalid audio such as silence and short non-speech audio segments, thereby obtaining standardized voice data.

[0080] S222. Use end-to-end automatic speech recognition to convert the standardized voice data into text to obtain text data.

[0081] Furthermore, use the end-to-end automatic speech recognition (ASR) technology to convert the standardized voice data into text. The end-to-end ASR model is usually based on a deep neural network architecture. Exemplarily, a combination of a convolutional neural network (CNN) and a recurrent neural network (RNN) is adopted. The CNN is used to extract the local features of the voice signal, and the RNN is responsible for capturing the temporal information of the voice. By layer-by-layer learning and mapping of the acoustic features of the voice signal, the voice signal is directly converted into the corresponding text sequence to obtain the preliminary text data.

[0082] S223. Perform error correction on the text data based on the sequence processing method to obtain the user interaction text.

[0083] Schematically, the CTC / Attention mechanism is used to correct the text data obtained by end-to-end automatic speech recognition to improve the recognition accuracy. Among them, CTC (Connectionist Temporal Classification) can directly calculate the probability distribution from the speech feature sequence to the label sequence without prior knowledge of the time alignment relationship between the speech signal and the corresponding text label. Specifically, by introducing a "blank" label, it can handle situations such as repeated labels and silent segments that may exist in the speech signal. Further, in the decoding stage, CTC will, based on the probability distribution obtained through training, find the most likely label sequence through algorithms such as Beam Search. Due to the complexity and diversity of speech signals, there may be certain errors in the directly decoded results. CTC can evaluate the reliability of the decoding results by calculating the probability scores of different paths. For paths with lower scores or some segments that do not conform to language rules in the decoding results, the CTC mechanism can mark them as areas that may contain errors.

[0084] Take the preliminary text result decoded by CTC as the input of the Attention mechanism. The Attention mechanism can dynamically allocate weights, enabling the model to focus on the key information related to the current output when processing the input sequence. In the speech recognition error correction task, the Attention mechanism can search for relevant important information in the input speech feature sequence or the previous decoding results according to the text segment that needs to be corrected currently. Specifically, the Attention mechanism can make full use of the context information of the text for error correction, consider the entire text sequence globally, analyze the semantic and syntactic relationships between words, and thus discover some recognition errors caused by speech similarity or grammar errors. Output this corrected text sequence as the final user interaction text.

[0085] S23. If the classification result is of the text data type, perform natural language processing on the user interaction information to obtain the user interaction text.

[0086] When the classification result is of the text data type, perform natural language processing (NLP) on the user interaction information. Specifically, perform lexical analysis to segment the text into words or morphemes and label the part of speech; perform syntactic analysis to construct the syntactic structure tree of the sentence to understand the components of the sentence and their relationships; perform semantic analysis to eliminate ambiguity and extract the true semantics of the text through multi-level parsing of lexical semantics, sentence semantics, and discourse semantics. In this process, techniques such as named entity recognition and anaphora resolution can also be used to identify and process the entities and anaphoric relationships in the text, thereby obtaining the user interaction text that has been deeply understood and processed.

[0087] In one embodiment, as Figure 2 shown, the digital life form obtains a personalized response through the following method:

[0088] S201. Perform sentiment analysis on the user interaction text to obtain a sentiment analysis result; the sentiment analysis result includes a sentiment label and a target person.

[0089] Process the user interaction text based on a sentiment classification model that combines Bi-LSTM (Bidirectional Long Short-Term Memory) and Transformer (attention mechanism). Among them, Bi-LSTM can learn the text sequence from both the forward and reverse directions, effectively capturing the long-distance dependencies in the text. Transformer, through the multi-head attention mechanism, can concurrently focus on the information at different positions in the text, comprehensively analyze the vocabulary, phrases, and sentence structures in the text. In particular, the F1-score (a binary classification metric) of the sentiment label ≥ 0.91.

[0090] The sentiment classification model is trained on a large text dataset with sentiment annotations to learn the sentiment feature patterns expressed by different vocabulary, grammatical structures, and semantic combinations. When applied, the input is the user interaction text. The model converts each word or phrase in the text into a vector representation, and then through multiple layers of calculations of Bi-LSTM and Transformer, extracts the sentiment features of the text. Finally, the model outputs the corresponding sentiment label, such as 12 emotion labels including happy, sad, angry, nostalgic, etc., and at the same time uses named entity recognition technology to extract information related to the target person from the text, determines the target person, and completes the generation of the sentiment analysis result.

[0091] S202. Index in the memory repository according to the time stamp based on the user interaction text and the sentiment analysis result to obtain personalized memories; the memory repository includes a long-term memory repository and a user knowledge graph.

[0092] Use the keywords in the user interaction text, the sentiment label and the target person in the sentiment analysis result as retrieval conditions, and search in the memory repository using the time stamp indexing mechanism. Schematically, in the long-term memory repository, each interaction record is marked with a time stamp, and the interaction record contains the user's past interaction content, sentiment state, and related event information. Exemplarily, the sentiment label is "nostalgic" and the target person is "grandma", that is, according to the time stamp, find the historical interaction records related to grandma and with a nostalgic sentiment, which may include the conversation content with grandma and the special times spent together, etc.

[0093] Schematically, the user knowledge graph stores structured knowledge such as various attributes, hobbies, and interpersonal relationships of the user. According to the association relationships of the target person in the knowledge graph, the retrieval scope is further expanded to obtain more comprehensive information related to the target person, such as the life story of the grandmother, family stories, etc. The information retrieved from the long-term memory library and the user knowledge graph is integrated to form personalized memories.

[0094] S203. Perform sentiment enhancement on the user interaction text and personalized memories based on the behavior pattern database to obtain a personalized text response.

[0095] The behavior pattern database stores behavior pattern information of the user in different emotional states and interaction scenarios, including language style, commonly used vocabulary, expression habits, etc. The user interaction text and personalized memories are input into a sentiment classification model combined with Bi-LSTM and Transformer, that is, a sentiment enhancement model based on the LLM (Large Language Model), which analyzes the sentiment tendency and key content in the user interaction text and personalized memories, and searches for similar emotional scenarios and interaction patterns in the behavior pattern database. According to the found behavior patterns, the module supplements and adjusts the original user interaction text, and can add some modal particles, rhetorical devices that conform to the user's emotional state in the response, or adjust the sentence structure to make the response better reflect the user's current emotional depth and personality characteristics, and generate a personalized text response that better meets the user's emotional needs.

[0096] S204. Perform voice cloning on the personalized text response according to the pre-stored voiceprint data of the target person to obtain a personalized response.

[0097] Schematically, a 256-dimensional voiceprint embedding combined with GAN (Generative Adversarial Network) training is used to achieve voice cloning, that is, the personalized text response is used to generate audio data with the unique voice characteristics of the target person, such as tone, pitch, speech rate, intonation, etc. Among them, the pre-stored voiceprint data of the target person is obtained by extracting features from a large number of voice samples of the target person.

[0098] Specifically, the 256-dimensional voiceprint embedding vector can accurately represent the voiceprint characteristics of the target person. During the GAN training process, the generator combining 4-layer CNN and GRU attempts to generate speech similar to the voiceprint of the target person, and the discriminator makes multi-scale spectral contrast judgments on the generated speech to distinguish whether it is the real speech of the target person or the forged speech generated by the generator. Through continuous adversarial training of the generator and the discriminator, the generator can gradually generate samples closer to the real speech of the target person. For personalized text responses, according to the semantics, emotions, and grammatical structures of the text, combined with the voiceprint characteristics of the target person, parameters such as the prosody, rhythm, intonation, speech rate, and pause of the speech are adjusted.

[0099] In one embodiment, according to the user interaction text and the result of sentiment analysis, indexing in the memory repository by timestamp to obtain personalized memories, including:

[0100] S31. Calculate a semantic vector according to the user interaction text and the result of sentiment analysis.

[0101] Illustratively, use natural language processing technology to deeply understand the semantics of the user interaction text, and further consider the influence of emotional factors on semantics in combination with the result of sentiment analysis. Exemplarily, the text information after semantic understanding and emotion weighting processing is encoded into a low-dimensional semantic vector through a deep learning model based on the Transformer architecture, which contains the semantic content of the user interaction text and the included emotional information.

[0102] S32. Calculate the similarity between the semantic vector and the emotional feature vector in the long-term memory repository to obtain the semantic similarity.

[0103] A large number of emotional feature vectors are stored in the long-term memory repository. The emotional feature vectors are obtained by performing sentiment analysis and feature extraction on the user's interaction information such as speech and text during the user's previous interactions. Each emotional feature vector represents the emotional state and related semantic information of the user at a specific moment.

[0104] The semantic similarity is obtained through the following formula:

[0105]

[0106] where Q is the semantic vector; D is the emotional feature vector; n is the number of sub-parts into which the semantic vector and the emotional feature vector are divided; Q i ,D i correspond to the i-th sub-part of the semantic vector and the emotional feature vector respectively; α i is the weight corresponding to each sub-part.

[0107] S33. Determine the historical memories in the long-term memory library corresponding to the semantic similarity exceeding the set threshold as historical conversation data based on the timestamp index.

[0108] When the calculated semantic similarity exceeds the set threshold, it indicates that the current user interaction text has a high similarity with some historical emotional information in the long-term memory library. Using the timestamp index mechanism, quickly locate the memory records at a specific time point, and search for the historical memories corresponding to the high-similarity emotional feature vectors in the long-term memory library. The historical memories contain information such as the user's past conversation content, interaction scenarios, and emotional states. Exemplarily, the similarity threshold is 90%.

[0109] S34. Search for the knowledge graph corresponding to the target person in the user knowledge graph based on the timestamp index to obtain historical interest data.

[0110] The user knowledge graph stores various attributes, interests, interpersonal relationships, and knowledge information related to different people of the user. After determining the target person, search for information related to the target person in the user knowledge graph based on the timestamp index. The timestamp index can help quickly locate the knowledge information related to the target person at different time points, including the life stories, interests, and common experiences with the user of the target person, which is called historical interest data.

[0111] S35. Obtain personalized memories according to the historical conversation data and historical interest data.

[0112] Integrate the historical conversation data and historical interest data to construct personalized memories. Exemplarily, the past conversation content between the user and the grandfather, such as the conversation about opera, can be combined with the historical interest data that the grandfather likes opera to form a more complete and personalized memory. This personalized memory not only contains the specific conversation content but also incorporates background information such as the interests of the target person, which can more comprehensively reflect the relationship and interaction history between the user and the target person and provide strong support for the subsequent generation of personalized responses.

[0113] In one of the embodiments, an interactive video corresponding to the personalized response is generated based on the digital human rendering, including:

[0114] S41. Synchronize the virtual image with the audio data based on the phoneme interpolation algorithm to obtain a video stream; the virtual image corresponds to the target person.

[0115] The phoneme interpolation algorithm is the core technology for synchronizing the virtual image with the audio data. Schematically, using speech recognition technology and acoustic models, the audio is parsed into a series of phoneme sequences. A phoneme is the smallest distinguishable unit in speech, and each phoneme corresponds to a specific pronunciation action and mouth shape change.

[0116] The virtual character has a pre - set lip - sync action library, which contains the standard lip - shapes corresponding to various phonemes and the key frames of lip - shape changes. During the synchronization process, the phoneme interpolation algorithm will select the corresponding lip - shape key frames in the lip - sync action library of the virtual character according to the duration and sequence of each phoneme in the audio.

[0117] Furthermore, there may be some slight variations and transitions in the pronunciation of the actual audio, which do not exactly correspond to the standard lip - shapes one by one. The phoneme interpolation algorithm generates intermediate transitional lip - shapes through interpolation calculations between key frames, making the lip - shape changes more natural and smooth.

[0118] Optionally, after calculating the lip - shape of each frame, combined with the facial bone animation and expression system of the virtual character, drive the virtual character to make corresponding actions and expressions. The facial bone animation of the virtual character will be adjusted according to the lip - shape changes to ensure that the movements of parts such as the lips and cheeks conform to the pronunciation actions. At the same time, the expression system will adjust the expressions of parts such as the virtual character's eyes and eyebrows according to the emotional information and conversation content in the audio, making the performance of the virtual character more vivid.

[0119] Combining the continuously changing virtual character frames at a certain frame rate forms a video stream synchronized with the audio data.

[0120] S42: Combine and render the audio data with the video stream to obtain an interactive video.

[0121] Schematically, the rendering engine will load each frame image and the corresponding audio data in the video stream. For the images in the video stream, the rendering engine can further process them according to factors such as the material of the virtual character, lighting conditions, and scene layout. Furthermore, the rendering engine precisely matches the audio data with the video stream in terms of time to ensure that the playback of the audio is completely synchronized with the lip - shapes and actions of the virtual character in the video, avoiding the situation of inconsistent sound and picture.

[0122] After the above - mentioned series of processes, the rendering engine fuses the video stream and the audio data together to generate the final interactive video.

[0123] In one of the embodiments, the digital life form is constructed by the following method:

[0124] S51: Obtain user - personalized data; the user - personalized data includes voice interaction data, text interaction data, and behavior pattern data.

[0125] User personalized data is the data input by users when constructing digital life forms, including continuously collecting information such as voiceprint, intonation, speech rate, and common expression patterns during user voice interactions; information such as the user's past chat content, various events mentioned, and preferred topics; continuously monitoring and analyzing a large number of user interaction behaviors to summarize behavioral characteristics such as the user's common conversation style, common ways of asking questions, and interaction habits.

[0126] S52. Preprocess the user personalized data to obtain user knowledge entries.

[0127] After obtaining the user personalized data, to ensure the accuracy and efficiency of subsequent analysis and applications, the original data is transformed into high-quality, structured user knowledge entries.

[0128] S53. Perform sentiment classification on the user knowledge entries based on a sentiment classification model and construct a long-term memory library in the order of timestamp indexing; the long-term memory library contains sentiment feature vectors.

[0129] Exemplarily, based on the CNN-RNN sentiment classification model, for each user knowledge entry, the model identifies the contained sentiment tendency, such as happy, sad, angry, surprised, etc., and assigns a corresponding sentiment label to obtain a sentiment feature vector. The user knowledge entries with sentiment feature vectors are stored in the long-term memory library in the order of timestamp indexing.

[0130] S54. Extract entities and corresponding relationships in the user knowledge entries based on a graph convolutional network and an attention mechanism, and record them in the order of timestamp indexing to obtain a user knowledge graph.

[0131] By combining a graph convolutional network (GCN) and an attention mechanism, various entities such as people, places, events, concepts, etc. are identified from the user knowledge entries, and their relationships such as ownership, causal relationship, time relationship, etc. are determined. Then, the entities and relationships are recorded in the user knowledge graph in the order of timestamp indexing. The user knowledge graph shows the user's knowledge system and cognitive structure in a structured manner. As new user knowledge entries are continuously added, the knowledge graph is also continuously updated and improved, providing strong support for understanding the user's interests, intentions, and knowledge background.

[0132] S55. Use a long short-term memory network to analyze the user interaction habits of the user knowledge entries to obtain a behavior pattern database; the behavior pattern database includes personalized conversation styles.

[0133] Schematically, user knowledge entries are input into an LSTM (Long Short-Term Memory) network in chronological order. Through learning from historical interaction data, the network analyzes the user's interaction behavior patterns in different scenarios, including language expression habits such as common vocabulary, sentence patterns, and topic preferences, as well as user behavior patterns such as question frequency, response time interval, and attention to different types of information. The personalized conversation style in the behavior pattern database not only reflects the user's language habits but also includes the user's behavior tendencies in different emotional states and interaction scenarios.

[0134] S56. Use a generative adversarial network and a 256-dimensional voiceprint vector to perform individual voice cloning of user knowledge entries to obtain a set of voiceprint data for the corresponding person.

[0135] The generative adversarial network and the 256-dimensional voiceprint vector are an accurate representation of the unique voice characteristics of the user or the target person, containing information in multiple aspects such as timbre, pitch, and speech rate.

[0136] S57. Obtain a digital life form based on the long-term memory library, the user knowledge graph, and the set of voiceprint data for the corresponding person.

[0137] The digital life form is a virtual entity that synthesizes various aspects of user information. The long-term memory library, the user knowledge graph, and the set of voiceprint data are the key elements for constructing the digital life form. The long-term memory library provides the user's emotional history and interaction records for the digital life form, enabling it to understand the user's emotional changes and give more considerate responses in interactions. The user knowledge graph constructs the digital life form's cognition of the user's knowledge system and interests, helping the digital life form understand the user's focus, professional field, and life background, and thus provide more targeted information and suggestions in conversations. The set of voiceprint data for the corresponding person endows the digital life form with a personalized voice. When interacting with the user, the digital life form can use the unique timbre of the target person for voice responses, enhancing the authenticity and immersion of the interaction. Integrating these three parts of information, the digital life form has the ability to understand the user's emotions, knowledge, and intentions, as well as the ability to interact with a personalized voice, thereby providing highly personalized interaction services and realizing the personalized simulation and interaction of digital life.

[0138] In one embodiment, the method further includes:

[0139] S61. Combine the user interaction information corresponding to each obtained user input instruction and the personalized response to form new interaction data.

[0140] Record the user interaction in each realization of digital life personalization to form new interaction data.

[0141] S62. Timestamp the new interaction data and store it in the long-term memory repository.

[0142] Store the new interaction data with timestamps in the long-term memory repository to accumulate the user's interaction history. Over time, by analyzing the historical data, we can understand the changes in the user's behavior patterns, emotional evolution, and the development trend of needs, so as to continuously improve the understanding of the user and the service quality.

[0143] S63. Use a graph convolutional network and an attention mechanism to extract new entities and corresponding relationships from the new interaction data and update the user knowledge graph.

[0144] Deeply analyze the new interaction data and integrate the extracted new entities and relationships into the existing user knowledge graph. During the update process, the relevance of the new information to the existing nodes and relationships in the knowledge graph will be checked to avoid duplicate addition and ensure the consistency and accuracy of the knowledge graph. By continuously updating the user knowledge graph, it can timely reflect the new interest points, new cognitions, and new relationships shown by the user during the interaction, so as to understand the user's knowledge system and interest preferences more comprehensively and accurately.

[0145] S64. Optimize the personalized dialogue style based on the long short-term memory network according to the new interaction data.

[0146] The LSTM network optimizes and adjusts the original personalized dialogue style model. Exemplarily, if the new interaction data shows that the user often uses a specific catchphrase when expressing excitement, the LSTM network will incorporate this feature into the personalized dialogue style model so that when generating responses later, it can more naturally simulate the user's dialogue style. By continuously optimizing the personalized dialogue style according to the new interaction data, the generated responses can better fit the user's language habits and emotional states, enhancing the naturalness and personalization of the interaction.

[0147] S65. Distributively store the long-term memory repository and the user knowledge graph.

[0148] Distributed storage technology disperses the data of the long-term memory repository and the user knowledge graph across multiple storage nodes. The storage nodes can be located in different geographical locations and different server devices, and form a storage cluster through network connections. There are many advantages to using distributed storage. First, it improves the reliability and fault tolerance of storage. When a storage node fails, other nodes can continue to provide data services to ensure that the data in the long-term memory repository and the user knowledge graph is not lost or interrupted. Second, distributed storage can enhance the scalability of storage. As the number of users increases and the interaction data continues to accumulate, the data volume of the long-term memory repository and the user knowledge graph will continue to grow. By adding new storage nodes, distributed storage can easily expand the storage capacity to meet the continuously growing data storage requirements.

[0149] In terms of data reading and writing, distributed storage improves the efficiency of data access through data redundancy and load balancing technologies. Multiple nodes can process data requests in parallel, reducing the waiting time for data reading and writing, and ensuring that the system can quickly respond to various query and update operations. This storage method provides strong guarantee for the stable operation and efficient management of the long-term memory library and the user knowledge graph, and supports the system to continuously and stably provide personalized interaction services in scenarios with a large number of users and massive data.

[0150] Optionally, reserved API interfaces support the access of third-party models, facilitating function upgrade and optimization.

[0151] The digital life personality realization method of this application can be applied in the following scenarios:

[0152] When the user has a need for cross-time-and-space emotional companionship, the user inputs "I want to talk to my dad", and obtains the user's voice input "I want to talk to my dad". The FSMN-VAD technology is used for noise suppression and invalid audio filtering to ensure that the input voice is clear and accurate. Then, the voice is converted into text through the end-to-end automatic speech recognition (ASR) module, and error correction is performed using the CTC / Attention mechanism to obtain the accurate user interaction text, that is, to ensure receiving "I want to talk to my dad".

[0153] Subsequently, the system performs emotional analysis on the interaction text, adopts an emotional classification model combining Bi-LSTM and transformer, identifies the emotional label of "nostalgia" of the user, and determines the target person as "dad" at the same time. According to the result of this emotional analysis and the user interaction text, the system indexes in the long-term memory library according to the time stamp, and extracts personalized memories related to dad, such as historical conversations and memories, which contain information such as dad's voice tone, speaking habits, and interaction scenarios with the user in the past.

[0154] Using this information, the emotion-enhanced LLM model, that is, the emotional classification model, based on a hybrid architecture of Transformer and Bi-LSTM, combines the long-term memory library to generate a comforting reply in line with dad's tone. When generating the reply, the model fully considers the user's current emotional state and the conversation style in memory, making the reply more natural and meeting the user's needs. At the same time, the system adopts 256-dimensional voiceprint embedding + GAN training according to the pre-stored voiceprint data of dad to generate a voice tone matching dad, and adjusts the intonation, speech rate, and pause in combination with emotional features to ensure that the speech is natural and fluent.

[0155] In the digital immortality service scenario, assuming that the user wants to build a personal digital avatar, the user needs to input daily conversation and behavior data for three consecutive months into the system, covering voice, text, preferences and other information. The system first pre-processes this data, using FSMN-VAD technology to process voice data, remove noise and invalid segments, and clean and segment text data, and standardize behavior data to obtain standardized user knowledge items.

[0156] Next, the system classifies the user's knowledge items based on the sentiment classification model, and builds a long-term memory library in timestamp index order to record the user's emotional state and interaction content at different times. The system uses graph convolutional networks and attention mechanisms to extract entities and corresponding relationships in user knowledge items, build and continuously update user knowledge graphs, and comprehensively present information such as user interests, hobbies, and interpersonal relationships. The system analyzes the interaction habits of user knowledge items through long-term and short-term memory networks to obtain a behavior pattern database containing personalized conversation styles.

[0157] In terms of voice cloning, the system uses generative adversarial networks and 256-dimensional voiceprint vectors to clone the individual timbre and voice of user knowledge items, and obtains the user's unique voiceprint data set. Based on the long-term memory library, user knowledge graph and voiceprint data set, the system builds a highly simulated personal digital avatar.

[0158] This digital avatar can participate in various activities on behalf of the user. For example, in a meeting, the digital avatar can accurately express the user's views and opinions based on the user's knowledge graph and conversation style, and communicate naturally and smoothly with the participants, with a response delay of <200ms. During the interaction process, the newly generated interaction data will be collected and processed in a timely manner to continuously optimize the performance of the digital avatar and make it more in line with the real user.

[0159] In the cultural heritage innovation scenario, take the example of the ancestors telling the family history in dialect. When the ancestors start to tell the family history, the system's microphone will obtain voice input in real time. The system uses the multi-language voice-to-text function, uses FSMN-VAD for noise suppression and invalid audio filtering, and then converts the dialect voice into text through the end-to-end automatic speech recognition (ASR) module. The dialect recognition accuracy rate is ≥88%. After that, the CTC / Attention mechanism is used for error correction to ensure the accuracy of the text.

[0160] The system extracts entities and relationships in the text, such as family members, historical events, time and locations, etc., based on graph convolutional network and attention mechanism, records them in the order of timestamp indexing, constructs a knowledge graph related to the family, and automatically generates a timeline graph of key events and figures. This information is stored in the family digital archive, using AES-256 encryption and distributed storage, so that the retrieval latency < 1s, ensuring the security and efficient access of data.

[0161] When the next generation wants to learn about the family history, they can wake up the "Family Digital Archive" by voice. After receiving the wake-up command, the system processes the voice to recognize the user's intention. Then it retrieves relevant information from the long-term memory and the family knowledge graph, and uses an emotion-enhanced LLM model to generate responses in the tone and style of the ancestors. It uses pre-stored voiceprint data of the ancestors for voice cloning to synthesize voices with the dialect characteristics of the ancestors. At the same time, based on the phoneme interpolation algorithm, it synchronizes the voice with the lip animation of the virtual ancestor image. The lip synchronization adapts to the dialect pronunciation characteristics, and the matching accuracy ≥ 94%. It is presented to the descendants through the video stream to achieve virtual conversations between grandparents and grandchildren, enabling the vivid inheritance of family oral history. In this process, the newly generated interaction data will also be timestamped and stored in the long-term memory, and the family knowledge graph will be updated according to the new information, accumulating richer materials for the inheritance of family culture.

[0162] Finally, based on the phoneme interpolation algorithm, the virtual image is synchronized with the generated audio data to generate a video stream, and then the audio data and the video stream are combined and rendered to obtain an interactive video for presentation to the user. During the entire interaction process, the newly generated interaction data, such as the user's questions and the system's responses, will be timestamped and stored in the long-term memory for further optimizing the system's understanding of the user's emotions and interaction habits.

[0163] It should be understood that although each step in the flowcharts involved in the above-described embodiments is shown in sequence according to the indication of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same moment, but can be executed at different moments. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.

[0164] Based on the same inventive concept, an embodiment of the present application further provides a digital life personalization implementation device for implementing the digital life personalization implementation method involved above. The implementation solution provided by this device to solve problems is similar to the implementation solution described in the above method. Therefore, the specific limitations in one or more embodiments of the digital life personalization implementation device provided below can refer to the limitations on the digital life personalization implementation method in the above text, and will not be elaborated here.

[0165] In an exemplary embodiment, as Figure 3 shown, a digital life personalization implementation device is provided, including:

[0166] A data acquisition module, configured to acquire user interaction information in response to the acquired user input instruction;

[0167] A data processing module, configured to preprocess the user interaction information according to a set mode to obtain a user interaction text;

[0168] A personalization module, configured to input the user interaction text into a digital life body to obtain a personalized reply; the personalized reply includes a target person and audio data;

[0169] An interaction module, configured to generate an interaction video corresponding to the personalized reply based on digital human rendering.

[0170] In one of the embodiments, it further includes a data type recognition module, configured to recognize the data type of the user interaction information to obtain a classification result.

[0171] In one of the embodiments, it further includes an emotion analysis module, a memory repository indexing module, and an emotion enhancement module;

[0172] The emotion analysis module is configured to perform emotion analysis on the user interaction text to obtain an emotion analysis result;

[0173] The memory repository indexing module is configured to index in the memory repository according to the user interaction text and the emotion analysis result by timestamp to obtain personalized memories;

[0174] The emotion enhancement module is configured to perform emotion enhancement on the user interaction text and the personalized memories based on the behavior pattern database to obtain a personalized text reply.

[0175] In one of the embodiments, it further includes an audio-visual synchronization module, configured to synchronize the virtual image and the audio data based on a phoneme interpolation algorithm to obtain a video stream.

[0176] In an embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps in the above method embodiments are implemented.

[0177] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0178] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to the partial descriptions of the method embodiments. The device embodiments described above are only illustrative. The components described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the present disclosure solution. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0179] The above embodiments only represent several implementation manners of the embodiments of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the embodiments of the application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the embodiments of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the embodiments of the present application.

Claims

1. A method for realizing digital life personalization, characterized in that, The method includes: Upon receiving a user input instruction, obtaining user interaction information; Preprocessing the user interaction information in a set mode to obtain a user interaction text; Inputting the user interaction text into a digital life form to obtain a personalized response; the personalized response includes a target person and audio data; Generating an interaction video corresponding to the personalized response based on digital human rendering.

2. The method according to claim 1, wherein The preprocessing the user interaction information in a set mode to obtain a user interaction text includes: Identifying the data type of the user interaction information to obtain a classification result; the classification result includes a voice data type and a text data type; If the classification result is the voice data type, processing the user interaction information according to a voice processing mode to obtain the user interaction text; If the classification result is the text data type, performing natural language processing on the user interaction information to obtain the user interaction text; Wherein, the voice processing mode corresponds to the following operation steps: Performing noise suppression and invalid audio filtering on the user interaction information according to voice activity detection based on a memory network to obtain standardized voice data; Using end-to-end automatic speech recognition to convert the standardized voice data into text to obtain text data; Performing error correction on the text data based on a sequence processing method to obtain the user interaction text.

3. The method according to claim 1, characterized in that The digital life form obtains the personalized response through the following method: Performing sentiment analysis on the user interaction text to obtain a sentiment analysis result; the sentiment analysis result includes a sentiment label and a target person; Indexing according to the time stamp in a memory repository based on the user interaction text and the sentiment analysis result to obtain personalized memories; the memory repository includes a long-term memory repository and a user knowledge graph; Performing sentiment enhancement on the user interaction text and the personalized memories based on a behavior pattern database to obtain a personalized text response; Performing voice cloning on the personalized text response according to the pre-stored voiceprint data of the target person to obtain the personalized response.

4. The method according to claim 3, wherein The indexing according to the time stamp in a memory repository based on the user interaction text and the sentiment analysis result to obtain personalized memories includes: Calculating a semantic vector based on the user interaction text and the sentiment analysis result; Calculating the similarity between the semantic vector and the sentiment feature vector in the long-term memory repository to obtain a semantic similarity; Determining the historical memory in the long-term memory repository corresponding to the semantic similarity exceeding a set threshold as historical dialogue data based on time stamp indexing; Searching in the user knowledge graph for the knowledge graph corresponding to the target person based on time stamp indexing to obtain historical interest data; Obtaining personalized memories according to the historical dialogue data and the historical interest data; The semantic similarity is obtained through the following formula: Wherein, Q is the semantic vector; D is the sentiment feature vector; n is the number of sub - parts into which the semantic vector and the sentiment feature vector are divided; Q i , D i correspond to the i - th sub - part of the semantic vector and the sentiment feature vector respectively; α i is the weight corresponding to each sub - part.

5. The method according to claim 1, characterized in that The generating an interaction video corresponding to the personalized response based on digital human rendering includes: Synchronizing the virtual image with the audio data based on a phoneme interpolation algorithm to obtain a video stream; the virtual image corresponds to the target person; Render the audio data in combination with the video stream to obtain the interactive video.

6. The method according to any one of claims 1 to 5, characterized in that, The digital life form is constructed by the following method: Obtain user personalized data; the user personalized data includes voice interaction data, text interaction data, and behavior pattern data; Preprocess the user personalized data to obtain user knowledge entries; Based on an emotion classification model, perform emotion classification on the user knowledge entries and construct a long-term memory bank in the order of timestamp indexing; The long-term memory bank contains emotion feature vectors; Based on a graph convolutional network and an attention mechanism, extract the entities and corresponding relationships in the user knowledge entries and record them in the order of timestamp indexing to obtain a user knowledge graph; Use a long short-term memory network to analyze the user interaction habits of the user knowledge entries to obtain a behavior pattern database; the behavior pattern database includes personalized dialogue styles; Use a generative adversarial network and 256-dimensional voiceprint vectors to perform individual voiceprint sound cloning on the user knowledge entries to obtain a set of voiceprint data for the corresponding person; Obtain a digital life form based on the long-term memory bank, the user knowledge graph, and the set of voiceprint data for the corresponding person.

7. The method according to claim 6, characterized in that, The method further includes: Construct new interaction data from the user interaction information and the personalized reply corresponding to each user input instruction obtained in response; Annotate the new interaction data with timestamps and store it in the long-term memory bank; Use a graph convolutional network and an attention mechanism to extract new entities and corresponding relationships from the new interaction data and update the user knowledge graph; Based on a long short-term memory network, optimize the personalized dialogue style according to the new interaction data; Distributively store the long-term memory bank and the user knowledge graph.

8. A digital life personalization implementation device, characterized in that, The device includes: A data acquisition module, configured to obtain user interaction information in response to a user input instruction obtained; A data processing module, configured to preprocess the user interaction information in a set mode to obtain a user interaction text; A personalization module, configured to input the user interaction text into a digital life form to obtain a personalized reply; the personalized reply includes a target person and audio data; An interaction module, configured to generate an interactive video corresponding to the personalized reply based on digital human rendering.

9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Digital human high-quality video generation method and system based on data driving

    CN120689477A

  • Configuration method and device of digital life entity, equipment and storage medium

    CN121033918A

  • A digital life body configuration method, device, equipment and storage medium

    CN121033918B

  • Digital NPC role emotion expression and interaction system

    CN121304874A

  • Digital npc character emotion expression and interaction system

    CN121304874B