Digital human interaction video synthesis method and device, electronic equipment and storage medium
By acquiring data and sentiment analysis of the target object, digital human video data is generated to achieve natural interaction, which solves the problems of insufficient intelligence and adaptability of digital human systems and realizes a closed-loop experience of emotion perception and personalized expression.
Patent Information
- Application Number
- CN202511682293.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-17
- Publication Date
- 2026-02-17
AI Technical Summary
Current digital human systems are insufficient in terms of intelligence, naturalness of interaction, and adaptability to applications, and cannot achieve a closed-loop experience of emotion perception, personalized expression, and natural interaction.
By acquiring initial data about the target object and determining its emotional data, video data for digital human responses, including facial expressions and movements, is generated based on the emotional data to ensure the naturalness and personalization of the interaction.
It achieves a closed-loop experience of emotion perception, personalized expression, and natural interaction, improving the intelligence level and naturalness of interaction of the digital human system.
Smart Images

Figure CN121547665A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method, apparatus, electronic device, and storage medium for synthesizing digital human interactive videos. Background Technology
[0002] The application of digital human technology in question-and-answer systems is becoming increasingly widespread. By combining cutting-edge technologies such as artificial intelligence, speech recognition, natural language processing, computer vision, and 3D rendering, it provides users with a more intelligent, intuitive, and human-like interactive experience. For example, in the field of intelligent customer service (such as customer service Q&A in client-side apps), digital humans, acting as virtual customer service representatives, can interact with users in real time via voice or video to answer questions and provide assistance. However, current digital human systems still have shortcomings in terms of intelligence level, naturalness of interaction, and application adaptability, and further improvements are needed. Summary of the Invention
[0003] This invention provides a digital human interactive video synthesis method, apparatus, electronic device, and storage medium to address the current digital human systems' shortcomings in intelligence level, naturalness of interaction, and application adaptability, thereby achieving a closed-loop experience of "emotion perception - personalized expression - natural interaction".
[0004] According to one aspect of the present invention, a method for synthesizing interactive digital human videos is provided, the method comprising:
[0005] Acquire first data sent by the target object, and determine first emotional data of the target object based on the first data; the first data includes text content and / or audio data; the first data is data formed by questions raised by the target object in response to different needs;
[0006] Based on the first data and the first sentiment data, second data is determined, and based on the second data, a first duration is determined; the second data is the response text to the question in the first data; the second data includes second sentiment data guiding the digital human to respond to the first data; the first duration is used to describe the duration used by the digital human to express the second data;
[0007] Based on the second emotion data in the second data, third data is determined, and first video data is formed based on the third data; the third data is used to describe the expressions and actions of the digital human in expressing the second data; the first video data is a video of the digital human completely displaying a set of third data;
[0008] Based on the second data, the first duration, and the first video data, second video data is determined so that the digital human can respond and display based on the second video data; the second video data is a video of the digital human expressing the second data.
[0009] According to another aspect of the present invention, a digital human interactive video synthesis apparatus is provided, the apparatus comprising:
[0010] The first information processing module is used to acquire first data sent by the target object and determine the first emotional data of the target object based on the first data; the first data includes text content and / or audio data; the first data is data formed by the questions raised by the target object in response to different needs.
[0011] The second information processing module is used to determine second data based on the first data and the first emotional data, and to determine a first duration based on the second data; the second data is a response text to the question in the first data; the second data includes second emotional data guiding the digital human to respond to the first data; the first duration is used to describe the duration used by the digital human to express the second data;
[0012] The first video data determination module is used to determine third data based on the second emotional data in the second data, and to form first video data based on the third data; the third data is used to describe the expressions and actions of the digital human in expressing the second data; the first video data is a video of the digital human completely displaying a set of third data;
[0013] The second video data determination module is used to determine second video data based on the second data, the first duration, and the first video data, so that the digital human can respond and display based on the second video data; the second video data is a video of the digital human expressing the second data.
[0014] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:
[0015] At least one processor; and
[0016] A memory communicatively connected to the at least one processor; wherein,
[0017] The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the digital human interactive video synthesis method according to any embodiment of the present invention.
[0018] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the digital human interactive video synthesis method according to any embodiment of the present invention.
[0019] The technical solution of this invention involves acquiring first data sent by a target object, and determining first emotional data of the target object based on the first data. The first data includes text content and / or audio data. The first data is data formed by questions raised by the target object in response to different needs. This enables accurate positioning of the emotional state of the target object when asking questions. Then, based on the first data and the first emotional data, second data for answering the questions in the first data is determined, and based on the second data, a first duration for describing the digital human's expression of the second data is determined, so that video generation can be accurately performed subsequently. Because the second data contains second emotional data guiding the digital human to answer the first data, third data for describing the digital human's facial expressions and actions in expressing the second data can be accurately determined based on the second emotional data in the second data. At the same time, a first video data for the digital human to completely display a set of third data is formed based on the third data. The second video data for the digital human to express the second data is determined based on the second data, the first duration, and the first video data, so that the digital human can respond based on the second video data. This solves the problems of insufficient intelligence, naturalness of interaction, and application adaptability in current digital human systems, and achieves a closed-loop experience of "emotion perception - personalized expression - natural interaction".
[0020] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 This is a flowchart of a digital human interactive video synthesis method provided by an embodiment of the present invention;
[0023] Figure 2 This is a schematic diagram of a digital human display interface applicable to an embodiment of the present invention;
[0024] Figure 3 This is a flowchart of a digital human interactive video synthesis method provided by an embodiment of the present invention;
[0025] Figure 4 This is a schematic diagram of the structure of a digital human interactive video synthesis device according to an embodiment of the present invention;
[0026] Figure 5This is a schematic diagram of the structure of an electronic device that implements the digital human interactive video synthesis method of the present invention, according to an embodiment of the present invention. Detailed Implementation
[0027] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0028] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0029] Example 1
[0030] Figure 1 This is a flowchart illustrating a digital human interactive video synthesis method provided in an embodiment of the present invention. This embodiment is applicable to situations involving digital human interactive video synthesis during user-digital human interaction in various scenarios. Examples of application scenarios include, but are not limited to, online education, telemedicine, financial services, smart government services, mental health services (emotional companion robots), and e-commerce live streaming (virtual anchors). This method can be executed by a digital human interactive video synthesis device, which can be implemented in hardware and / or software and can be configured in any electronic device with network communication capabilities. Figure 1 As shown, the digital human interactive video synthesis method of the present invention may include:
[0031] S110. Obtain the first data sent by the target object, and determine the first emotional data of the target object based on the first data; the first data includes text content and / or audio data; the first data is data formed by the questions raised by the target object in response to different needs.
[0032] The first piece of data can be used to indicate the information of the questions that need to be answered and the service area to which they belong.
[0033] Specifically, the target object is the object that interacts with the digital human through an interactive device. The target object can send initial data by triggering controls on the display interface of the interactive device, and the digital human is then displayed on the display interface of the interactive device. For example, Figure 2 As shown. When not performing a response task, the digital human will periodically display preset videos with preset expressions and preset actions. To reduce complexity, the digital human can only have an upper body, and the digital human's dynamics can be limited to voice, facial expressions, mouth movements, and hand movements (including subtle body movements associated with hand movements).
[0034] Preset videos can embed periodic timestamp watermarks. These watermarks can be embedded into the video's luminance components using a DCT domain steganography algorithm, with an extraction error rate of <0.1%. This ensures no impact on visual quality and allows for quick verification of the video's source using dedicated tools. Simultaneously, the metadata of the preset video includes "base video version number + action template library version number" to distinguish different versions of the base material, improving the system's material management efficiency. The timestamp watermark can be a SHA-256 hash value based on the base video version number and the generated timestamp. The periodicity can be set according to requirements, such as 30 seconds.
[0035] When the first data is text content, determining the first sentiment data of the target object based on the first data may include: performing sentiment text recognition processing on the first data using a text recognition model to determine first reference sentiment data; the text recognition model is used to describe the mapping relationship between text features and sentiment data. For example, the text recognition model can be the BERT model.
[0036] Alternatively, a domain sentiment knowledge base can be obtained. This knowledge base uses a "triple" structure, where each triple includes a sentiment word, its corresponding trigger phrase, and the service domain. Based on the domain sentiment knowledge base and the first data, first reference sentiment data is determined. Specifically, the trigger phrases and service domains in the first data that match the triples in the domain sentiment knowledge base are identified, thereby determining the sentiment words associated with those trigger phrases and using them as the first reference sentiment data.
[0037] When the first data consists of text content and audio data, determining the first sentiment data of the target object based on the first data may include: extracting text from the audio data using a speech recognition model to obtain second text data; obtaining intonation information from the audio data and determining second reference sentiment data based on the intonation information; performing sentiment text recognition processing on the text content using a text recognition model to determine third reference sentiment data; performing sentiment text recognition processing on the second text data using a text recognition model to determine fourth reference sentiment data; and performing weighted fusion processing on the second, third, and fourth reference sentiment data to obtain the first sentiment data of the target object.
[0038] S120. Based on the first data and the first emotional data, determine the second data, and based on the second data, determine the first duration; the second data is the response text to the question in the first data; the second data includes the second emotional data that guides the digital human to respond to the first data; the first duration is used to describe the duration used by the digital human to express the second data.
[0039] The second emotional data can be understood as the emotions that the digital human needs to display when responding to the first data.
[0040] Specifically, the first data indicates the question information that needs to be answered and the service domain to which it belongs. This allows for the extraction of response text—the second data—from the domain knowledge base of the corresponding service domain, based on the first data and the first sentiment data. The inclusion of the first sentiment data further ensures that the generated second data contains more accurate sentiment data.
[0041] Furthermore, extracting the response text that can answer the first data from the domain knowledge base of the corresponding service domain may include: calculating the vector similarity between the first data and the content in the domain knowledge base; when the similarity is greater than a preset similarity threshold, it can be determined as the response text that can answer the first data; further, an extraction-based summarization algorithm is used to summarize the response text that can answer the first data, retaining the core entities and solution logic, to obtain the second data; at the same time, in order to ensure that the response time is not too long and affects the user experience, a greedy truncation method can be used to control the number of words in the second data.
[0042] Determining the first duration based on the second data may include: there is a first preset correlation between the number of characters in the response text and the response duration; the first duration can be determined based on the first preset correlation and the number of characters in the second data. The first preset correlation may be a linear function.
[0043] S130. Based on the second emotional data in the second data, determine the third data, and form the first video data based on the third data; the third data is used to describe the expressions and actions of the digital human in expressing the second data; the first video data is a video of the digital human completely displaying a set of the third data.
[0044] Specifically, a second pre-defined correlation exists between emotional data and facial expressions and actions. Based on this second pre-defined correlation and the second emotional data, third data is determined. Actions in the third data are connected using a skeletal keypoint trajectory clustering method, such as the DBSCAN algorithm, to ensure smoothness. Simultaneously, morphological target (MorphTargets) technology is used to adjust the facial expressions in the third data. The facial expressions and actions in a set of third data are arranged according to the time dimension to accurately form the first video data. Furthermore, optical flow consistency is checked on the first video data, requiring that the optical flow vector difference between consecutive pre-defined frames is less than a pre-defined pixel value, the subjective score for visual jumps in transition areas is reduced to a pre-defined score, and ambient audio uses spectral envelope alignment technology to avoid abrupt auditory changes.
[0045] In an embodiment of the present invention, optionally, the second emotional data includes emotional values of different emotional data, and determining the third data based on the second emotional data may include steps A1-A3:
[0046] Step A1: Based on the first mapping rule and the second emotional data, determine the reference facial expressions and reference actions of the digital human; the first mapping rule is used to describe the mapping relationship between emotional data and facial expressions and actions.
[0047] For example, the first mapping rule could include: anger corresponds to frowning + waving hands, happiness corresponds to smiling + nodding, and doubt corresponds to tilting the head + blinking at a preset frequency, etc., mapping relationships between emotional data and facial expressions and actions.
[0048] Step A2: Based on the second mapping rule and the sentiment value in the second sentiment data, determine the expression amplitude of the reference expression and the motion amplitude of the reference action; the second mapping rule uses a non-linear mapping function to describe the relationship between the sentiment value and the motion intensity of the sentiment data.
[0049] Specifically, the second mapping rule uses a non-linear mapping function to describe the relationship between the emotional value and the intensity of the emotional data, avoiding distortion of actions under extreme emotions; for example, the frowning amplitude corresponding to the anger value x can be 20%×√x.
[0050] Furthermore, the Bezier curve dynamically adjusts the parameters, using low curvature for smooth movements, such as a curvature of 0.3 for nodding, and high curvature for emphasizing movements, such as a curvature of 0.7 for waving hands. At the same time, a breathing design is introduced, inserting micro-pauses of 50-100ms between continuous movements to simulate the relaxed state of real muscles, which further improves the naturalness score of the movements by 18%.
[0051] Step A3: Determine the reference facial expression, reference action, the facial expression range of the reference facial expression, and the action range of the reference action as the third data.
[0052] Optionally, the present invention may design an emotion decay mechanism to automatically reset the emotion value of an object that has not interacted for more than a preset number of minutes to a preset value. The preset value is used to indicate that the object's emotional state is neutral, that is, a stable state, so as to avoid the accumulation of emotional misjudgments.
[0053] In this embodiment of the invention, reference expressions and reference actions for a digital human are determined based on a first mapping rule and second emotional data. The first mapping rule describes the mapping relationship between emotional data and expressions and actions, ensuring the accuracy of the determined reference expressions and actions. Then, based on the second mapping rule and the emotional values in the second emotional data, the amplitude of the reference expression and the amplitude of the reference action are determined. The introduction of the second mapping rule, which uses a non-linear mapping function to describe the relationship between emotional values and the intensity of actions in emotional data, avoids distortion of actions under extreme emotions. The reference expression, reference action, amplitude of the reference expression, and amplitude of the reference action are determined as third data to make the digital human's reactions appear as natural and realistic as possible.
[0054] S140. Based on the second data, the first duration, and the first video data, determine the second video data so that the digital human can respond and display based on the second video data; the second video data is a video of the digital human expressing the second data.
[0055] Specifically, the first duration determines the minimum duration of the second video data, the second data determines the information the digital human replies with, the first video data determines the actions and expressions the digital human displays while narrating the second data, and the first video data is periodically superimposed to form the expressions and actions of the second video data.
[0056] Optionally, in an embodiment of the present invention, determining the second video data based on the second data, the first duration, and the first video data may include: determining a video period based on the first duration and the second duration of the first video data; reversing the first video data to obtain the third video data; and determining the second video data based on the second data, the first video data, the third video data, and the video period.
[0057] Specifically, the first video data is reversed using frame order reversal and motion blur compensation. Then, according to the video cycle, the first video data and the third video data are cyclically spliced together until the number of video cycles is reached.
[0058] That is, if the video period is even, the number of the first video data and the number of the third video data are video period / 2 respectively. Then the first video data and the third video data can be spliced together to form the fourth video data, and then the video period / 2 of the fourth video data can be spliced together to form the spliced video data.
[0059] If the video period is odd, the number of videos in the first video data is the video period / 2 rounded up, and the number of videos in the second video data is the video period / 2 rounded down. The first and third video data are concatenated to form the fourth video data. The number of videos in the second video data is then concatenated with the fourth video data to form the fifth video data. Finally, the fifth video data is concatenated with the first video data to form the final concatenated video data. A 0.5s crossfade-in / fade-out transition (linear change in the alpha channel) can be added during concatenation to ensure inter-frame optical flow continuity (optical flow vector difference < 5 pixels / frame).
[0060] Furthermore, the spliced video data is then combined with the spoken content from the second data to form the second video data.
[0061] In this embodiment of the invention, determining the second video data based on the second data, the first duration, and the first video data may include: determining the video period based on the first duration and the second duration of the first video data; reversing the first video data to obtain the third video data; and determining the second video data based on the second data, the first video data, the third video data, and the video period, thereby achieving accurate generation of the third video data played when the digital human responds to the first data.
[0062] Optionally, in an embodiment of the present invention, after determining the second video data, the method further includes: determining the end time point for generating the second video data, and determining a dynamic buffer time; the dynamic buffer time is used to adjust the insertion time point of the second video data; determining the video insertion time point based on the end time point and the dynamic buffer time; determining a video splicing strategy based on the device type of the target object; and inserting the second video data at the video insertion time point of a preset video based on the video splicing strategy; the preset video is a video showing preset expressions and preset actions when the digital human does not perform a response task.
[0063] The dynamic buffer time can be preset to ensure that the video insertion time is slightly later than the end time of generating the second video data. The dynamic buffer time is generally 1-3 seconds.
[0064] The dynamic buffering time can be dynamically adjusted based on the device performance of the target object's interactive device and / or the scene complexity corresponding to the second video data. Device performance can include CPU utilization, historical stuttering records, and device type; historical stuttering records can be used to describe the number of stutters per minute. Different scene complexities are associated with corresponding dynamic buffering times.
[0065] For example, when the CPU utilization rate is greater than the preset utilization rate (80%), the dynamic buffer time is the first preset time (3s); otherwise, the dynamic buffer time is the second preset time (1s); the first preset time is greater than the second preset time. Historical stuttering records can be weighted, meaning the dynamic buffer time is updated using historical stuttering records and the weighting factor. The updated dynamic buffer time = historical stuttering record × weighting factor + dynamic buffer time. Device type can also be used to adjust the dynamic buffer time. For example, if the device type is a basic mobile phone, a preset increment can be added to the original dynamic buffer time.
[0066] The scene complexity corresponding to the second video data can be calculated using the improved Canny operator. If the scene complexity is greater than the first scene high threshold, then the scene complexity is a dynamic scene; if the scene complexity is less than the second scene high threshold, then the scene complexity is a static scene; and if the first scene high threshold is greater than the second scene high threshold.
[0067] In addition, scene complexity can be determined based on the dynamic pixel ratio calculated by the edge detection operator. The dynamic pixel ratio can be understood as the ratio of dynamic pixels to the total number of pixels. Segments with a dynamic pixel ratio of <10% can be selected as insertion points.
[0068] In addition, a preset number of alternative insertion time points (with a 2-second interval) can be reserved to prevent the main insertion time point from failing due to sudden performance degradation.
[0069] Video stitching strategy can be understood as a stitching strategy determined based on a device grading table. The device grading table can be divided according to device performance, which can be the model and GPU performance. For example, the model is a terminal for elderly users, which is matched to Level 1 (720P resolution + simplified effects) by default, and adopts "progressive decoding" (decodes I-frames first and then completes P / B frames), which reduces the insertion time from 1.2s to 0.6s; energy consumption optimization is achieved by dynamically adjusting the encoding bitrate (reduced to 1.5Mbps for elderly users), combined with a sleep-wake mechanism (GPU sleeps for 200ms after insertion), which extends the battery life of elderly users by 1.8 hours in actual tests; the preloading mechanism combines the user's historical interaction time (if the user's historical interaction time is greater than 5 minutes, a 60s segment is loaded) to avoid waiting and stuttering when inserting again.
[0070] After determining the second video data, this invention determines the end time point for generating the second video data and the dynamic buffer time. The dynamic buffer time is used to adjust the insertion time point of the second video data. Based on the end time point and the dynamic buffer time, the video insertion time point is determined. Based on the device type of the target object, a video splicing strategy is determined. Based on the video splicing strategy, the second video data is inserted at the preset video insertion time point to ensure that the digital human can display actions continuously and to ensure the continuity and naturalness of the entire video playback process.
[0071] Optionally, after the second video data is inserted, the system automatically generates an insertion log containing "insertion timestamp + video A / B splicing parameter hash + terminal device identifier". The log is stored using AES-256 encryption and uploaded to the blockchain node (block generation interval ≤ 10s) to ensure that the log is tamper-proof and traceable. At the same time, "insertion algorithm version number (including video A / B duration ratio parameter)" is written to the metadata area of the synthesized video (the video with the second video data inserted).
[0072] The technical solution of this invention involves acquiring first data sent by a target object, and determining first emotional data of the target object based on the first data. The first data includes text content and / or audio data. The first data is data formed by questions raised by the target object in response to different needs. This enables accurate positioning of the emotional state of the target object when asking questions. Then, based on the first data and the first emotional data, second data for answering the questions in the first data is determined, and based on the second data, a first duration for describing the digital human's expression of the second data is determined, so that video generation can be accurately performed subsequently. Because the second data contains second emotional data guiding the digital human to answer the first data, third data for describing the digital human's facial expressions and actions in expressing the second data can be accurately determined based on the second emotional data in the second data. At the same time, a first video data for the digital human to completely display a set of third data is formed based on the third data. The second video data for the digital human to express the second data is determined based on the second data, the first duration, and the first video data, so that the digital human can respond based on the second video data. This solves the problems of insufficient intelligence, naturalness of interaction, and application adaptability in current digital human systems, and achieves a closed-loop experience of "emotion perception - personalized expression - natural interaction".
[0073] Example 2
[0074] Figure 3 This is a flowchart of a digital human interactive video synthesis method provided by an embodiment of the present invention. The technical solution of this embodiment further optimizes the process of S140 in the aforementioned embodiments based on the above embodiments. This embodiment can be combined with various optional solutions in one or more of the above embodiments. Figure 3As shown, the digital human interactive video synthesis method of the present invention may include:
[0075] S210. Obtain the first data sent by the target object, and determine the first emotional data of the target object based on the first data; the first data is audio data; the first data is data formed by the questions raised by the target object in response to different needs.
[0076] Specifically, text extraction is performed on the audio data based on a speech recognition model to obtain first text data; intonation information is obtained from the audio data, and first reference sentiment data is determined based on the intonation information; sentiment text recognition processing is performed on the first text data based on a text recognition model to determine second reference sentiment data; the text recognition model is used to describe the mapping relationship between text features and sentiment data; the text recognition model can be a BERT model. The first and second reference sentiment data are then weighted and fused to obtain the first sentiment data of the target object.
[0077] Specifically, intonation information can include fundamental frequency, short-time energy, and speech rate features. A CNN-LSTM model can be used to classify the first and second reference sentiment data, assign a first weight to the first reference sentiment data, and assign a second weight to the second reference sentiment data. Then, based on the first and second reference sentiment data, the first and second weights, the first sentiment data of the target object can be determined.
[0078] S220. Obtain voiceprint feature information from audio data, and determine the age characteristics of the target object based on the voiceprint feature information; determine the second data based on the language big model, age characteristics and first sentiment data; the language big model includes a sentiment-speech mapping layer; the sentiment-speech mapping layer is used to describe the mapping relationship between different sentiment data and response text.
[0079] Among them, the large language model can be a conversational model with a parameter scale of 10B or more, which is connected to the domain knowledge base through RAG technology. For example, the large language model can be LLaMA-2-13B.
[0080] Specifically, a feature library linking age features and voiceprint features can be introduced to determine the age features of the target object based on the voiceprint features and the feature library. Furthermore, based on the language model, age features, and primary sentiment data, it is possible to better generate response texts and response text structures that correspond to the age features, i.e., secondary data.
[0081] For example, the angry target uses a three-part structure of "apology + solution + compensation promise" (average sentence length ≤ 8 characters), while the elderly target uses a "terminology explanation library" (which can be understood as colloquial expressions), increasing the proportion of explanatory statements to 40%. At the same time, a dynamic truncation mechanism is designed so that when it is detected that the target asks the same question twice in a row, the reply text is automatically shortened to 60% of the original length and an example is added, reducing the repeated questioning rate by 29%.
[0082] S230. Obtain the digital human's response speed, and determine the first duration based on the response speed and the number of words in the second data.
[0083] Specifically, the response rate can be understood as the number of words replied per second. The product of the response rate and the number of words in the second data is the first duration.
[0084] S240. Based on the second emotional data in the second data, determine the third data, and form the first video data based on the third data; the third data is used to describe the expressions and actions of the digital human in expressing the second data; the first video data is a video of the digital human completely displaying a set of the third data.
[0085] S250. Based on the second data, the first duration, and the first video data, determine the second video data so that the digital human can respond and display based on the second video data; the second video data is a video of the digital human expressing the second data.
[0086] The technical solution of this invention involves acquiring first data sent by a target object, and determining first emotional data of the target object based on the first data. The first data is audio data, which is data formed by questions raised by the target object in response to different needs. Voiceprint feature information is acquired from the audio data, and the age characteristics of the target object are determined based on the voiceprint feature information. Second data is determined based on a language model, age characteristics, and the first emotional data. The language model includes an emotion-speech mapping layer. The emotion-speech mapping layer, used to describe the mapping relationship between different emotional data and response text, can automatically generate accurate and suitable answers and control the length of the response text to ensure that the response speech is concise and refined. The introduction of age characteristics takes into account the needs of users of different age groups, especially the elderly. This design allows more users to conveniently use the digital human customer service system, expanding the potential user base. Furthermore, the response speech rate of the digital human is acquired, and the first duration is accurately determined based on the response speech rate and the number of words in the second data, so that video generation can be accurately performed subsequently. Based on the second emotional data in the second data set, third data is determined, and first video data is formed based on the third data. The third data is used to describe the expressions and actions of the digital human in expressing the second data. The first video data is a video of the digital human completely displaying a set of third data. Based on the second data, the first duration, and the first video data, second video data is determined so that the digital human can respond based on the second video data. The second video data is a video of the digital human expressing the second data, in order to solve the problems of insufficient intelligence, naturalness of interaction, and application adaptability of current digital human systems, and to achieve a closed-loop experience of "emotion perception - personalized expression - natural interaction".
[0087] Example 3
[0088] Figure 4 This is a schematic diagram of a digital human interactive video synthesis device provided in an embodiment of the present invention. This embodiment is applicable to the synthesis of digital human interactive videos during user-digital human interaction in different scenarios. The digital human interactive video synthesis device can be implemented in hardware and / or software, and can be configured in any electronic device with network communication capabilities. Figure 4 As shown, the digital human interactive video synthesis device of the present invention includes:
[0089] The first information processing module 310 is used to acquire first data sent by the target object and determine the first emotional data of the target object based on the first data; the first data includes text content and / or audio data; the first data is data formed by the questions raised by the target object in response to different needs.
[0090] The second information processing module 320 is used to determine second data based on the first data and the first emotional data, and to determine a first duration based on the second data; the second data is a response text to the question in the first data; the second data includes second emotional data guiding the digital human to respond to the first data; the first duration is used to describe the duration used by the digital human to express the second data;
[0091] The first video data determination module 330 is used to determine third data based on the second emotional data in the second data, and to form first video data based on the third data; the third data is used to describe the expressions and actions of the digital human in expressing the second data; the first video data is a video of the digital human completely displaying a set of third data;
[0092] The second video data determination module 340 is used to determine second video data based on the second data, the first duration, and the first video data, so that the digital human can respond and display based on the second video data; the second video data is a video of the digital human expressing the second data.
[0093] Based on the above embodiments, optionally, the first information processing module is configured to: extract text from the audio data based on a speech recognition model to obtain first text data; obtain intonation information from the audio data and determine first reference sentiment data based on the intonation information; perform sentiment text recognition processing on the first text data based on a text recognition model to determine second reference sentiment data; the text recognition model is used to describe the mapping relationship between text features and sentiment data; and perform weighted fusion processing on the first reference sentiment data and the second reference sentiment data to obtain the first sentiment data of the target object.
[0094] Based on the above embodiments, optionally, the second information processing module includes a second data determination unit, which is used to: acquire voiceprint feature information in the audio data, determine the age feature of the target object based on the voiceprint feature information; determine second data based on the language big model, the age feature and the first emotion data; the language big model includes an emotion-speech mapping layer; the emotion-speech mapping layer is used to describe the mapping relationship between different emotion data and response text.
[0095] Optionally, based on the above embodiments, the second information processing module includes a duration determination unit, which is used to: obtain the digital human's response speech rate, and determine the first duration based on the response speech rate and the number of words in the second data.
[0096] Based on the above embodiments, optionally, the second emotion data includes emotion values of different emotion data, and the first video data determination module is used to: determine the reference expression and reference action of the digital human based on the first mapping rule and the second emotion data; the first mapping rule is used to describe the mapping relationship between emotion data and expression and action; based on the second mapping rule and the emotion values in the second emotion data, determine the expression amplitude of the reference expression and the action amplitude of the reference action; the second mapping rule uses a nonlinear mapping function to describe the relationship between the emotion value and the action intensity of the emotion data; and determine the reference expression, the reference action, the expression amplitude of the reference expression, and the action amplitude of the reference action as third data.
[0097] Optionally, based on the above embodiments, the second video data determination module is used to: determine the video period based on the first duration and the second duration of the first video data; reverse the first video data to obtain the third video data; and determine the second video data based on the second data, the first video data, the third video data, and the video period.
[0098] Optionally, based on the above embodiments, the digital human interactive video synthesis device further includes a video stitching module. The video stitching module is used to: determine the end time point for generating the second video data and determine a dynamic buffer time; the dynamic buffer time is used to adjust the insertion time point of the second video data; determine the video insertion time point based on the end time point and the dynamic buffer time; determine a video stitching strategy based on the device type of the target object; and insert the second video data at the video insertion time point of a preset video based on the video stitching strategy; the preset video is a video in which the digital human displays preset expressions and preset actions when it does not perform a response task.
[0099] The digital human interactive video synthesis apparatus provided in the embodiments of the present invention can execute the digital human interactive video synthesis method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of executing the method.
[0100] Example 4
[0101] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0102] Figure 5A schematic diagram of an electronic device that can be used to implement the digital human interactive video synthesis method of embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0103] like Figure 5 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0104] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0105] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as digital human interactive video synthesis methods.
[0106] In some embodiments, the digital human interactive video synthesis method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via read-only memory (ROM) 12 and / or communication unit 19. When the computer program is loaded into random access memory (RAM) 13 and executed by processor 11, one or more steps of the digital human interactive video synthesis method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the digital human interactive video synthesis method by any other suitable means (e.g., by means of firmware).
[0107] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0108] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0109] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0110] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0111] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0112] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0113] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0114] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A method for synthesizing interactive digital human videos, characterized in that, The method includes: Acquire first data sent by the target object, and determine first emotional data of the target object based on the first data; the first data includes text content and / or audio data; the first data is data formed by questions raised by the target object in response to different needs; Based on the first data and the first sentiment data, second data is determined, and based on the second data, a first duration is determined; the second data is the response text to the question in the first data; the second data includes second sentiment data guiding the digital human to respond to the first data; the first duration is used to describe the duration used by the digital human to express the second data; Based on the second emotion data in the second data, third data is determined, and first video data is formed based on the third data; the third data is used to describe the expressions and actions of the digital human in expressing the second data; the first video data is a video of the digital human completely displaying a set of third data; Based on the second data, the first duration, and the first video data, second video data is determined so that the digital human can respond and display based on the second video data; the second video data is a video of the digital human expressing the second data.
2. The method according to claim 1, characterized in that, The first data is audio data. Based on the first data, the first emotional data of the target object is determined, including: The audio data is extracted based on a speech recognition model to obtain the first text data. Obtain intonation information from the audio data, and determine first reference sentiment data based on the intonation information; Based on a text recognition model, the first text data is processed for sentiment recognition to determine the second reference sentiment data; the text recognition model is used to describe the mapping relationship between text features and sentiment data. The first reference sentiment data and the second reference sentiment data are weighted and fused to obtain the first sentiment data of the target object.
3. The method according to claim 2, characterized in that, Based on the first data and the first sentiment data, the second data is determined, including: Obtain voiceprint feature information from the audio data, and determine the age characteristics of the target object based on the voiceprint feature information; Based on the language big data model, the age features, and the first sentiment data, the second data is determined; the language big data model includes a sentiment-recitation mapping layer; the sentiment-recitation mapping layer is used to describe the mapping relationship between different sentiment data and response text.
4. The method according to claim 3, characterized in that, The first duration is determined based on the second data, including: The response speed of the digital human is obtained, and the first duration is determined based on the response speed and the number of words in the second data.
5. The method according to claim 1, characterized in that, The second sentiment data includes sentiment values from different sentiment data. Based on the second sentiment data in the second data, the third data is determined, including: Based on the first mapping rule and the second emotion data, reference expressions and reference actions of the digital human are determined; the first mapping rule is used to describe the mapping relationship between emotion data and expressions and actions; Based on the second mapping rule and the sentiment values in the second sentiment data, the expression amplitude of the reference facial expression and the motion amplitude of the reference action are determined; the second mapping rule uses a non-linear mapping function to describe the relationship between the sentiment value and the motion intensity of the sentiment data; The reference expression, the reference action, the amplitude of the reference expression, and the amplitude of the reference action are determined as the third data.
6. The method according to claim 1, characterized in that, Based on the second data, the first duration, and the first video data, the second video data is determined, including: The video period is determined based on the first duration and the second duration of the first video data; The first video data is reversed to obtain the third video data; The second video data is determined based on the second data, the first video data, the third video data, and the video period.
7. The method according to claim 1, characterized in that, After determining the second video data, the method further includes: The end time point for generating the second video data is determined, as well as the dynamic buffer time; the dynamic buffer time is used to adjust the insertion time point of the second video data. The video insertion time point is determined based on the end time point and the dynamic buffer time. The video stitching strategy is determined based on the device type of the target object. Based on the video stitching strategy, the second video data is inserted at the video insertion time point of the preset video. The preset video is a video in which the digital human displays preset expressions and preset actions when it does not perform a response task.
8. A digital human interactive video synthesis device, characterized in that, The device includes: The first information processing module is used to acquire first data sent by the target object and determine the first emotional data of the target object based on the first data; the first data includes text content and / or audio data; the first data is data formed by the questions raised by the target object in response to different needs. The second information processing module is used to determine second data based on the first data and the first emotional data, and to determine a first duration based on the second data; the second data is a response text to the question in the first data; the second data includes second emotional data guiding the digital human to respond to the first data; the first duration is used to describe the duration used by the digital human to express the second data; The first video data determination module is used to determine third data based on the second emotional data in the second data, and to form first video data based on the third data; the third data is used to describe the expressions and actions of the digital human in expressing the second data; the first video data is a video of the digital human completely displaying a set of third data; The second video data determination module is used to determine second video data based on the second data, the first duration, and the first video data, so that the digital human can respond and display based on the second video data; the second video data is a video of the digital human expressing the second data.
9. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the digital human interactive video synthesis method according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that, when executed by a processor, implement the digital human interactive video synthesis method according to any one of claims 1-7.