Real-time interactive synthesis processing method and system based on duplicated digital human
By obtaining target historical figures, users and equipment data and synthesizing and updating digital people, the problem that digital people cannot be personalized in the existing technology is solved, the digital life generation and interaction efficiency is improved, and user satisfaction is improved.
Patent Information
- Application Number
- CN202510779729.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-12
AI Technical Summary
Existing digital human creations cannot meet users' personalized needs, lack of flexibility in interaction, and it is difficult to adjust according to users' real-time needs, resulting in low human-computer interaction efficiency.
A real-time interactive synthesis processing method based on replica digital people is provided. By obtaining data from target historical figures, users and devices, synthesizing initial replica digital people, and updating in response to user adjustment requests, to generate target replica digital people that meet user needs.
It has achieved the personalization and interaction efficiency of digital human generation, improved the user's satisfaction with virtual digital people, and enhanced the human-computer interaction experience.
Smart Images

Figure CN120336572A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a real-time interactive synthesis processing method and system based on replicated digital humans. Background Art
[0002] With the development of technology, digital human technology has great application potential in fields such as culture and entertainment. In terms of cultural dissemination, digital humans that reproduce historical figures can promote cultural inheritance. However, most existing digital human creations are in a general mode and cannot meet the personalized needs of users.
[0003] However, the current interaction of digital humans lacks flexibility. Once created, it is difficult to adjust according to the real-time needs of users, and users need to regenerate digital humans, resulting in low efficiency of human-computer interaction. Summary of the Invention
[0004] The embodiments of this application provide a real-time interactive synthesis processing method and system based on replicated digital humans, which can generate digital humans that better meet the needs of users and improve the efficiency of human-computer interaction. The technical solutions are as follows: On the one hand, a real-time interactive synthesis processing method based on replicated digital humans is provided. The method includes: In response to a replication request from a target user for a target historical figure, obtain the first historical figure data of the target historical figure, the first user data of the target user, and the device performance data of the device that sent the replication request. The first historical figure data includes the life information and image records of the target historical figure, and the first user data includes the user image data and user action data of the target user; Based on the first historical figure data of the target historical figure, the first user data, and the device performance data, synthesize and display an initial replicated digital human; In response to an adjustment request from the target user for the initial replicated digital human, obtain the second user data of the target user, the device environment data of the sending device, and the second historical figure data of the target historical figure. The second user data includes the physiological data and adjustment suggestion data of the target user, and the second historical figure data includes the personality description and public evaluation of the target historical figure; Based on the second user data, the device environment data, and the second historical figure data, update the initial replicated digital human to obtain a target replicated digital human.
[0005] On the one hand, a real-time interactive synthesis processing device based on replicated digital humans is provided. The device includes: An acquisition module, configured to acquire first historical figure data of the target historical figure, first user data of the target user, and device performance data of the sending device of the reproduction request in response to a reproduction request of the target user for the target historical figure, where the first historical figure data includes life information and image records of the target historical figure, and the first user data includes user image data and user action data of the target user; A synthesis and display module, configured to synthesize and display an initial reproduced digital human based on the first historical figure data of the target historical figure, the first user data, and the device performance data; The acquisition module is further configured to acquire second user data of the target user, device environment data of the sending device, and second historical figure data of the target historical figure in response to an adjustment request of the target user for the initial reproduced digital human, where the second user data includes physiological data and adjustment suggestion data of the target user, and the second historical figure data includes character descriptions and public evaluations of the target historical figure; An update module, configured to update the initial reproduced digital human based on the second user data, the device environment data, and the second historical figure data to obtain a target reproduced digital human.
[0006] On the one hand, a computer device is provided, where the computer device includes one or more processors and one or more memories, and at least one computer program is stored in the one or more memories, and the computer program is loaded and executed by the one or more processors to implement the real-time interaction synthesis processing method based on the reproduced digital human.
[0007] On the one hand, a computer-readable storage medium is provided, where at least one computer program is stored in the computer-readable storage medium, and the computer program is loaded and executed by a processor to implement the real-time interaction synthesis processing method based on the reproduced digital human.
[0008] On the one hand, a computer program product or a computer program is provided, where the computer program product or the computer program includes program code, the program code is stored in a computer-readable storage medium, a processor of a computer device reads the program code from the computer-readable storage medium, and the processor executes the program code, so that the computer device executes the above-mentioned real-time interaction synthesis processing method based on the reproduced digital human.
[0009] Through the technical solution provided by the embodiments of the present application, in response to a reproduction request of a target user for a target historical figure, the first historical figure data of the target historical figure, the first user data of the target user, and the device performance data of the sending device of the reproduction request are obtained. Based on the first historical figure data, the first user data, and the device performance data, an initial reproduced digital human is synthesized and displayed, thereby realizing the preliminary generation and display of the reproduced digital human. In response to an adjustment request of the target user for the initial reproduced digital human, the second user data of the target user, the device environment data of the sending device, and the second historical figure data of the target historical figure are obtained. Based on the second user data, the device environment data, and the second historical figure data, the initial reproduced digital human is updated to obtain a target reproduced digital human, realizing convenient update of the reproduced digital human. Subsequently, the target reproduced digital human can be used to interact with the target user, improving the satisfaction of the target user with the target virtual digital human, thereby improving the user experience. Brief Description of the Drawings
[0010] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings. Figure 1 It is a schematic diagram of the implementation environment of a real-time interactive synthesis processing method based on a reproduced digital human provided by the embodiments of the present application; Figure 2 It is a flowchart of a real-time interactive synthesis processing method based on a reproduced digital human provided by the embodiments of the present application; Figure 3 It is a flowchart of another real-time interactive synthesis processing method based on a reproduced digital human provided by the embodiments of the present application; Figure 4 It is a schematic diagram of the structure of a real-time interactive synthesis processing device based on a reproduced digital human provided by the embodiments of the present application; Figure 5 It is a schematic diagram of the structure of a server provided by the embodiments of the present application. Detailed Description of the Embodiments
[0011] To make the purpose, technical solutions, and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail in conjunction with the drawings.
[0012] In the present application, terms such as "first" and "second" are used to distinguish the same items or similar items with basically the same functions. It should be understood that there is no logical or chronological dependency between "first", "second", and "nth", nor are the quantity and execution order limited.
[0013] Digital Human: A digital human is a virtual human image generated using computer technology. It combines information science and life science and aims to conduct virtual simulations of the human body at different levels. As an emerging technological product, a digital human has multiple human characteristics, including appearance features, human performance capabilities, and interaction capabilities, etc.
[0014] Replicated Digital Human: A replicated digital human refers to a virtual image with interactive capabilities created by simulating the appearance and movements of real humans through computer technology.
[0015] Augmented Reality: Augmented reality is a technology that integrates digital information into the real world. Through a device with a camera function, users can interact with both their physical space and computer-generated content simultaneously.
[0016] Virtual Reality: Virtual reality is a technology that creates a realistic three-dimensional simulated environment where users can interact with the virtual environment through VR devices (such as helmets and controllers).
[0017] Metaverse: The metaverse is a virtual shared space, usually composed of multiple virtual reality worlds or games, where users can conduct activities such as social interaction, asset trading, and virtual tourism, etc.
[0018] In the embodiments of this application, digital humans can be applied in augmented reality, virtual reality, and the metaverse.
[0019] Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.
[0020] Machine Learning (ML) is a multi-disciplinary field that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills, and reorganize existing knowledge sub-models to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning from demonstration.
[0021] Semantic feature: A feature used to represent the semantics expressed by text. Different texts can correspond to the same semantic feature. For example, the texts "What's the weather like today" and "How's the weather today" can correspond to the same semantic feature. A computer device can map the characters in a text to character vectors, and based on the relationships between the characters, combine and operate on the character vectors to obtain the semantic feature of the text. For example, the computer device can adopt the Bidirectional Encoder Representations from Transformers (BERT).
[0022] Normalization: Maps a sequence of numbers with different value ranges to the interval (0, 1) for convenient data processing. In some cases, the normalized values can be directly implemented as probabilities.
[0023] Embedded Coding: Embedded Coding represents a correspondence in mathematics, that is, mapping the data in the X space to the Y space through a function F, where the function F is an injective function, and the result of the mapping is structure-preserving. The injective function means that the data after mapping corresponds uniquely to the data before mapping, and structure-preserving means that the size relationship of the data before mapping is the same as that of the data after mapping. For example, there are data X1 and X2 before mapping, and Y1 corresponding to X1 and Y2 corresponding to X2 are obtained after mapping. If the data X1 > X2 before mapping, then correspondingly, the data Y1 after mapping is greater than Y2. For words, it is to map the words to another space for subsequent machine learning and processing.
[0024] Attention weight: Can represent the importance of a certain data during the training or prediction process. Importance indicates the magnitude of the influence of the input data on the output data. Data with high importance has a higher corresponding attention weight value, and data with low importance has a lower corresponding attention weight value. In different scenarios, the importance of data is not the same, and the process of training the attention weight of the model is also the process of determining the importance of data.
[0025] It should be noted that the information involved in this application (including but not limited to user device information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.), and signals are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions. For example, the first user data and the second user data involved in this application are obtained under full authorization.
[0026] Figure 1It is a schematic diagram of the implementation environment of a real-time interaction synthesis processing method based on a replicated digital human provided by an embodiment of the present application. Refer to Figure 1 In this implementation environment, a terminal 110 and a server 140 may be included.
[0027] The terminal 110 is connected to the server 140 through a wireless network or a wired network. Optionally, the terminal 110 is a smart phone, a tablet computer, a laptop computer, a desktop computer, etc., but is not limited thereto. An application program that supports real-time interaction synthesis processing based on a replicated digital human is installed and run on the terminal 110.
[0028] The server 140 is an independent physical server, or a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. The server 140 provides background services for the application program running on the terminal 110.
[0029] Those skilled in the art can know that the number of the above terminals may be more or less. For example, there is only one of the above terminals, or there are dozens or hundreds of the above terminals, or a larger number. In this case, other terminals are also included in the above implementation environment. The embodiments of the present application do not limit the number and device types of the terminals.
[0030] After introducing the implementation environment of the embodiments of the present application, the application scenarios of the embodiments of the present application will be introduced below. In the scenario of creating a replicated digital human, the technical solution provided by the embodiments of the present application enables a user to use the technical solution provided by the embodiments of the present application to create a desired digital human, and then be able to interact with the created digital human.
[0031] After introducing the implementation environment and application scenarios of the embodiments of the present application, the technical solution provided by the embodiments of the present application will be described below. Figure 2 It is a flowchart of a real-time interaction synthesis processing method based on a replicated digital human provided by an embodiment of the present application. Refer to Figure 2 Taking the server as the execution subject, the method includes the following steps.
[0032] 201. In response to a replication request of a target user for a target historical figure, the server obtains first historical figure data of the target historical figure, first user data of the target user, and device performance data of the sending device of the replication request. The first historical figure data includes biographical information and image records of the target historical figure, and the first user data includes user image data and user action data of the target user.
[0033] Among them, the target user is the user who uses the services related to the virtual digital human. The replicated digital human is a virtual character image generated by using computer technology, and different replicated digital humans have different appearances and actions. The target historical figure is the historical figure selected by the target user and is the object of digital human replication. Correspondingly, the replication request is used to request the replication of the target historical figure to generate the replicated digital human corresponding to the target historical figure. The sending device is the computer device used by the target user. After the replicated virtual object is generated, it will be displayed through the sending device. The device performance data of the sending device is used to represent the strength of the rendering ability of the sending device. Since real-time rendering by the sending device is required when displaying the replicated virtual human, the stronger the rendering ability, the higher the rendering fineness and action complexity of the virtual human that can be rendered. The life information and image records are all obtained from the historical figure database, and the historical figure database stores the life information and image records of multiple candidate historical figures, and the target historical figure belongs to the multiple candidate historical figures. The user image data includes the image data related to the target user, and the user action data includes the action data related to the target user.
[0034] 202. The server synthesizes and displays the initial replicated digital human based on the first historical figure data of the target historical figure, the first user data, and the device performance data.
[0035] Among them, the initial replicated digital human is a digital human generated by using the first historical figure data, the first user data, and the device performance data. The initial replicated digital human is relatively matched with the target historical figure, the target user, and the sending device. Displaying the initial replicated digital human means displaying the initial replicated digital human and controlling the initial replicated digital human to speak and perform actions. Since the server cannot directly display the initial replicated digital human, the above display is actually that the server displays the initial replicated digital human through the sending device, and the target user can view the initial replicated digital human through the initial replicated digital human.
[0036] 203. In response to the adjustment request of the target user for the initial replicated digital human, the server obtains the second user data of the target user, the device environment data of the sending device, and the second historical figure data of the target historical figure. The second user data includes the physiological data and adjustment suggestion data of the target user, and the second historical figure data includes the personality description and public evaluation of the target historical figure.
[0037] Among them, the adjustment request is used to request an adjustment to the initial replicated digital human. The appearance of the adjustment request indicates that the target user is not satisfied with the initial replicated digital human. At this time, other data will be obtained to adjust the initial replicated digital human. The personality description and public evaluation of the target historical figure are also stored in the historical figure database and can be directly obtained from the historical figure database. The physiological data of the target user is used to represent the physiological state of the target user. The adjustment suggestion data is carried by the adjustment request. That is, when sending the adjustment request, the adjustment suggestion data will be carried. The adjustment suggestion data is used to indicate the suggestions for the target user to adjust the initial replicated digital human. The device environment data is used to represent the environmental conditions of the sending device.
[0038] 204. The server updates the initial replicated digital human based on the second user data, the device environment data, and the second historical figure data to obtain a target replicated digital human.
[0039] Among them, the target replicated digital human is the updated replicated digital human, and the target user can interact with the target replicated digital human subsequently.
[0040] Through the technical solution provided by the embodiments of the present application, in response to the replication request of the target user for the target historical figure, the first historical figure data of the target historical figure, the first user data of the target user, and the device performance data of the sending device of the replication request are obtained. Based on the first historical figure data, the first user data, and the device performance data, an initial replicated digital human is synthesized and displayed, thereby realizing the preliminary generation and display of the replicated digital human. In response to the adjustment request of the target user for the initial replicated digital human, the second user data of the target user, the device environment data of the sending device, and the second historical figure data of the target historical figure are obtained. The initial replicated digital human is updated based on the second user data, the device environment data, and the second historical figure data to obtain a target replicated digital human, realizing convenient update of the replicated digital human. Subsequently, the target replicated digital human can be used to interact with the target user, improving the satisfaction of the target user with the target virtual digital human, thereby improving the user experience.
[0041] The above steps 201-204 are a brief introduction to the real-time interaction synthesis processing method based on replicated digital humans provided by the embodiments of the present application. Below, some examples will be combined to more clearly illustrate the real-time interaction synthesis processing method based on replicated digital humans provided by the embodiments of the present application. See Figure 3 , taking the execution subject as the server as an example, the method includes the following steps.
[0042] 301. In response to a reproduction request from a target user for a target historical figure, the server obtains the first historical figure data of the target historical figure, the first user data of the target user, and the device performance data of the sending device of the reproduction request. The first historical figure data includes the life information and image records of the target historical figure, and the first user data includes the user image data and user action data of the target user.
[0043] Among them, the target user is a user who uses the virtual digital human-related services. The reproduced digital human is a virtual human image generated by computer technology, and different reproduced digital humans have different appearances and actions. The target historical figure is the historical figure selected by the target user and is the object of digital human reproduction. Correspondingly, the reproduction request is used to request the reproduction of the target historical figure to generate the reproduced digital human corresponding to the target historical figure. The sending device is the computer device used by the target user, and after the reproduced virtual object is generated, it will be displayed through the sending device. The device performance data of the sending device is used to represent the strength of the rendering ability of the sending device. Since real-time rendering by the sending device is required when displaying the reproduced virtual human, the stronger the rendering ability, the higher the rendering fineness and action complexity of the virtual human that can be rendered. The life information and image records are both obtained from the historical figure database, and the historical figure database stores the life information and image records of multiple candidate historical figures, and the target historical figure belongs to the multiple candidate historical figures. The user image data includes the image data related to the target user, and the user action data includes the action data related to the target user.
[0044] In a possible implementation manner, in response to a reproduction request from a target user for a target historical figure, the server queries in the historical figure database based on the character identifier of the target historical figure to obtain the life information and image records of the target historical figure. The server parses the reproduction request to obtain the user image data, user action data of the target user, and the device performance data of the sending device.
[0045] Among them, the life information is in text form, and the image records include image pictures and image description texts. The user image data includes the facial image of the target user and the expected image description text for the target historical figure. The user action data includes the current user action data and the historical user action data.
[0046] 302. The server synthesizes and displays the initial reproduced digital human based on the first historical figure data of the target historical figure, the first user data, and the device performance data.
[0047] Among them, the initial replicated digital human is a digital human generated using the first historical figure data, the first user data, and the device performance data. The initial replicated digital human is relatively matched with the target historical figure, the target user, and the sending device. Displaying the initial replicated digital human means showing the initial replicated digital human and controlling the initial replicated digital human to speak and perform actions. Since the server cannot directly display the initial replicated digital human, the above display is actually the server displaying the initial replicated digital human through the sending device, and the target user can view the initial replicated digital human through the initial replicated digital human.
[0048] In a possible implementation, the server determines the first replicated digital human image data based on the image record of the target historical figure and the user image data of the target user. The server determines the first replicated digital human action data based on the life information and the user action data. The server determines the second replicated digital human image data based on the device performance data, the life information, and the first replicated digital human image data. The server generates the initial replicated digital human based on the first replicated digital human action data and the second replicated digital human image data. The server displays the initial replicated digital human.
[0049] Among them, the first replicated digital human image data is the replicated digital human image data obtained by combining the image record of the target historical figure and the user image data of the target user. It is equivalent to using the image record of the target historical figure as the benchmark and the user image data for correction, and is relatively matched with the image record of the target historical figure and the user image data of the target user. The first replicated digital human action data is combined with the life information of the target historical figure and the user action data of the target user. It is equivalent to using the life information as the benchmark and the user action data for correction, and is relatively matched with the life information of the target historical figure and the user action data of the target user. The process of obtaining the first replicated digital human image data and the first replicated digital human action data not only considers the basic image record and life information of the target historical figure, but also combines the user image data and user action data related to the target user, so that the subsequent generated replicated digital human can not only restore the target historical figure, but also combine the relevant information of the target user. The second replicated digital human image data combines the device performance data, the life information, and the first replicated digital human image data, which is equivalent to using the device performance data and the life information to correct the first replicated digital human image data. The initial replicated digital human is generated by combining the first replicated digital human action data and the second replicated digital human image data. It is equivalent that the image of the initial replicated digital human is generated based on the second replicated digital human image data, and the action is generated based on the first replicated digital human action data. The server displaying the initial replicated digital human is the server displaying the initial replicated digital human through the sending device.
[0050] To illustrate the above - mentioned embodiments more clearly, the above - mentioned embodiments will be described in several parts below.
[0051] Part 1: The server determines the first replicated digital human image data based on the image record of the target historical figure and the user image data of the target user.
[0052] In a possible implementation, the image record includes an image and an image description text, and the user image data includes a first user facial image, a second user facial image, and an expected image description text of the target historical figure. The second user facial image is obtained by the target user editing the first user facial image. The server determines the first reference image data of the target historical figure based on the image, the image description text, and the expected image description text. The server determines the second reference image data of the target user based on the first user facial image and the second user facial image. The server determines the first replicated digital human image data based on the first reference image data and the second reference image data.
[0053] Among them, both the first user facial image and the second user facial image are provided by the target user. The first user facial image is a user facial image directly captured using an image acquisition device, and the second user facial image is a user facial image obtained by the target user editing the first user facial image. The change from the first user facial image to the second user facial image can reflect the target user's preference for the facial image, thus guiding the generation of the replicated digital human's image. The expected image description text is provided by the target user and is used to reflect the target user's expectation for the image of the target historical figure. In the embodiments of this application, the expected image description text is used to assist in the generation of the replicated digital human. The first reference image data is generated by combining the image, the image description text, and the expected image description text, and can reflect the basic image of the target historical figure and the target user's expectation for the image of the target historical figure. The second reference image data is determined based on the first user facial image and the second user facial image, and can reflect the target user's preference for the facial image.
[0054] In the above - mentioned embodiments, the first reference image data of the target historical figure is determined using the image, the image description text, and the expected image description text, and the second reference image data of the target user is determined using the first user facial image and the second user facial image. Combining the first reference image data and the second reference image data to determine the first replicated digital human image data has relatively high accuracy.
[0055] To illustrate the above - mentioned embodiments more clearly, the above - mentioned embodiments will be described in several parts below again.
[0056] A server determines first reference image data of the target historical figure based on the image, the image description text, and the expected image description text.
[0057] In a possible implementation, the server determines basic image data of the target historical figure based on the image and the image description text. The server determines expected image data of the target historical figure based on the image and the expected image description text. The server determines the first reference image data of the target historical figure based on the basic image data and the expected image data of the target historical figure.
[0058] For example, the server respectively extracts features from the image and the image description text to obtain the image features of the image and the image description text features of the image description text. The server fuses the image features and the image description text features to obtain basic image features. The server decodes the basic image features to obtain the basic image data. The server respectively extracts features from the image and the expected image description text to obtain the image features of the image and the expected image description text features of the expected image description text. The server fuses the image features and the expected image description text features to obtain expected image features. The server decodes the expected image features to obtain the expected image data. The server performs weighted fusion on the basic image data and the expected image data of the target historical figure to obtain the first reference image data of the target historical figure.
[0059] Among them, the weights for weighted fusion are set by technicians according to the actual situation. Generally speaking, the weight corresponding to the basic image data is higher than the weight corresponding to the expected image data, so that the first reference image data can be more biased towards the basic image data, that is, more of the original image of the target historical figure is retained.
[0060] For example, the server encodes the image of the figure and the text description of the figure respectively based on the attention mechanism, and obtains the figure image features of the image of the figure and the figure description text features of the text description of the figure. The server fuses the figure image features and the figure description text features to obtain the basic figure features. The server performs multiple rounds of iterative decoding on the basic figure features based on the attention mechanism to obtain the basic figure data. The server encodes the image of the figure and the expected figure description text respectively based on the attention mechanism, and obtains the figure image features of the image of the figure and the expected figure description text features of the expected figure description text. The server fuses the figure image features and the expected figure description text features to obtain the expected figure features. The server performs multiple rounds of iterative decoding on the expected figure features based on the attention mechanism to obtain the expected figure data. The server performs weighted fusion on the basic figure data and the expected figure data of the target historical figure to obtain the first reference figure data of the target historical figure.
[0061] Among them, the above encoding and decoding processes are semantic encoding and semantic decoding processes, which are equivalent to the process of data mining and fusion.
[0062] B. The server determines the second reference figure data of the target user based on the first user's facial image and the second user's facial image.
[0063] In a possible implementation manner, the server determines a facial image difference description text for the change from the first user's facial image to the second user's facial image based on the first user's facial image and the second user's facial image. The server determines the second reference figure data of the target user based on the facial image difference description text and the second user's facial image.
[0064] Among them, the facial image difference description text is used to represent the image difference between the first user's facial image and the second user's facial image, so as to reflect the target user's preference for the facial image.
[0065] For example, the server extracts features from the first user's facial image and the second user's facial image respectively, and obtains the first facial features of the first user's facial image and the second facial features of the second user's facial image. The server determines the facial image difference description text based on the first facial features and the second facial features. The server fuses the facial image difference description text and the second user's facial image to obtain the second reference figure data of the target user.
[0066] For example, the server encodes the first user's facial image and the second user's facial image respectively based on the attention mechanism, and obtains the first facial feature of the first user's facial image and the second facial feature of the second user's facial image. The server determines the facial difference feature between the first facial feature and the second facial feature. The server performs multiple rounds of iterative decoding on the facial difference feature based on the attention mechanism, and obtains the facial image difference description text. The server extracts features from the facial image difference description text respectively, and obtains the difference description text feature of the facial image difference description text. The server fuses the difference description text feature and the second facial feature of the second user's facial image, and obtains the second reference image feature. The server performs multiple rounds of iterative decoding on the second reference image feature based on the attention mechanism, and obtains the second reference image data of the target user.
[0067] C. The server determines the first replicated digital human image data based on the first reference image data and the second reference image data.
[0068] In a possible implementation manner, the server encodes the first reference image data and the second reference image data based on the attention mechanism, and obtains the first reference image data feature of the first reference image data and the second reference image data feature of the second reference image data. The server fuses the first reference image data feature and the second reference image data feature, and obtains the first replicated digital human image feature. The server performs multiple rounds of iterative decoding on the first replicated digital human image feature based on the attention mechanism, and obtains the first replicated digital human image data.
[0069] Wherein, the first replicated digital human image data is in text form or image form, and the embodiments of the present application do not make any limitation thereto.
[0070] It should be noted that the above feature extraction processes are all implemented by a multi-modal feature extractor based on the attention mechanism, so that data of different modalities can be mapped to the same feature space, which is convenient for subsequent processing.
[0071] Second part: The server determines the first replicated digital human action data based on the life information and the user action data.
[0072] In a possible implementation manner, the user action data includes current user action data and historical user action data. The server performs data recognition on the life information, and obtains the action description data of the target historical figure in the life information. The server determines the reference action data based on the action description data and the current user action data. The server determines the first replicated digital human action data based on the historical user action data and the reference action data.
[0073] Among them, the current user action data is the data generated by capturing the actions of the target user during the current reproduction digital life generation cycle. The sending device will display action prompt text. After the action prompt text is displayed, the actions performed by the target user will be captured, thereby generating the current action data. The historical action data is the data generated by historically capturing the actions of the target user. The capture of user action data is achieved through an action capture device. The action data is used to guide the action execution of the generated reproduced virtual human.
[0074] To illustrate the above embodiments more clearly, the above embodiments will be described in several parts below.
[0075] A. The server performs data recognition on the life information to obtain the action description data of the target historical figure in the life information.
[0076] In a possible implementation manner, the server extracts features from the life information to obtain the life information features of the life information. The server determines the action description data from the life information based on the life information features.
[0077] For example, the server divides the life information into multiple paragraphs. The server encodes the multiple paragraphs based on the attention mechanism to obtain the paragraph features of each paragraph. The paragraph features of the multiple paragraphs form the life information features of the life information. The server classifies each paragraph based on the paragraph features of each paragraph to obtain the paragraph type of each paragraph. The server combines the paragraphs with the paragraph type of action description in the multiple paragraphs to obtain the action description data.
[0078] For instance, the server divides the life information into multiple paragraphs according to punctuation marks. The server encodes the multiple paragraphs based on the attention mechanism to obtain the paragraph features of each paragraph. The server performs a fully connected and normalization operation on the paragraph features of each paragraph to obtain a probability set for each paragraph. The probability set includes multiple probabilities, and one probability corresponds to a candidate paragraph type. For any paragraph in the multiple paragraphs, the server determines the candidate paragraph type corresponding to the highest probability in the probability set of the paragraph as the paragraph type of the paragraph. The server combines the paragraphs with the paragraph type of action description in the multiple paragraphs to obtain the action description data.
[0079] B. The server determines reference action data based on the action description data and the current user action data.
[0080] In a possible implementation, the server respectively extracts features from the action description data and the current user action data to obtain the action description data features of the action description data and the current user action features of the current user action data. The server fuses the action description data features and the current user action features to obtain reference action features. The server decodes the reference action features to obtain the reference action data.
[0081] Among them, the reference action data is in text form or implicit feature form, which is not limited in the embodiments of the present application. The implicit feature is similar to the definition of the implicit feature h in the long short-term memory network (LSTM).
[0082] For example, the server encodes the action description data and the current user action data respectively based on the attention mechanism to obtain the action description data features of the action description data and the current user action features of the current user action data. The server performs weighted fusion on the action description data features and the current user action features to obtain reference action features. The server performs multiple rounds of iterative decoding on the reference action features based on the attention mechanism to obtain the reference action data.
[0083] Among them, the weights of the weighted fusion are set by those skilled in the art according to the actual situation, which is not limited in the embodiments of the present application. The above encoding and decoding processes are semantic encoding and semantic decoding processes, which are equivalent to the process of data mining and fusion.
[0084] C. The server determines the first replicated digital human action data based on the historical user action data and the reference action data.
[0085] In a possible implementation, the server splices the historical user action data and the reference action data to obtain the first replicated digital human action data.
[0086] Among them, splicing is to completely retain the data.
[0087] Part Three: The server determines the second replicated digital human image data based on the device performance data, the life information, and the first replicated digital human image data.
[0088] In a possible implementation, the server extracts features from the life information to obtain the life information features of the life information. The server determines the third replicated digital human image data based on the life information features and the first replicated digital human image data. The server determines the second replicated digital human image data based on the device performance data and the third replicated digital human image data.
[0089] For example, the server encodes the life data based on the attention mechanism to obtain the life data features of the life data. The server encodes the first replicated digital human image data based on the attention mechanism to obtain the first replicated digital human image data features of the first replicated digital human image data. The server fuses the life data features and the first replicated digital human image data features to obtain the third replicated digital human image data features. The server decodes the third replicated digital human image data features to obtain the third replicated digital human image data. The server determines the upper limit information of the digital human rendering effect based on the device performance data. The server determines the second replicated digital human image data based on the upper limit information of the digital human rendering effect and the third replicated digital human image data.
[0090] Among them, the upper limit information of the digital human rendering effect represents the maximum digital human rendering ability of the sending device, and the upper limit information of the digital human rendering effect includes the upper limit of the rendering resolution, the upper limit of the rendering sampling number, the upper limit of reflection and refraction, etc.
[0091] The following describes the method by which the server determines the upper limit information of the digital human rendering effect based on the device performance data in the above example.
[0092] In some embodiments, the server performs multiple fully connected operations on the device performance parameters to obtain the device performance features of the device performance parameters. The server performs a fully connected operation and normalization on the device performance features to obtain the upper limit information of the digital human rendering effect.
[0093] The following describes the method by which the server determines the second replicated digital human image data based on the upper limit information of the digital human rendering effect and the third replicated digital human image data in the above example.
[0094] In some embodiments, the server fuses the upper limit information of the digital human rendering effect and the third replicated digital human image data to obtain the second replicated digital human image data.
[0095] Part Four: The server generates the initial replicated digital human based on the first replicated digital human action data and the second replicated digital human image data.
[0096] In a possible implementation manner, the server generates a plurality of skeleton points based on the second replicated digital human image data and renders the surface formed by the plurality of skeleton points to obtain the digital human virtual image of the initial replicated digital human. The server generates the digital human action mode of the initial replicated digital human based on the first replicated digital human action data, and the digital human action mode is used to constrain the movement mode of the plurality of skeleton points.
[0097] Among them, the digital human virtual image is a visual three-dimensional model, which is a skeletal skin model. The skeleton refers to the multiple skeletal points, and the skin refers to the surface formed by the multiple skeletal points.
[0098] For example, the server encodes the second replicated digital human image data based on the attention mechanism to obtain the second replicated digital human image data features. The server performs multiple rounds of iterative decoding on the second replicated digital human image data features to obtain the positions of multiple skeletal points and rendering parameters. The server creates an initial three-dimensional model based on the positions of the multiple skeletal points, and renders the surface formed by the multiple skeletal points using the rendering parameters to obtain the digital human virtual image of the initial replicated digital human. The server encodes the first replicated digital human action data and the positions of the multiple skeletal points based on the attention mechanism to obtain the first replicated digital human action data features of the first replicated digital human action data and the position features of the multiple skeletal points. The server decodes the first replicated digital human action data features and the position features of the multiple skeletal points based on the attention mechanism to obtain the digital human action mode of the initial replicated digital human.
[0099] Part Five: The server displays the initial replicated digital human.
[0100] In a possible implementation manner, the server sends the initial replicated digital human display data of the initial replicated digital human to the sending device, so that the sending device displays the initial replicated digital human based on the initial replicated digital human display data.
[0101] Among them, the initial replicated digital human display data includes the digital human virtual image and the digital human action mode of the initial replicated digital human.
[0102] 303. In response to the adjustment request of the target user for the initial replicated digital human, the server obtains the second user data of the target user, the device environment data of the sending device, and the second historical figure data of the target historical figure. The second user data includes the physiological data and adjustment suggestion data of the target user, and the second historical figure data includes the personality description and public evaluation of the target historical figure.
[0103] Among them, the adjustment request is used to request an adjustment to the initial replicated digital human. The occurrence of this adjustment request indicates that the target user is not satisfied with the initial replicated digital human. At this time, other data will be obtained to adjust the initial replicated digital human. The personality description and public evaluation of the target historical figure are also stored in the historical figure database and can be directly obtained from this historical figure database. The physiological data of the target user is used to represent the physiological state of the target user. The adjustment suggestion data is carried by this adjustment request, that is, when sending this adjustment request, this adjustment suggestion data will be carried. This adjustment suggestion data is used to indicate suggestions for the target user to adjust the initial replicated digital human. The device environment data is used to represent the environmental conditions of the sending device.
[0104] In a possible implementation manner, in response to the adjustment request of the target user for the initial replicated digital human, the server obtains the second user data and the device environment data from this adjustment request. The server queries in the historical figure database based on the person identifier of the target historical figure to obtain the second historical figure data of the target historical figure.
[0105] Among them, this historical figure database is a database maintained by the server correspondingly.
[0106] 304. The server updates the initial replicated digital human based on the second user data, the device environment data, and the second historical figure data to obtain the target replicated digital human.
[0107] Among them, the target replicated digital human is the replicated digital human after update, and the subsequent target user can interact with this target replicated digital human.
[0108] In a possible implementation manner, the server determines the subjective feeling data of the target user for the initial replicated digital human based on the physiological data and the adjustment suggestion data. The server determines the first adjustment data for the initial replicated digital human based on the adjustment suggestion data and the device environment data, and the device environment data is used to describe the environment where the sending device is located. The server determines the second adjustment data for the initial replicated digital human based on the personality description and the public evaluation. The server determines the replicated digital human update data based on the subjective feeling data, the first adjustment data, and the second adjustment data. The server uses this replicated digital human update data to update the initial replicated digital human to obtain this target replicated digital human.
[0109] Among them, combining the device environment data is to make the displayed replicated digital human match the environment where the sending device is located. Combining the adjustment suggestion data is to adjust the replicated digital human in the way indicated by the user. Combining the personality description data and the public evaluation is to enrich the image of the replicated digital human.
[0110] To illustrate the above embodiments more clearly, the above embodiments will be described in several parts below.
[0111] First part: The server determines the subjective feeling data of the target user for the initial replicated digital human based on the physiological data and the adjustment suggestion data.
[0112] In a possible implementation manner, the physiological data includes first physiological data and second physiological data. The first physiological data is the physiological data collected before presenting the initial replicated digital human, and the second physiological data is the physiological data collected after presenting the initial replicated digital human. The server determines the physiological state change data of the target user based on the first physiological data and the second physiological data. The server determines the satisfaction degree of the target user for the initial replicated digital human based on the adjustment suggestion data. The server determines the subjective feeling data based on the physiological state change data and the satisfaction degree.
[0113] To illustrate the above embodiments more clearly, the above embodiments will be described in several parts below again.
[0114] A. The server determines the physiological state change data of the target user based on the first physiological data and the second physiological data.
[0115] In a possible implementation manner, the server subtracts the second physiological data from the first physiological data to obtain the physiological state change data.
[0116] B. The server determines the satisfaction degree of the target user for the initial replicated digital human based on the adjustment suggestion data.
[0117] In a possible implementation manner, the server extracts features from the adjustment suggestion data to obtain the adjustment suggestion data features of the adjustment suggestion data. The server determines the satisfaction degree of the target user for the initial replicated digital human based on the adjustment suggestion data features.
[0118] For example, the server encodes the adjustment suggestion data based on the attention mechanism to obtain the adjustment suggestion data features. The server performs a fully connected operation and normalization on the adjustment suggestion data features to obtain a satisfaction degree classification value. The server determines the candidate satisfaction degree corresponding to the classification value interval to which the satisfaction degree classification value belongs as the satisfaction degree of the target user for the initial replicated digital human.
[0119] Among them, there are multiple classification value intervals, and one classification value interval corresponds to a candidate satisfaction level. The division of the multiple classification value intervals and the correspondence between the classification value intervals and the candidate satisfaction levels are set by technicians according to the actual situation, and the embodiments of the present application do not limit this. For example, the multiple candidate satisfaction levels include very dissatisfied, dissatisfied, average, satisfied, and very satisfied.
[0120] C. The server determines the subjective feeling data based on the physiological state change data and the satisfaction level.
[0121] In a possible implementation manner, the server extracts features from the physiological state change data to obtain physiological state change features. The server determines the physiological preference information of the target user for the initial replicated digital human based on the physiological state change features. The server determines the subjective feeling data based on the physiological preference information and the satisfaction level.
[0122] For example, the server performs temporal encoding on the physiological state change data to obtain the physiological state change features. The server decodes the physiological state change features to obtain the physiological preference information. The server splices the physiological preference information and the satisfaction level to obtain the subjective feeling data.
[0123] Second part: The server determines the first adjustment data for the initial replicated digital human based on the adjustment suggestion data and the device environment data, and the device environment data is used to describe the environment where the sending device is located.
[0124] In a possible implementation manner, the device environment data includes device location information and device environment light information. The server determines the first image adjustment data and the first action adjustment data of the initial replicated digital human based on the adjustment suggestion data. The server determines reference device adjustment data based on the device location information and the device environment light information, and the reference device adjustment data matches the device location information and the device environment light information. The server determines the first adjustment data based on the first image adjustment data, the first action adjustment data, and the reference device adjustment data.
[0125] To illustrate the above implementation manners more clearly, the above implementation manners will be further described in several parts below.
[0126] A. The server determines the first image adjustment data and the first action adjustment data of the initial replicated digital human based on the adjustment suggestion data.
[0127] In a possible implementation, the server encodes the adjustment suggestion data based on the attention mechanism to obtain the adjustment suggestion data features of the adjustment suggestion data. The server performs multiple rounds of iterative decoding on the adjustment suggestion data features based on the attention mechanism to obtain the first image adjustment data and the first action adjustment data.
[0128] Among them, the above encoding and decoding are implemented through a first adjustment data determination model. The first adjustment data determination model can convert the adjustment suggestion data into image adjustment data and action adjustment data. The first adjustment data determination model is trained based on multiple sample adjustment suggestion data and the corresponding labeled image adjustment data and labeled action adjustment data. The first adjustment data determination model is an encoding and decoding model based on the attention mechanism. The embodiments of the present application do not limit the structure and training method of the first adjustment data determination model.
[0129] B. The server determines the reference device adjustment data based on the device location information and the device ambient light information.
[0130] In a possible implementation, the server queries using the device location information to obtain the location adjustment data corresponding to the device location information. The server queries using the device ambient light information to obtain the ambient light adjustment data corresponding to the device ambient light information. The server determines the reference device adjustment data based on the location adjustment data and the ambient light adjustment data.
[0131] For example, the server queries in the first relationship table using the device location information to obtain the location adjustment data corresponding to the device location information. The server queries in the second relationship table using the device ambient light information to obtain the ambient light adjustment data corresponding to the device ambient light information. The server fuses the location adjustment data and the ambient light adjustment data to obtain the reference device adjustment data.
[0132] Among them, the first relationship table stores multiple device location information and position adjustment data corresponding to each device information, and the second relationship table stores multiple device ambient light information and ambient light adjustment data corresponding to each device ambient light information. The position adjustment data is the replica object adjustment data corresponding to the position device information, which is used to indicate the way to adjust the image of the replicated digital human; the ambient light adjustment data is the replica object adjustment data corresponding to the ambient light information, which is used to indicate the way to adjust the image of the replicated digital human. The reason for adjusting the initial replica object in combination with the device location information is that users in the same location have similar preferences for the replica object, so the initial replica object will be adjusted in combination with the device location information. The reason for adjusting the initial replica object in combination with the ambient light adjustment data is that the same rendering effect may have different visual effects under different ambient lights, so the initial replica object will be adjusted in combination with the light environment data. The first relationship table and the second relationship table are set by technicians according to the actual situation, and the embodiments of the present application do not limit this.
[0133] C. The server determines the first adjustment data based on the first image adjustment data, the first action adjustment data, and the reference device adjustment data.
[0134] In a possible implementation manner, the server fuses the first image adjustment data and the reference device adjustment data to obtain intermediate image adjustment data. The server splices the intermediate image adjustment data and the reference device adjustment data to obtain the first adjustment data.
[0135] Third part: The server determines the second adjustment data for the initial replicated digital human based on the personality description and the public evaluation.
[0136] In a possible implementation manner, the server determines the second action adjustment data of the initial replicated digital human based on the personality description. The server determines the third action adjustment data and the second image adjustment data of the initial replicated digital human based on the public evaluation. The server determines the second adjustment data of the initial replicated digital human based on the second action adjustment data, the third action adjustment data, and the second image adjustment data.
[0137] Among them, personality is usually associated with actions, so the second action adjustment data can be determined using the personality description. Public evaluation usually includes evaluations of the image and actions, so it can be used to adjust the image and actions.
[0138] To illustrate the above implementation manner more clearly, the above implementation manner will be further described in several parts below.
[0139] A. The server determines the second action adjustment data of the initial replicated digital human based on the personality description.
[0140] In a possible implementation, the server extracts features from the personality description to obtain the personality description features of the personality description. The server determines the second action adjustment data based on the personality description features.
[0141] For example, the server encodes the personality description based on the attention mechanism to obtain the personality description features of the personality description. The server performs multiple rounds of iterative decoding on the personality description features to obtain the second action adjustment data.
[0142] Among them, the above encoding and decoding are implemented through a second adjustment data determination model. The second adjustment data determination model can convert the personality description into the second action adjustment data. The second adjustment data determination model is trained based on multiple sample personality descriptions and the corresponding labeled action adjustment data for each sample personality description. The second adjustment data determination model is an encoding and decoding model based on the attention mechanism. The embodiments of the present application do not limit the structure and training method of the second adjustment data determination model.
[0143] B. The server determines the third action adjustment data and the second image adjustment data of the initial replicated digital human based on the public evaluation.
[0144] In a possible implementation, the server encodes the public evaluation based on the attention mechanism to obtain the public evaluation features of the public evaluation. The server performs multiple rounds of iterative decoding on the public evaluation features to obtain the second image adjustment data and the third action adjustment data.
[0145] Among them, the above encoding and decoding are implemented through a third adjustment data determination model. The third adjustment data determination model can convert the public evaluation into the image adjustment data and the action adjustment data. The third adjustment data determination model is trained based on multiple sample public evaluations and the corresponding labeled image adjustment data and labeled action adjustment data for each sample public evaluation. The third adjustment data determination model is an encoding and decoding model based on the attention mechanism. The embodiments of the present application do not limit the structure and training method of the third adjustment data determination model.
[0146] C. The server determines the second adjustment data of the initial replicated digital human based on the second action adjustment data, the third action adjustment data, and the second image adjustment data.
[0147] In a possible implementation, the server fuses the second action adjustment data and the third action adjustment data to obtain the intermediate action adjustment data. The server splices the intermediate action adjustment data and the second image adjustment data to obtain the second adjustment data of the initial replicated digital human.
[0148] Part Four: The server determines the updated data of the replicated digital human based on the subjective feeling data, the first adjustment data, and the second adjustment data.
[0149] In a possible implementation manner, the server determines an adjustment coefficient based on the subjective feeling data, and the adjustment data is used to control the adjustment amplitude. The server fuses the first adjustment data and the second adjustment data to obtain target adjustment data. The server fuses the adjustment coefficient with the target adjustment data to obtain the updated data of the replicated digital human.
[0150] For example, the server extracts features from the subjective feeling data to obtain the subjective feeling data features of the subjective feeling data. The server performs full connection and normalization on the subjective feeling data features to obtain the adjustment coefficient. The server performs weighted fusion on the first adjustment data and the second adjustment data to obtain target adjustment data. The server multiplies the adjustment coefficient by the target adjustment data to obtain the updated data of the replicated digital human.
[0151] Among them, the weights for weighted fusion are set by technicians according to the actual situation, and the embodiments of this application do not limit this.
[0152] Part Five: The server updates the initial replicated digital human with the updated data of the replicated digital human to obtain the target replicated digital human.
[0153] In a possible implementation manner, the server updates the initial replicated digital human display data of the initial replicated digital human with the updated data of the replicated digital human to obtain the target replicated digital human display data of the target replicated digital human. The server generates a virtual image of the target replicated digital human based on the target replicated digital human display data.
[0154] Among them, the initial replicated digital human display data includes multiple skeleton points, rendering parameters, and action modes of the initial replicated digital human. Updating the initial replicated digital human display data includes at least one of adjusting the positions of multiple skeleton points, adjusting the rendering parameters, and adjusting the action modes.
[0155] The methods for adjusting the positions of multiple skeleton points, adjusting the rendering parameters, and adjusting the action modes are described below respectively.
[0156] In some embodiments, the server obtains skeleton point update data from the updated data of the replicated digital human. The server uses the skeleton point update data to adjust the positions of the multiple skeleton points to obtain the updated multiple skeleton points.
[0157] In some embodiments, the server obtains rendering parameter update data from the replicated digital human update data. The server uses the rendering parameter update data to update the rendering data, obtaining updated rendering parameters.
[0158] In some embodiments, the server obtains action mode adjustment data from the replicated digital human update data. The server uses the action mode adjustment data to adjust the action mode, obtaining an updated action mode.
[0159] It should be noted that the above adjustment data are all used to indicate the adjustment method, and the form of the adjustment data is set by those skilled in the art according to the actual situation. The embodiments of the present application do not limit this.
[0160] 305. The server sends the target digital human display data of the target replicated digital human to the sending device, so that the sending device displays the target replicated digital human based on the target digital human display data.
[0161] Among them, after the target replicated digital human is displayed through the sending device, the sending device displays the target replicated digital human based on the target digital human display data. The target digital human display data includes multiple skeleton points, rendering parameters, and action modes of the target replicated digital human. After the sending device displays the target replicated digital human, the target user can interact with the target replicated digital human. For example, the target user can control the target replicated digital human to perform specific actions and can also have a conversation with the target replicated digital human. Action execution can be achieved by using the method of controlling the digital human by controlling the skeleton points in related technologies, and the conversation is achieved by the conversation model in related technologies. The embodiments of the present application do not limit this.
[0162] All of the above optional technical solutions can be combined arbitrarily to form optional embodiments of the present application, which will not be elaborated here one by one.
[0163] Through the technical solution provided by the embodiments of the present application, in response to the reproduction request of the target user for the target historical figure, the first historical figure data of the target historical figure, the first user data of the target user, and the device performance data of the sending device of the reproduction request are obtained. Based on the first historical figure data, the first user data, and the device performance data, an initial reproduced digital human is synthesized and displayed, thereby realizing the preliminary generation and display of the reproduced digital human. In response to the adjustment request of the target user for the initial reproduced digital human, the second user data of the target user, the device environment data of the sending device, and the second historical figure data of the target historical figure are obtained. Based on the second user data, the device environment data, and the second historical figure data, the initial reproduced digital human is updated to obtain the target reproduced digital human, realizing convenient update of the reproduced digital human. Subsequently, the target reproduced digital human can be used to interact with the target user, improving the satisfaction of the target user with the target virtual digital human, thereby improving the user experience.
[0164] Figure 4 is a structural schematic diagram of a real-time interaction synthesis processing device based on a reproduced digital human provided by the embodiments of the present application. Refer to Figure 4 The device includes: an acquisition module 401, a synthesis and display module 402, and an update module 403.
[0165] The acquisition module 401 is configured to, in response to the reproduction request of the target user for the target historical figure, obtain the first historical figure data of the target historical figure, the first user data of the target user, and the device performance data of the sending device of the reproduction request. The first historical figure data includes the life information and image records of the target historical figure, and the first user data includes the user image data and user action data of the target user.
[0166] The synthesis and display module 402 is configured to synthesize and display an initial reproduced digital human based on the first historical figure data of the target historical figure, the first user data, and the device performance data.
[0167] The acquisition module 401 is further configured to, in response to the adjustment request of the target user for the initial reproduced digital human, obtain the second user data of the target user, the device environment data of the sending device, and the second historical figure data of the target historical figure. The second user data includes the physiological data and adjustment suggestion data of the target user, and the second historical figure data includes the personality description and public evaluation of the target historical figure.
[0168] The update module 403 is configured to update the initial reproduced digital human based on the second user data, the device environment data, and the second historical figure data to obtain the target reproduced digital human.
[0169] In a possible implementation, the synthesis display module 402 is configured to determine first replicated digital human image data based on the image record of the target historical figure and the user image data of the target user. Determine first replicated digital human motion data based on the life information and the user motion data. Determine second replicated digital human image data based on the device performance data, the life information, and the first replicated digital human image data. Generate the initial replicated digital human based on the first replicated digital human motion data and the second replicated digital human image data. Display the initial replicated digital human.
[0170] In a possible implementation, the image record includes an image and an image description text, the user image data includes a first user facial image, a second user facial image, and an expected image description text of the target historical figure. The second user facial image is obtained by the target user editing the first user facial image. The synthesis display module 402 is configured to determine first reference image data of the target historical figure based on the image, the image description text, and the expected image description text. Determine second reference image data of the target user based on the first user facial image and the second user facial image. Determine the first replicated digital human image data based on the first reference image data and the second reference image data.
[0171] In a possible implementation, the user motion data includes current user motion data and historical user motion data. The synthesis display module 402 is configured to perform data identification on the life information to obtain motion description data of the target historical figure in the life information. Determine reference motion data based on the motion description data and the current user motion data. Determine the first replicated digital human motion data based on the historical user motion data and the reference motion data.
[0172] In a possible implementation, the synthesis display module 402 is configured to extract features from the life information to obtain life information features of the life information. Determine third replicated digital human image data based on the life information features and the first replicated digital human image data. Determine the second replicated digital human image data based on the device performance data and the third replicated digital human image data.
[0173] In a possible implementation, the update module 403 is configured to determine subjective feeling data of the target user for the initial replicated digital human based on the physiological data and the adjustment suggestion data. Determine first adjustment data for the initial replicated digital human based on the adjustment suggestion data and the device environment data, where the device environment data is used to describe the environment where the sending device is located. Determine second adjustment data for the initial replicated digital human based on the personality description and the public evaluation. Determine replicated digital human update data based on the subjective feeling data, the first adjustment data, and the second adjustment data. Update the initial replicated digital human with the replicated digital human update data to obtain the target replicated digital human.
[0174] In a possible implementation, the physiological data includes first physiological data and second physiological data. The first physiological data is the physiological data collected before presenting the initial replicated digital human, and the second physiological data is the physiological data collected after presenting the initial replicated digital human. The update module 403 is configured to determine physiological state change data of the target user based on the first physiological data and the second physiological data. Determine the satisfaction degree of the target user with the initial replicated digital human based on the adjustment suggestion data. Determine the subjective feeling data based on the physiological state change data and the satisfaction degree.
[0175] In a possible implementation, the device environment data includes device location information and device environment light information. The update module 403 is configured to determine first image adjustment data and first action adjustment data of the initial replicated digital human based on the adjustment suggestion data. Determine reference device adjustment data based on the device location information and the device environment light information, where the reference device adjustment data matches the device location information and the device environment light information. Determine the first adjustment data based on the first image adjustment data, the first action adjustment data, and the reference device adjustment data.
[0176] In a possible implementation, the update module 403 is configured to determine second action adjustment data of the initial replicated digital human based on the personality description. Determine third action adjustment data and second image adjustment data of the initial replicated digital human based on the public evaluation. Determine the second adjustment data of the initial replicated digital human based on the second action adjustment data, the third action adjustment data, and the second image adjustment data.
[0177] It should be noted that: when the real-time interactive synthesis processing device based on the replicated digital human provided in the above embodiment performs replicated digital human processing, only the division of the above-mentioned functional modules is used for illustration. In practical applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the computer device is divided into different functional modules to complete all or part of the functions described above. In addition, the real-time interactive synthesis processing device based on the replicated digital human provided in the above embodiment and the embodiment of the real-time interactive synthesis processing method based on the replicated digital human belong to the same concept. For the specific implementation process, please refer to the method embodiment and will not be elaborated here.
[0178] Through the technical solution provided by the embodiment of the present application, in response to the replication request of the target user for the target historical figure, the first historical figure data of the target historical figure, the first user data of the target user, and the device performance data of the sending device of the replication request are obtained. Based on the first historical figure data, the first user data, and the device performance data, an initial replicated digital human is synthesized and displayed, thus realizing the preliminary generation and display of the replicated digital human. In response to the adjustment request of the target user for the initial replicated digital human, the second user data of the target user, the device environment data of the sending device, and the second historical figure data of the target historical figure are obtained. Based on the second user data, the device environment data, and the second historical figure data, the initial replicated digital human is updated to obtain the target replicated digital human, realizing the convenient update of the replicated digital human. Subsequently, the target replicated digital human can be used to interact with the target user, improving the satisfaction of the target user with the target virtual digital human, thereby improving the user experience.
[0179] Figure 5 It is a schematic structural diagram of a server provided by an embodiment of the present application. The server 500 may vary greatly due to different configurations or performances, and may include one or more processors (Central Processing Units, CPU) 501 and one or more memories 502. Among them, at least one computer program is stored in the one or more memories 502, and the at least one computer program is loaded and executed by the one or more processors 501 to implement the methods provided by the above various method embodiments. Of course, the server 500 may also have components such as wired or wireless network interfaces, keyboards, and input / output interfaces for input / output. The server 500 may also include other components for implementing device functions, which will not be elaborated here.
[0180] In an exemplary embodiment, a computer-readable storage medium is further provided, such as a memory including a computer program, and the computer program can be executed by a processor to complete the real-time interactive synthesis processing method based on the replicated digital human in the above embodiment. For example, the computer-readable storage medium can be a Read-Only Memory (ROM), a Random Access Memory (RAM), a Compact Disc Read-Only Memory (CD-ROM), magnetic tape, floppy disk, and optical data storage device, etc.
[0181] In an exemplary embodiment, a computer program product or a computer program is further provided. The computer program product or the computer program includes program code, and the program code is stored in a computer-readable storage medium. The processor of the computer device reads the program code from the computer-readable storage medium, and the processor executes the program code, so that the computer device executes the above real-time interactive synthesis processing method based on the replicated digital human.
[0182] In some embodiments, the computer program involved in the embodiments of the present application can be deployed to be executed on one computer device, or on multiple computer devices located at one place. Or, it can be executed on multiple computer devices distributed at multiple places and interconnected through a communication network. The multiple computer devices distributed at multiple places and interconnected through a communication network can form a blockchain system.
[0183] Those of ordinary skill in the art can understand that all or part of the steps to implement the above embodiments can be completed by hardware, or can be completed by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium, and the above-mentioned storage medium can be a read-only memory, a magnetic disk, or an optical disc, etc.
[0184] The above are only optional embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A real-time interactive synthesis processing method based on replicated digital humans, characterized in that, The method includes: In response to a reproduction request of a target user for a target historical figure, obtaining first historical figure data of the target historical figure, first user data of the target user, and device performance data of the sending device of the reproduction request, where the first historical figure data includes biographical information and image records of the target historical figure, and the first user data includes user image data and user action data of the target user; Based on the first historical figure data of the target historical figure, the first user data, and the device performance data, synthesizing and displaying an initial reproduced digital human; In response to an adjustment request of the target user for the initial reproduced digital human, obtaining second user data of the target user, device environment data of the sending device, and second historical figure data of the target historical figure, where the second user data includes physiological data and adjustment suggestion data of the target user, and the second historical figure data includes character descriptions and public evaluations of the target historical figure; Based on the second user data, the device environment data, and the second historical figure data, updating the initial reproduced digital human to obtain a target reproduced digital human.
2. The method according to claim 1, characterized in that The synthesizing and displaying an initial reproduced digital human based on the first historical figure data of the target historical figure, the first user data, and the device performance data includes: Based on the image record of the target historical figure and the user image data of the target user, determining first reproduced digital human image data; Based on the biographical information and the user action data, determining first reproduced digital human action data; Based on the device performance data, the biographical information, and the first reproduced digital human image data, determining second reproduced digital human image data; Based on the first reproduced digital human action data and the second reproduced digital human image data, generating the initial reproduced digital human; Displaying the initial reproduced digital human.
3. The method according to claim 2, wherein The image record includes an image and an image description text, the user image data includes a first user facial image, a second user facial image, and an expected image description text of the target historical figure, the second user facial image is obtained by the target user editing the first user facial image, and the determining first reproduced digital human image data based on the image record of the target historical figure and the user image data of the target user includes: Based on the image, the image description text, and the expected image description text, determining first reference image data of the target historical figure; Based on the first user facial image and the second user facial image, determining second reference image data of the target user; Based on the first reference image data and the second reference image data, determining the first reproduced digital human image data.
4. The method according to claim 2, wherein The user action data includes current user action data and historical user action data, and the determining first reproduced digital human action data based on the biographical information and the user action data includes: Identify the data of the life information to obtain the action description data of the target historical figure in the life information; Determine the reference action data based on the action description data and the current user action data; Determine the first replicated digital human action data based on the historical user action data and the reference action data.
5. The method according to claim 2, wherein The determining the second replicated digital human image data based on the device performance data, the life information, and the first replicated digital human image data includes: Extract features from the life information to obtain the life information features of the life information; Determine the third replicated digital human image data based on the life information features and the first replicated digital human image data; Determine the second replicated digital human image data based on the device performance data and the third replicated digital human image data.
6. The method according to claim 1, wherein The updating the initial replicated digital human based on the second user data, the device environment data, and the second historical figure data to obtain the target replicated digital human includes: Determine the subjective feeling data of the target user for the initial replicated digital human based on the physiological data and the adjustment suggestion data; Determine the first adjustment data for the initial replicated digital human based on the adjustment suggestion data and the device environment data, where the device environment data is used to describe the environment where the sending device is located; Determine the second adjustment data for the initial replicated digital human based on the personality description and the public evaluation; Determine the replicated digital human update data based on the subjective feeling data, the first adjustment data, and the second adjustment data; Update the initial replicated digital human with the replicated digital human update data to obtain the target replicated digital human.
7. The method according to claim 6, wherein The physiological data includes first physiological data and second physiological data. The first physiological data is the physiological data collected before displaying the initial replicated digital human, and the second physiological data is the physiological data collected after displaying the initial replicated digital human. The determining the subjective feeling data of the target user for the initial replicated digital human based on the physiological data and the adjustment suggestion data includes: Determine the physiological state change data of the target user based on the first physiological data and the second physiological data; Determine the satisfaction degree of the target user for the initial replicated digital human based on the adjustment suggestion data; Determine the subjective feeling data based on the physiological state change data and the satisfaction degree.
8. The method according to claim 6, characterized in that, The device environment data includes device location information and device environment light information. The determining the first adjustment data for the initial replicated digital human based on the adjustment suggestion data and the device environment data includes: Determine the first image adjustment data and the first action adjustment data of the initial replicated digital human based on the adjustment suggestion data; Determine the reference device adjustment data based on the device location information and the device environment light information, and the reference device adjustment data matches the device location information and the device environment light information; Determine the first adjustment data based on the first image adjustment data, the first motion adjustment data, and the reference device adjustment data.
9. The method according to claim 6, wherein Determining the second adjustment data for the initial replicated digital human based on the personality description and the public evaluation includes: Determine the second motion adjustment data of the initial replicated digital human based on the personality description; Determine the third motion adjustment data and the second image adjustment data of the initial replicated digital human based on the public evaluation; Determine the second adjustment data of the initial replicated digital human based on the second motion adjustment data, the third motion adjustment data, and the second image adjustment data.
10. A real-time interactive synthesis processing system based on a replicated digital human, characterized in that, The system includes: An acquisition module, configured to, in response to a replication request of a target user for a target historical figure, acquire the first historical figure data of the target historical figure, the first user data of the target user, and the device performance data of the device where the replication request is sent. The first historical figure data includes the life information and image records of the target historical figure, and the first user data includes the user image data and user motion data of the target user; A synthesis and display module, configured to synthesize and display an initial replicated digital human based on the first historical figure data of the target historical figure, the first user data, and the device performance data; The acquisition module is further configured to, in response to an adjustment request of the target user for the initial replicated digital human, acquire the second user data of the target user, the device environment data of the sending device, and the second historical figure data of the target historical figure. The second user data includes the physiological data and adjustment suggestion data of the target user, and the second historical figure data includes the personality description and public evaluation of the target historical figure; An update module, configured to update the initial replicated digital human based on the second user data, the device environment data, and the second historical figure data to obtain a target replicated digital human.
Citation Information
Patent Citations
Virtual digital human construction method and system
CN117828320A
Human-computer interaction method and device, electronic equipment and computer storage medium
CN118567602A
KR20250014307A
Cited By
Digital human emotion processing method and system based on AI model
CN120895060A