A real-time interactive synthesis processing method and system based on replicating digital humans

By obtaining target historical figures and user data, combining equipment performance data to synthesize and update digital people, the problem that digital people cannot be personalized in the existing technology is solved, and human-computer interaction efficiency and user satisfaction are improved.

CN120336572BActive Publication Date: 2025-08-19HANGZHOU BEIMING AURORA TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510779729.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-08-19
Estimated Expiration
2045-06-12

AI Technical Summary

Technical Problem

Existing digital human creations cannot meet users' personalized needs, and the interaction lacks flexibility, making it difficult to adjust according to users' real-time needs, resulting in low human-computer interaction efficiency.

Method used

A real-time interactive synthesis processing method based on replica digital people is provided. By obtaining the data of target historical figures and users, synthesizing the initial digital people with equipment performance data, and updating it in response to user adjustment requests, to generate a target replica digital people that meets user needs.

Benefits of technology

It realizes personalization and flexibility of digital human generation, improves human-computer interaction efficiency, and improves users' satisfaction with virtual digital people.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336572B_ABST
    Figure CN120336572B_ABST
Patent Text Reader

Abstract

The present application discloses a real-time interactive synthesis processing method and system based on a replica digital human, belonging to the field of computer technology. In response to a target user's request for a replica of a target historical figure, the first historical figure data of the target historical figure, the first user data of the target user, and the device performance data are obtained. Based on the first historical figure data, the first user data, and the device performance data, an initial replica digital human is synthesized and displayed. In response to the target user's request to adjust the initial replica digital human, the second user data of the target user, the device environment data of the sending device, and the second historical figure data of the target historical figure are obtained. Based on the second user data, the device environment data, and the second historical figure data, the initial replica digital human is updated to obtain a target replica digital human, thereby achieving convenient replica digital human updates and improving the target user's satisfaction with the target virtual digital human.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a real-time interactive synthesis processing method and system based on replicating digital humans. Background Art

[0002] With the development of science and technology, digital human technology has great potential for application in fields such as culture and entertainment. In terms of cultural dissemination, digital humans recreating historical figures can promote cultural heritage. However, existing digital humans are mostly created using generic models that cannot meet the personalized needs of users.

[0003] However, current digital human interactions lack flexibility. Once created, they are difficult to adjust according to real-time user needs, requiring users to regenerate the digital human, resulting in low efficiency in human-computer interaction. Summary of the Invention

[0004] The present application provides a real-time interactive synthesis processing method and system based on a replica digital human, which can generate a digital human that better meets user needs and improve the efficiency of human-computer interaction. The technical solution is as follows:

[0005] In one aspect, a real-time interactive synthesis processing method based on a replica digital human is provided, the method comprising:

[0006] In response to a target user's request for a reproduction of a target historical figure, obtaining first historical figure data of the target historical figure, first user data of the target user, and device performance data of a device sending the reproduction request, wherein the first historical figure data includes biographical information and an image record of the target historical figure, and the first user data includes user image data and user action data of the target user;

[0007] synthesizing and displaying an initial replica digital human based on the first historical figure data of the target historical figure, the first user data, and the device performance data;

[0008] In response to the target user's request for adjustment of the initially replicated digital human, obtaining second user data of the target user, device environment data of the sending device, and second historical figure data of the target historical figure, wherein the second user data includes physiological data of the target user and adjustment suggestion data, and the second historical figure data includes a personality description and public evaluation of the target historical figure;

[0009] Based on the second user data, the device environment data and the second historical figure data, the initial replica digital human is updated to obtain a target replica digital human.

[0010] In one aspect, a real-time interactive synthesis processing device based on a replica digital human is provided, the device comprising:

[0011] an acquisition module, configured to, in response to a target user's request for replicating a target historical figure, acquire first historical figure data of the target historical figure, first user data of the target user, and device performance data of a device sending the replica request, wherein the first historical figure data includes biographical information and an image record of the target historical figure, and the first user data includes user image data and user action data of the target user;

[0012] a synthesis and display module, configured to synthesize and display an initial replica digital human based on the first historical figure data of the target historical figure, the first user data, and the device performance data;

[0013] The acquisition module is further configured to, in response to the target user's request for adjustment of the initially replicated digital human, acquire second user data of the target user, device environment data of the sending device, and second historical figure data of the target historical figure, wherein the second user data includes physiological data of the target user and adjustment suggestion data, and the second historical figure data includes a personality description and public evaluation of the target historical figure;

[0014] An updating module is used to update the initial replica digital human based on the second user data, the device environment data and the second historical figure data to obtain a target replica digital human.

[0015] On the one hand, a computer device is provided, which includes one or more processors and one or more memories, wherein at least one computer program is stored in the one or more memories, and the computer program is loaded and executed by the one or more processors to implement the real-time interactive synthesis processing method based on replicating digital humans.

[0016] On the one hand, a computer-readable storage medium is provided, in which at least one computer program is stored. The computer program is loaded and executed by a processor to implement the real-time interactive synthesis processing method based on replicating digital humans.

[0017] On the one hand, a computer program product or computer program is provided, which includes a program code, which is stored in a computer-readable storage medium. A processor of a computer device reads the program code from the computer-readable storage medium, and the processor executes the program code, so that the computer device performs the above-mentioned real-time interactive synthesis processing method based on replicating digital humans.

[0018] Through the technical solution provided by the embodiment of the present application, in response to a target user's request to replicate a target historical figure, the first historical figure data of the target historical figure, the first user data of the target user, and the device performance data of the device sending the replication request are obtained. Based on the first historical figure data, the first user data, and the device performance data, an initial replicated digital human is synthesized and displayed, thereby achieving the preliminary generation and display of the replicated digital human. In response to the target user's request to adjust the initial replicated digital human, the second user data of the target user, the device environment data of the sending device, and the second historical figure data of the target historical figure are obtained. Based on the second user data, the device environment data, and the second historical figure data, the initial replicated digital human is updated to obtain a target replicated digital human, thereby achieving convenient replication of the replicated digital human. The target replicated digital human can subsequently be used to interact with the target user, thereby increasing the target user's satisfaction with the target virtual digital human, thereby improving the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present application. Those skilled in the art can also derive other drawings based on these drawings without inventive effort.

[0020] Figure 1 Schematic diagram of an implementation environment of a real-time interactive synthesis processing method based on a replica digital human provided in an embodiment of the present application;

[0021] Figure 2 This is a flow chart of a real-time interactive synthesis processing method based on replicating digital humans provided by an embodiment of the present application;

[0022] Figure 3 This is a flowchart of another real-time interactive synthesis processing method based on replicating digital humans provided by an embodiment of the present application;

[0023] Figure 4 This is a structural diagram of a real-time interactive synthesis processing device based on a replica digital human provided by an embodiment of the present application;

[0024] Figure 5 This is a structural diagram of a server provided in an embodiment of the present application. DETAILED DESCRIPTION

[0025] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.

[0026] In this application, the terms "first", "second", etc. are used to distinguish identical or similar items with substantially the same effects and functions. It should be understood that there is no logical or temporal dependency between "first", "second", and "nth", nor is there any limitation on the quantity and execution order.

[0027] Digital Human: A digital human is a virtual human created using computer technology. It integrates information science and life science, aiming to simulate the human body at various levels. As an emerging technology, digital humans possess multiple human characteristics, including physical appearance, performance abilities, and interactive capabilities.

[0028] Replicating digital humans: Replicating digital humans refers to simulating the appearance and movements of real humans through computer technology to create a virtual image with interactive capabilities.

[0029] Augmented Reality: Augmented reality is a technology that integrates digital information into the real world. Through a device with a camera, users can interact with their physical space and computer-generated content simultaneously.

[0030] Virtual Reality: Virtual reality is a technology that creates realistic three-dimensional simulated environments, allowing users to interact with the virtual environment through VR devices such as helmets and controllers.

[0031] Metaverse: The metaverse is a virtual shared space, usually composed of multiple virtual reality worlds or games, in which users can engage in social interactions, asset transactions, virtual tourism and other activities.

[0032] In the embodiments of the present application, digital humans can be used in augmented reality, virtual reality, and the metaverse.

[0033] Artificial Intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to achieve the best results.

[0034] Machine learning (ML) is a multidisciplinary field that encompasses probability theory, statistics, approximation theory, convex analysis, and algorithmic complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize existing knowledge sub-models to continuously improve their performance. Machine learning is at the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications span all areas of AI. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning through demonstration.

[0035] Semantic features: Features used to represent the semantic meaning of a text. Different texts can correspond to the same semantic features. For example, the text "How's the weather today?" and the text "How is the weather today?" can both correspond to the same semantic feature. Computers can map characters in a text into character vectors and, based on the relationships between characters, combine and operate on these character vectors to obtain the semantic features of the text. For example, computers can use Bidirectional Encoder Representations from Transformers (BERT).

[0036] Normalization: Mapping sequences of numbers with different value ranges to the interval (0, 1) facilitates data processing. In some cases, the normalized values can be directly implemented as probabilities.

[0037] Embedded Coding: Mathematically, embedded coding represents a correspondence relationship, mapping data in X space to Y space using a function F, where F is an injective function. The mapping result preserves structure. An injective function indicates that the data after mapping uniquely corresponds to the data before mapping. Structural preservation means that the size relationship of the data before mapping is the same as the size relationship of the data after mapping. For example, if there are data X1 and X2 before mapping, the data Y1 corresponding to X1 and Y2 corresponding to X2 will be obtained after mapping. If the data X1 before mapping is greater than X2, then the data Y1 after mapping is greater than Y2. For words, this means mapping the words to another space to facilitate subsequent machine learning and processing.

[0038] Attention weight: This indicates the importance of a piece of data during training or prediction. Importance indicates the impact of the input data on the output data. Highly important data has a higher attention weight, while lowly important data has a lower attention weight. Data importance varies in different scenarios, and the process of training a model's attention weights is also the process of determining data importance.

[0039] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, storage, and display, etc.), and signals involved in this application are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the first user data and the second user data involved in this application were both obtained with full authorization.

[0040] Figure 1This is a schematic diagram of the implementation environment of a real-time interactive synthesis processing method based on a replica digital human provided by an embodiment of the present application, see Figure 1 , the implementation environment may include a terminal 110 and a server 140.

[0041] The terminal 110 is connected to the server 140 via a wireless network or a wired network. Optionally, the terminal 110 is a smart phone, a tablet computer, a laptop computer, a desktop computer, etc., but is not limited thereto. The terminal 110 is installed and running an application that supports real-time interactive synthesis processing based on the replica digital human.

[0042] Server 140 is an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, a Content Delivery Network (CDN), and big data and artificial intelligence platforms. Server 140 provides backend services for applications running on terminal 110.

[0043] Those skilled in the art will appreciate that the number of terminals may be greater or less. For example, there may be only one terminal, or there may be dozens, hundreds, or even more terminals, in which case the implementation environment may also include other terminals. The embodiments of this application do not limit the number or device types of terminals.

[0044] After introducing the implementation environment of the embodiments of the present application, the following describes the application scenarios of the embodiments of the present application. In scenarios where the technical solutions provided by the embodiments of the present application can create replica digital humans, users can use the technical solutions provided by the embodiments of the present application to create the desired digital humans and subsequently interact with the created digital humans.

[0045] After introducing the implementation environment and application scenarios of the embodiments of the present application, the technical solutions provided by the embodiments of the present application are described below. Figure 2 This is a flowchart of a real-time interactive synthesis processing method based on replicating digital humans provided by an embodiment of the present application. Figure 2 Taking the execution subject as a server as an example, the method includes the following steps.

[0046] 201. In response to a target user's request to replicate a target historical figure, the server obtains the first historical figure data of the target historical figure, the first user data of the target user, and the device performance data of the device sending the replication request, wherein the first historical figure data includes the biographical information and image record of the target historical figure, and the first user data includes the user image data and user action data of the target user.

[0047] The target user is the user using the virtual digital human service. The replica digital human is a virtual human image generated using computer technology, with different replica digital humans having different appearances and movements. The target historical figure is the historical figure selected by the target user and the subject of the digital human replica. Accordingly, the replica request is used to request a replica of the target historical figure to generate a replica digital human corresponding to the target historical figure. The sending device is the computer device used by the target user. After the replica virtual object is generated, it is displayed through the sending device. The device performance data of the sending device indicates the rendering capabilities of the sending device. Since displaying the replica virtual figure requires real-time rendering on the sending device, a higher rendering capability allows for rendering of a virtual figure with greater detail and movement complexity. Biographical information and image records are obtained from a historical figure database, which stores the biographical information and image records of multiple candidate historical figures. The target historical figure is one of these multiple candidate historical figures. User image data includes image data related to the target user, and user action data includes action data related to the target user.

[0048] 202. The server synthesizes and displays an initial replica digital human based on the first historical figure data of the target historical figure, the first user data, and the device performance data.

[0049] The initial replica digital human is generated using the first historical figure's data, the first user's data, and the device's performance data. This initial replica digital human closely matches the target historical figure, the target user, and the sending device. Displaying the initial replica digital human refers to displaying the initial replica digital human and controlling it to speak and perform actions. Since the server cannot directly display the initial replica digital human, the aforementioned display actually involves the server displaying the initial replica digital human through the sending device, and the target user can view the initial replica digital human through the initial replica digital human.

[0050] 203. In response to the target user's request to adjust the initial replica digital person, the server obtains the second user data of the target user, the device environment data of the sending device, and the second historical figure data of the target historical figure, the second user data including the physiological data of the target user and the adjustment suggestion data, and the second historical figure data including the personality description and public evaluation of the target historical figure.

[0051] Among them, the adjustment request is used to request adjustments to the initial replica digital human. The appearance of the adjustment request means that the target user is not satisfied with the initial replica digital human. At this time, other data will be obtained to adjust the initial replica digital human. The personality description and public evaluation of the target historical figure are also stored in the historical figure database and can be directly obtained from the historical figure database. The physiological data of the target user is used to represent the physiological state of the target user. The adjustment suggestion data is carried by the adjustment request, that is, when the adjustment request is sent, the adjustment suggestion data will be carried. The adjustment suggestion data is used to indicate the target user's suggestions for adjusting the initial replica digital human. The device environment data is used to represent the environmental conditions of the environment in which the sending device is located.

[0052] 204. The server updates the initial replica digital human based on the second user data, the device environment data and the second historical figure data to obtain a target replica digital human.

[0053] Among them, the target replica digital human is the updated replica digital human, and subsequent target users can interact with the target replica digital human.

[0054] Through the technical solution provided by the embodiment of the present application, in response to a target user's request to replicate a target historical figure, the first historical figure data of the target historical figure, the first user data of the target user, and the device performance data of the device sending the replication request are obtained. Based on the first historical figure data, the first user data, and the device performance data, an initial replicated digital human is synthesized and displayed, thereby achieving the preliminary generation and display of the replicated digital human. In response to the target user's request to adjust the initial replicated digital human, the second user data of the target user, the device environment data of the sending device, and the second historical figure data of the target historical figure are obtained. Based on the second user data, the device environment data, and the second historical figure data, the initial replicated digital human is updated to obtain a target replicated digital human, thereby achieving convenient replication of the replicated digital human. The target replicated digital human can subsequently be used to interact with the target user, thereby increasing the target user's satisfaction with the target virtual digital human, thereby improving the user experience.

[0055] The above steps 201-204 are a brief introduction to the real-time interactive synthesis processing method based on the replica digital human provided by the embodiment of the present application. The following will combine some examples to more clearly illustrate the real-time interactive synthesis processing method based on the replica digital human provided by the embodiment of the present application. Figure 3 Taking the execution subject as a server as an example, the method includes the following steps.

[0056] 301. In response to a target user's request to replicate a target historical figure, the server obtains the first historical figure data of the target historical figure, the first user data of the target user, and the device performance data of the device sending the replication request. The first historical figure data includes the biographical information and image record of the target historical figure, and the first user data includes the user image data and user action data of the target user.

[0057] The target user is the user using the virtual digital human service. The replica digital human is a virtual human image generated using computer technology, with different replica digital humans having different appearances and movements. The target historical figure is the historical figure selected by the target user and the subject of the digital human replica. Accordingly, the replica request is used to request a replica of the target historical figure to generate a replica digital human corresponding to the target historical figure. The sending device is the computer device used by the target user. After the replica virtual object is generated, it is displayed through the sending device. The device performance data of the sending device indicates the rendering capabilities of the sending device. Since displaying the replica virtual figure requires real-time rendering on the sending device, a higher rendering capability allows for rendering of a virtual figure with greater detail and movement complexity. Biographical information and image records are obtained from a historical figure database, which stores the biographical information and image records of multiple candidate historical figures. The target historical figure is one of these multiple candidate historical figures. User image data includes image data related to the target user, and user action data includes action data related to the target user.

[0058] In one possible implementation, in response to a target user's request to reproduce a target historical figure, the server queries a historical figure database based on the target historical figure's identifier to obtain the target historical figure's biographical information and image record. The server then parses the reproduction request to obtain the target user's user image data, user action data, and device performance data of the sending device.

[0059] The biographical data is in text form, and the image record includes an image and a text description of the image. The user image data includes the target user's facial image and a text description of the target historical figure's desired image. The user action data includes the current user action data and historical user action data.

[0060] 302. The server synthesizes and displays an initial replica digital human based on the first historical figure data of the target historical figure, the first user data, and the device performance data.

[0061] The initial replica digital human is generated using the first historical figure's data, the first user's data, and the device's performance data. This initial replica digital human closely matches the target historical figure, the target user, and the sending device. Displaying the initial replica digital human refers to displaying the initial replica digital human and controlling it to speak and perform actions. Since the server cannot directly display the initial replica digital human, the aforementioned display actually involves the server displaying the initial replica digital human through the sending device, and the target user can view the initial replica digital human through the initial replica digital human.

[0062] In one possible implementation, the server determines the image data of the first replica digital human based on the image record of the target historical figure and the user image data of the target user. The server also determines the motion data of the first replica digital human based on the biographical information and the user motion data. The server also determines the image data of the second replica digital human based on the device performance data, the biographical information, and the image data of the first replica digital human. The server generates the initial replica digital human based on the motion data of the first replica digital human and the image data of the second replica digital human. The server displays the initial replica digital human.

[0063] The first replica digital human image data is derived from the image record of the target historical figure and the user image data of the target user. This data is equivalent to being derived from the image record of the target historical figure and modified using the user image data, and closely matches the image record of the target historical figure and the user image data of the target user. The first replica digital human motion data is derived from the biographical information of the target historical figure and the user motion data of the target user. This data is equivalent to being derived from the biographical information and modified using the user motion data, and closely matches the biographical information and user motion data of the target historical figure. The above-mentioned process of obtaining the first replica digital human image data and the first replica digital human motion data not only considers the basic image record and biographical information of the target historical figure, but also incorporates the user image data and user motion data related to the target user. This ensures that the subsequently generated replica digital human not only restores the target historical figure but also incorporates relevant information about the target user. The second replica digital human image data is derived from the device performance data, biographical information, and the first replica digital human image data, and is equivalent to using the device performance data and biographical information to modify the first replica digital human image data. The initial replica digital human is generated by combining the motion data of the first replica digital human and the image data of the second replica digital human. This means that the image of the initial replica digital human is generated based on the image data of the second replica digital human, and the motions are generated based on the motion data of the first replica digital human. The server displays the initial replica digital human by sending the device.

[0064] In order to explain the above embodiment more clearly, the above embodiment will be described in several parts below.

[0065] In the first part, the server determines the first replica digital human image data based on the image record of the target historical figure and the user image data of the target user.

[0066] In one possible implementation, the image record includes an image image and image description text. The user image data includes a first user facial image, a second user facial image, and a desired image description text for the target historical figure. The second user facial image is obtained by the target user editing the first user facial image. The server determines first reference image data for the target historical figure based on the image image, the image description text, and the desired image description text. The server determines second reference image data for the target user based on the first user facial image and the second user facial image. The server determines the first replica digital human image data based on the first and second reference image data.

[0067] The first user facial image and the second user facial image are both provided by the target user. The first user facial image is a user facial image directly captured using an image acquisition device, and the second user facial image is a user facial image obtained by the target user through image editing of the first user facial image. The change between the first user facial image and the second user facial image can reflect the target user's preference for facial images, thereby guiding the generation of the replica digital human image. The expected image description text is provided by the target user and is used to reflect the target user's expectations for the image of the target historical figure. In the embodiment of the present application, the expected image description text is used to assist in the generation of the replica digital human. The first reference image data is generated by combining the image image, the image description text, and the expected image description text, and can reflect the basic image of the target historical figure and the target user's expectations for the image of the target historical figure. The second reference image data is determined based on the first user facial image and the second user facial image, and can reflect the target user's preferences for facial images.

[0068] In the above embodiment, the image image, image description text, and the desired image description text are used to determine the first reference image data of the target historical figure, and the facial image of the first user and the facial image of the second user are used to determine the second reference image data of the target user. The combination of the first reference image data and the second reference image data to determine the first replica digital human image data has a high accuracy.

[0069] In order to explain the above embodiment more clearly, the above embodiment will be described in several parts below.

[0070] A. The server determines first reference image data of the target historical figure based on the image, the image description text and the expected image description text.

[0071] In one possible implementation, the server determines basic image data of the target historical figure based on the image and the image description text. The server determines expected image data of the target historical figure based on the image and the expected image description text. The server determines first reference image data of the target historical figure based on the basic image data and the expected image data of the target historical figure.

[0072] For example, the server performs feature extraction on the image image and the image description text respectively to obtain the image image features of the image image and the image description text features of the image description text. The server fuses the image image features and the image description text features to obtain basic image features. The server decodes the basic image features to obtain the basic image data. The server performs feature extraction on the image image and the expected image description text respectively to obtain the image image features of the image image and the expected image description text features of the expected image description text. The server fuses the image image features and the expected image description text features to obtain expected image features. The server decodes the expected image features to obtain the expected image data. The server performs weighted fusion on the basic image data and the expected image data of the target historical figure to obtain the first reference image data of the target historical figure.

[0073] Among them, the weights of weighted fusion are set by technical personnel according to actual conditions. Generally speaking, the weight corresponding to the basic image data is higher than the weight corresponding to the expected image data, so that the first reference image data can be more biased towards the basic image data, that is, the original image of the target historical person is more retained.

[0074] For example, the server encodes the image image and the image description text separately based on the attention mechanism to obtain the image image features of the image image and the image description text features of the image description text. The server fuses the image image features and the image description text features to obtain basic image features. The server performs multiple rounds of iterative decoding on the basic image features based on the attention mechanism to obtain the basic image data. The server encodes the image image and the expected image description text separately based on the attention mechanism to obtain the image image features of the image image and the expected image description text features of the expected image description text. The server fuses the image image features and the expected image description text features to obtain expected image features. The server performs multiple rounds of iterative decoding on the expected image features based on the attention mechanism to obtain the expected image data. The server performs weighted fusion on the basic image data and the expected image data of the target historical figure to obtain the first reference image data of the target historical figure.

[0075] The above encoding and decoding processes are semantic encoding and semantic decoding processes, which are equivalent to data mining and fusion processes.

[0076] B. The server determines the second reference image data of the target user based on the first user's facial image and the second user's facial image.

[0077] In one possible implementation, the server determines, based on the first user's facial image and the second user's facial image, a text describing a facial image difference between the first user's facial image and the second user's facial image. The server determines second reference image data for the target user based on the text describing the facial image difference and the second user's facial image.

[0078] The facial image difference description text is used to represent the image difference between the first user's facial image and the second user's facial image, thereby reflecting the target user's preference for facial images.

[0079] For example, the server performs feature extraction on the facial image of the first user and the facial image of the second user, respectively, to obtain first facial features of the first user's facial image and second facial features of the second user's facial image. The server determines text describing the differences between the facial images based on the first and second facial features. The server then fuses the text describing the differences between the facial images with the facial image of the second user to obtain second reference image data for the target user.

[0080] For example, the server encodes the facial image of the first user and the facial image of the second user respectively based on the attention mechanism to obtain the first facial features of the first user facial image and the second facial features of the second user facial image. The server determines the facial difference features between the first facial features and the second facial features. The server performs multiple rounds of iterative decoding on the facial difference features based on the attention mechanism to obtain the facial image difference description text. The server performs feature extraction on the facial image difference description text respectively to obtain the difference description text features of the facial image difference description text. The server fuses the difference description text features and the second facial features of the second user facial image to obtain the second reference image features. The server performs multiple rounds of iterative decoding on the second reference image features based on the attention mechanism to obtain the second reference image data of the target user.

[0081] C. The server determines the first replica digital human image data based on the first reference image data and the second reference image data.

[0082] In one possible implementation, the server encodes the first reference image data and the second reference image data based on an attention mechanism to obtain first reference image data features of the first reference image data and second reference image data features of the second reference image data. The server fuses the first reference image data features with the second reference image data features to obtain first replica digital human image features. The server performs multiple rounds of iterative decoding on the first replica digital human image features based on the attention mechanism to obtain the first replica digital human image data.

[0083] The first replica digital human image data is in text form or image form, which is not limited in this embodiment of the present application.

[0084] It should be noted that the above feature extraction processes are all implemented through a multimodal feature extractor based on the attention mechanism, so that data of different modalities can be mapped to the same feature space, which is convenient for subsequent processing.

[0085] Part 2: The server determines the first replica digital human motion data based on the biographical information and the user motion data.

[0086] In one possible implementation, the user motion data includes current user motion data and historical user motion data. The server performs data recognition on the biographical data to obtain motion description data of the target historical figure in the biographical data. Based on the motion description data and the current user motion data, the server determines reference motion data. Based on the historical user motion data and the reference motion data, the server determines the first replica digital human motion data.

[0087] The current user action data is generated by capturing the target user's actions during the current replica digital human generation cycle. The sending device displays action prompt text. After the action prompt text is displayed, the target user's actions are captured, generating the current action data. The historical action data is generated by capturing the target user's actions in the past. This user action data is captured using a motion capture device. This action data is used to guide the execution of the generated replica virtual human's actions.

[0088] In order to explain the above embodiment more clearly, the above embodiment will be described in several parts below.

[0089] A. The server performs data recognition on the biographical data and obtains action description data of the target historical figure in the biographical data.

[0090] In a possible implementation, the server extracts features from the biographical data to obtain biographical features of the biographical data, and determines the action description data from the biographical data based on the biographical features.

[0091] For example, the server divides the biographical information into multiple paragraphs. The server encodes these paragraphs using an attention mechanism to obtain paragraph features for each paragraph. These features form the biographical information features for the biographical information. The server classifies each paragraph based on its features to obtain a paragraph type for each paragraph. The server combines the paragraphs with the action description type from the multiple paragraphs to obtain the action description information.

[0092] For example, the server divides the biographical information into multiple paragraphs according to punctuation marks. The server encodes the multiple paragraphs based on the attention mechanism to obtain the paragraph features of each paragraph. The server fully connects and normalizes the paragraph features of each paragraph to obtain a probability set for each paragraph. The probability set includes multiple probabilities, and one probability corresponds to a candidate paragraph type. For any paragraph among the multiple paragraphs, the server determines the candidate paragraph type corresponding to the highest probability in the probability set of the paragraph as the paragraph type of the paragraph. The server combines the paragraphs of the multiple paragraphs whose paragraph type is action description to obtain the action description information.

[0093] B. The server determines reference action data based on the action description information and the current user action data.

[0094] In one possible implementation, the server extracts features from the action description material and the current user action data to obtain action description material features of the action description material and current user action features of the current user action data. The server fuses the action description material features and the current user action features to obtain a reference action feature. The server decodes the reference action feature to obtain the reference action data.

[0095] The reference action data is in text form or latent feature form, which is not limited in the embodiment of the present application. The latent feature is similar to the definition of the latent feature h in the long short-term memory network (LSTM).

[0096] For example, the server encodes the action description data and the current user action data separately based on the attention mechanism to obtain action description data features of the action description data and current user action features of the current user action data. The server performs weighted fusion of the action description data features and the current user action features to obtain a reference action feature. The server performs multiple rounds of iterative decoding on the reference action feature based on the attention mechanism to obtain the reference action data.

[0097] The weights of weighted fusion are set by technicians according to actual conditions, and are not limited in this embodiment of the present application. The above encoding and decoding processes are semantic encoding and semantic decoding processes, which are equivalent to data mining and fusion processes.

[0098] C. The server determines the first replica digital human motion data based on the historical user motion data and the reference motion data.

[0099] In a possible implementation, the server splices the historical user motion data and the reference motion data to obtain the first replicated digital human motion data.

[0100] Among them, splicing is to preserve the data completely.

[0101] Part 3: The server determines the second replica digital human image data based on the device performance data, the biographical information and the first replica digital human image data.

[0102] In one possible implementation, the server extracts features from the biographical data to obtain biographical features of the biographical data. Based on the biographical features and the first replica digital human image data, the server determines third replica digital human image data. Based on the device performance data and the third replica digital human image data, the server determines second replica digital human image data.

[0103] For example, the server encodes the biographical information based on an attention mechanism to obtain biographical information features of the biographical information. The server encodes the first replica digital human image data based on an attention mechanism to obtain first replica digital human image data features of the first replica digital human image data. The server fuses the biographical information features with the first replica digital human image data features to obtain third replica digital human image data features. The server decodes the third replica digital human image data features to obtain third replica digital human image data. The server determines digital human rendering effect upper limit information based on the device performance data. The server determines second replica digital human image data based on the digital human rendering effect upper limit information and the third replica digital human image data.

[0104] The upper limit information of the digital human rendering effect indicates the maximum digital human rendering capability of the sending device. The upper limit information of the digital human rendering effect includes the upper limit of the rendering resolution, the upper limit of the number of rendering samples, the upper limit of reflection and refraction, etc.

[0105] The following describes how the server determines the upper limit of the digital human rendering effect based on the device performance data in the above example.

[0106] In some embodiments, the server performs multiple full connections on the device performance parameters to obtain device performance characteristics of the device performance parameters. The server performs full connections and normalization on the device performance characteristics to obtain the upper limit information of the digital human rendering effect.

[0107] The following describes how the server determines the second replica digital human image data based on the digital human rendering effect upper limit information and the third replica digital human image data in the above example.

[0108] In some embodiments, the server merges the digital human rendering effect upper limit information and the third replica digital human image data to obtain the second replica digital human image data.

[0109] Part 4: The server generates the initial replica digital human based on the first replica digital human motion data and the second replica digital human image data.

[0110] In one possible implementation, the server generates multiple skeletal points based on the second replica digital human image data and renders the surface formed by the multiple skeletal points to obtain a digital human virtual image of the initial replica digital human. Based on the first replica digital human motion data, the server generates a digital human motion pattern for the initial replica digital human. The digital human motion pattern is used to constrain the movement of the multiple skeletal points.

[0111] The digital human virtual image is a visualized three-dimensional model, which is a skeleton-skin model. The skeleton refers to the multiple skeleton points, and the skin refers to the surface formed by the multiple skeleton points.

[0112] For example, the server encodes the second replica digital human image data based on the attention mechanism to obtain the second replica digital human image data features. The server performs multiple rounds of iterative decoding on the second replica digital human image data features based on the attention mechanism to obtain the positions of multiple skeletal points and rendering parameters. The server creates an initial three-dimensional model based on the positions of the multiple skeletal points, and uses the rendering parameters to render the surface formed by the multiple skeletal points to obtain the digital human virtual image of the initial replica digital human. The server encodes the first replica digital human motion data and the positions of the multiple skeletal points based on the attention mechanism to obtain the first replica digital human motion data features of the first replica digital human motion data and the position features of the multiple skeletal points. The server decodes the first replica digital human motion data features and the position features of the multiple skeletal points based on the attention mechanism to obtain the digital human motion mode of the initial replica digital human.

[0113] Part 5: The server displays the initial replica digital human.

[0114] In a possible implementation, the server sends the initial replica digital human presentation data of the initial replica digital human to the sending device, so that the sending device presents the initial replica digital human based on the initial replica digital human presentation data.

[0115] The initial replica digital human display data includes the digital human virtual image and digital human motion pattern of the initial replica digital human.

[0116] 303. In response to the target user's request to adjust the initial replica digital person, the server obtains the second user data of the target user, the device environment data of the sending device, and the second historical figure data of the target historical figure, the second user data including the physiological data of the target user and the adjustment suggestion data, and the second historical figure data including the personality description and public evaluation of the target historical figure.

[0117] Among them, the adjustment request is used to request adjustments to the initial replica digital human. The appearance of the adjustment request means that the target user is not satisfied with the initial replica digital human. At this time, other data will be obtained to adjust the initial replica digital human. The personality description and public evaluation of the target historical figure are also stored in the historical figure database and can be directly obtained from the historical figure database. The physiological data of the target user is used to represent the physiological state of the target user. The adjustment suggestion data is carried by the adjustment request, that is, when the adjustment request is sent, the adjustment suggestion data will be carried. The adjustment suggestion data is used to indicate the target user's suggestions for adjusting the initial replica digital human. The device environment data is used to represent the environmental conditions of the environment in which the sending device is located.

[0118] In one possible implementation, in response to the target user's request to adjust the initial replica digital human, the server obtains the second user data and the device environment data from the adjustment request. The server then searches a historical figure database based on the target historical figure's identifier to obtain the second historical figure data for the target historical figure.

[0119] The historical figures database is a database maintained by the server.

[0120] 304. The server updates the initial replica digital human based on the second user data, the device environment data and the second historical figure data to obtain a target replica digital human.

[0121] Among them, the target replica digital human is the updated replica digital human, and subsequent target users can interact with the target replica digital human.

[0122] In one possible implementation, the server determines the target user's subjective perception data of the initial replica digital human based on the physiological data and the adjustment suggestion data. The server determines first adjustment data for the initial replica digital human based on the adjustment suggestion data and the device environment data, where the device environment data is used to describe the environment in which the sending device is located. The server determines second adjustment data for the initial replica digital human based on the personality description and the public evaluation. The server determines replica digital human update data based on the subjective perception data, the first adjustment data, and the second adjustment data. The server uses the replica digital human update data to update the initial replica digital human, obtaining the target replica digital human.

[0123] The integration of device environment data aims to ensure that the displayed replica digital human matches the environment of the sending device. The integration of adjustment suggestion data aims to adjust the replica digital human in accordance with user instructions. The integration of personality description data and public reviews aims to enrich the image of the replica digital human.

[0124] In order to explain the above embodiment more clearly, the above embodiment will be described in several parts below.

[0125] In the first part, the server determines the target user's subjective feeling data on the initial replica digital human based on the physiological data and the adjustment suggestion data.

[0126] In one possible implementation, the physiological data includes first physiological data and second physiological data. The first physiological data is collected before the initial replica digital human is displayed, and the second physiological data is collected after the initial replica digital human is displayed. The server determines the target user's physiological state change data based on the first and second physiological data. The server determines the target user's satisfaction with the initial replica digital human based on the adjustment suggestion data. The server determines the subjective perception data based on the physiological state change data and the satisfaction level.

[0127] In order to explain the above embodiment more clearly, the above embodiment will be described in several parts below.

[0128] A. The server determines physiological state change data of the target user based on the first physiological data and the second physiological data.

[0129] In a possible implementation, the server subtracts the second physiological data from the first physiological data to obtain physiological state change data.

[0130] B. The server determines the target user's satisfaction with the initial replica digital human based on the adjustment suggestion data.

[0131] In a possible implementation, the server extracts features from the adjustment suggestion data to obtain adjustment suggestion data features of the adjustment suggestion data. Based on the adjustment suggestion data features, the server determines the target user's satisfaction with the initial replica digital human.

[0132] For example, the server encodes the adjustment suggestion data using an attention mechanism to obtain features of the adjustment suggestion data. The server then performs full-connection and normalization on the features of the adjustment suggestion data to obtain a satisfaction classification value. The server then determines the candidate satisfaction level corresponding to the classification value interval to which the satisfaction classification value belongs as the target user's satisfaction level with the initial replica digital human.

[0133] There are multiple classification value intervals, each of which corresponds to a candidate satisfaction level. The division of the multiple classification value intervals and the correspondence between the classification value intervals and the candidate satisfaction levels are determined by technicians based on actual conditions and are not limited in this embodiment of the present application. For example, the multiple candidate satisfaction levels include very dissatisfied, dissatisfied, average, satisfied, and very satisfied.

[0134] C. The server determines the subjective feeling data based on the physiological state change data and the satisfaction level.

[0135] In one possible implementation, the server extracts features from the physiological state change data to obtain physiological state change features. Based on the physiological state change features, the server determines the target user's physiological preference information for the initial replica digital human. Based on the physiological preference information and the satisfaction level, the server determines the subjective perception data.

[0136] For example, the server performs time-series encoding on the physiological state change data to obtain the physiological state change characteristics. The server decodes the physiological state change characteristics to obtain the physiological preference information. The server combines the physiological preference information with the satisfaction level to obtain the subjective feeling data.

[0137] In the second part, the server determines the first adjustment data for the initial replica digital human based on the adjustment suggestion data and the device environment data, where the device environment data is used to describe the environment in which the sending device is located.

[0138] In one possible implementation, the device environment data includes device location information and device ambient light information. The server determines first image adjustment data and first motion adjustment data for the initial replica digital human based on the adjustment suggestion data. The server determines reference device adjustment data based on the device location information and the device ambient light information, where the reference device adjustment data matches the device location information and the device ambient light information. The server determines the first adjustment data based on the first image adjustment data, the first motion adjustment data, and the reference device adjustment data.

[0139] In order to explain the above embodiment more clearly, the above embodiment will be described in several parts below.

[0140] A. The server determines the first image adjustment data and the first action adjustment data of the initial replica digital human based on the adjustment suggestion data.

[0141] In one possible implementation, the server encodes the adjustment suggestion data based on an attention mechanism to obtain an adjustment suggestion data feature of the adjustment suggestion data. The server performs multiple rounds of iterative decoding on the adjustment suggestion data feature based on the attention mechanism to obtain the first image adjustment data and the first action adjustment data.

[0142] Among them, the above-mentioned encoding and decoding are realized through a first adjustment data determination model, which can convert the adjustment suggestion data into image adjustment data and action adjustment data. The first adjustment data determination model is trained based on multiple sample adjustment suggestion data and the labeled image adjustment data and labeled action adjustment data corresponding to each sample adjustment suggestion data. The first adjustment data determination model is an encoding and decoding model based on the attention mechanism. The embodiment of the present application does not limit the structure and training method of the first adjustment data determination model.

[0143] B. The server determines reference device adjustment data based on the device location information and the device ambient light information.

[0144] In one possible implementation, the server uses the device location information to query and obtain position adjustment data corresponding to the device location information. The server uses the device ambient light information to query and obtain ambient light adjustment data corresponding to the device ambient light information. The server determines the reference device adjustment data based on the position adjustment data and the ambient light adjustment data.

[0145] For example, the server uses the device location information to query the first relationship table to obtain the location adjustment data corresponding to the device location information. The server uses the device ambient light information to query the second relationship table to obtain the ambient light adjustment data corresponding to the device ambient light information. The server merges the location adjustment data with the ambient light adjustment data to obtain the reference device adjustment data.

[0146] The first relationship table stores multiple device location information and position adjustment data corresponding to each device information, and the second relationship table stores multiple device ambient light information and ambient light adjustment data corresponding to each device ambient light information. The position adjustment data is replica object adjustment data corresponding to the location device information, indicating how to adjust the replica digital human's image; the ambient light adjustment data is replica object adjustment data corresponding to the ambient light information, indicating how to adjust the replica digital human's image. The initial replica object is adjusted in conjunction with the device location information because users in the same location have similar preferences for replica objects, so the initial replica object is adjusted in conjunction with the device location information. The initial replica object is adjusted in conjunction with the ambient light adjustment data because the same rendering effect may have different visual effects under different ambient light conditions, so the initial replica object is adjusted in conjunction with the light environment data. The first and second relationship tables are configured by technicians based on actual conditions and are not limited in this embodiment of the present application.

[0147] C. The server determines the first adjustment data based on the first image adjustment data, the first action adjustment data, and the reference device adjustment data.

[0148] In a possible implementation, the server merges the first image adjustment data and the reference device adjustment data to obtain intermediate image adjustment data, and then splices the intermediate image adjustment data with the reference device adjustment data to obtain the first adjustment data.

[0149] Part three: The server determines second adjustment data for the initial replica digital human based on the personality description and the public evaluation.

[0150] In one possible implementation, the server determines second motion adjustment data for the initial replica digital human based on the personality description. The server determines third motion adjustment data and second image adjustment data for the initial replica digital human based on the public evaluation. The server determines second adjustment data for the initial replica digital human based on the second motion adjustment data, the third motion adjustment data, and the second image adjustment data.

[0151] Among them, personality is usually associated with action, so the second action adjustment data can be determined by using personality description. Public evaluation usually includes evaluation of image and evaluation of action, so it can be used to adjust image and action.

[0152] In order to explain the above embodiment more clearly, the above embodiment will be described in several parts below.

[0153] A. The server determines the second action adjustment data of the initial replica digital human based on the personality description.

[0154] In a possible implementation, the server extracts features from the personality description to obtain personality description features of the personality description, and determines the second action adjustment data based on the personality description features.

[0155] For example, the server encodes the personality description based on the attention mechanism to obtain the personality description feature of the personality description. The server performs multiple rounds of iterative decoding on the personality description feature based on the attention mechanism to obtain the second action adjustment data.

[0156] Among them, the above-mentioned encoding and decoding are realized through a second adjustment data determination model, which can convert the personality description into second action adjustment data. The second adjustment data determination model is trained based on multiple sample personality descriptions and the labeled action adjustment data corresponding to each sample personality description. The second adjustment data determination model is an encoding and decoding model based on the attention mechanism. The embodiment of the present application does not limit the structure and training method of the second adjustment data determination model.

[0157] B. The server determines the third action adjustment data and the second image adjustment data of the initial replica digital human based on the public evaluation.

[0158] In one possible implementation, the server encodes the public evaluation based on an attention mechanism to obtain a public evaluation feature of the public evaluation. The server performs multiple rounds of iterative decoding on the public evaluation feature based on the attention mechanism to obtain the second image adjustment data and the third action adjustment data.

[0159] Among them, the above-mentioned encoding and decoding are realized through a third adjustment data determination model, which can convert public evaluations into image adjustment data and action adjustment data. The third adjustment data determination model is trained based on multiple sample public evaluations and the labeled image adjustment data and labeled action adjustment data corresponding to each sample public evaluation. The third adjustment data determination model is an encoding and decoding model based on the attention mechanism. The embodiment of the present application does not limit the structure and training method of the third adjustment data determination model.

[0160] C. The server determines the second adjustment data of the initial replica digital human based on the second action adjustment data, the third action adjustment data and the second image adjustment data.

[0161] In one possible implementation, the server merges the second action adjustment data with the third action adjustment data to obtain intermediate action adjustment data, and then splices the intermediate action adjustment data with the second image adjustment data to obtain the second adjustment data of the initial replica digital human.

[0162] Part 4: The server determines the updated data of the replica digital human based on the subjective feeling data, the first adjustment data and the second adjustment data.

[0163] In one possible implementation, the server determines an adjustment coefficient based on the subjective perception data, and the adjustment data is used to control the adjustment range. The server fuses the first adjustment data with the second adjustment data to obtain target adjustment data. The server then fuses the adjustment coefficient with the target adjustment data to obtain the updated data for the replica digital human.

[0164] For example, the server performs feature extraction on the subjective perception data to obtain subjective perception data features of the subjective perception data. The server performs full connection and normalization on the subjective perception data features to obtain the adjustment coefficient. The server performs weighted fusion on the first adjustment data and the second adjustment data to obtain target adjustment data. The server multiplies the adjustment coefficient by the target adjustment data to obtain the updated data of the replica digital human.

[0165] Among them, the weights of weighted fusion are set by technical personnel according to actual conditions, and the embodiments of this application do not limit this.

[0166] Part 5: The server uses the replica digital human update data to update the initial replica digital human to obtain the target replica digital human.

[0167] In one possible implementation, the server uses the replica digital human update data to update the initial replica digital human display data of the initial replica digital human, obtaining the target replica digital human display data of the target replica digital human. Based on the target replica digital human display data, the server generates a virtual image of the target replica digital human.

[0168] The initial replica digital human display data includes multiple skeleton points, rendering parameters, and motion modes of the initial replica digital human. Updating the initial replica digital human display data includes at least one of adjusting the positions of the multiple skeleton points, adjusting the rendering parameters, and adjusting the motion modes.

[0169] The following describes how to adjust the positions of multiple skeleton points, the rendering parameters, and the action mode.

[0170] In some embodiments, the server obtains skeleton point update data from the replica digital human update data, and uses the skeleton point update data to adjust the positions of the plurality of skeleton points to obtain updated skeleton points.

[0171] In some embodiments, the server obtains rendering parameter update data from the replica digital human update data, and updates the rendering data using the rendering parameter update data to obtain updated rendering parameters.

[0172] In some embodiments, the server obtains motion mode adjustment data from the updated data of the replica digital human, and uses the motion mode adjustment data to adjust the motion mode to produce an updated motion mode.

[0173] It should be noted that the above adjustment data are all used to indicate the adjustment method. The form of the adjustment data is set by technical personnel according to actual conditions, and the embodiments of this application do not limit this.

[0174] 305. The server sends the target digital human display data of the target replica digital human to the sending device, so that the initiating device displays the target replica digital human based on the target digital human display data.

[0175] After the target replica digital human is displayed via the sending device, the sending device displays the target replica digital human based on the target digital human display data, which includes multiple skeletal points, rendering parameters, and action modes of the target replica digital human. After the sending device displays the target replica digital human, the target user can interact with the target replica digital human. For example, the target user can control the target replica digital human to perform specific actions or have a conversation with the target replica digital human. Action execution can be achieved by controlling the digital human using skeletal points as described in related technologies, and conversation can be achieved through conversation models as described in related technologies. This is not limited in the present embodiment.

[0176] All of the above optional technical solutions can be combined in any way to form optional embodiments of the present application, and will not be described in detail here.

[0177] Through the technical solution provided by the embodiment of the present application, in response to a target user's request to replicate a target historical figure, the first historical figure data of the target historical figure, the first user data of the target user, and the device performance data of the device sending the replication request are obtained. Based on the first historical figure data, the first user data, and the device performance data, an initial replicated digital human is synthesized and displayed, thereby achieving the preliminary generation and display of the replicated digital human. In response to the target user's request to adjust the initial replicated digital human, the second user data of the target user, the device environment data of the sending device, and the second historical figure data of the target historical figure are obtained. Based on the second user data, the device environment data, and the second historical figure data, the initial replicated digital human is updated to obtain a target replicated digital human, thereby achieving convenient replication of the replicated digital human. The target replicated digital human can subsequently be used to interact with the target user, thereby increasing the target user's satisfaction with the target virtual digital human, thereby improving the user experience.

[0178] Figure 4 This is a structural diagram of a real-time interactive synthesis processing device based on a replica digital human provided by an embodiment of the present application, see Figure 4 The device includes: an acquisition module 401, a synthesis and display module 402 and an update module 403.

[0179] Acquisition module 401 is used to respond to the target user's request for replicating the target historical figure, obtain the first historical figure data of the target historical figure, the first user data of the target user, and the device performance data of the device sending the replica request, the first historical figure data includes the life information and image record of the target historical figure, and the first user data includes the user image data and user action data of the target user.

[0180] The synthesis and display module 402 is used to synthesize and display an initial replica digital human based on the first historical figure data of the target historical figure, the first user data and the device performance data.

[0181] The acquisition module 401 is also used to respond to the target user's adjustment request for the initial replica digital human, obtain the target user's second user data, the device environment data of the sending device, and the second historical figure data of the target historical figure, wherein the second user data includes the target user's physiological data and adjustment suggestion data, and the second historical figure data includes the personality description and public evaluation of the target historical figure.

[0182] The updating module 403 is used to update the initial replica digital human based on the second user data, the device environment data and the second historical figure data to obtain a target replica digital human.

[0183] In one possible implementation, the synthesis and display module 402 is configured to determine first replica digital human image data based on the image record of the target historical figure and the user image data of the target user. First replica digital human motion data is determined based on the biographical information and the user motion data. Second replica digital human image data is determined based on the device performance data, the biographical information, and the first replica digital human image data. Based on the first replica digital human motion data and the second replica digital human image data, the initial replica digital human is generated. The initial replica digital human is displayed.

[0184] In one possible implementation, the image record includes an image image and image description text. The user image data includes a first user facial image, a second user facial image, and a desired image description text of the target historical figure. The second user facial image is obtained by the target user editing the first user facial image. The synthesis display module 402 is configured to determine first reference image data of the target historical figure based on the image image, the image description text, and the desired image description text. Second reference image data of the target user is determined based on the first user facial image and the second user facial image. The first replica digital human image data is determined based on the first reference image data and the second reference image data.

[0185] In one possible implementation, the user motion data includes current user motion data and historical user motion data. The synthesis and display module 402 is configured to perform data recognition on the biographical data to obtain motion description data of the target historical figure in the biographical data. Reference motion data is determined based on the motion description data and the current user motion data. The first replica digital human motion data is determined based on the historical user motion data and the reference motion data.

[0186] In one possible implementation, the synthesis and display module 402 is configured to extract features from the biographical data to obtain biographical features of the biographical data. Third-replica digital human image data is determined based on the biographical features and the first-replica digital human image data. Second-replica digital human image data is determined based on the device performance data and the third-replica digital human image data.

[0187] In one possible implementation, the update module 403 is configured to determine the target user's subjective perception data of the initial replica digital human based on the physiological data and the adjustment suggestion data. First adjustment data for the initial replica digital human is determined based on the adjustment suggestion data and the device environment data, where the device environment data describes the environment in which the sending device is located. Second adjustment data for the initial replica digital human is determined based on the personality description and the public evaluation. Replica digital human update data is determined based on the subjective perception data, the first adjustment data, and the second adjustment data. The initial replica digital human is updated using the replica digital human update data to obtain the target replica digital human.

[0188] In one possible implementation, the physiological data includes first physiological data and second physiological data, wherein the first physiological data is collected before the initial replica digital human is displayed, and the second physiological data is collected after the initial replica digital human is displayed. The updating module 403 is configured to determine the target user's physiological state change data based on the first physiological data and the second physiological data. Based on the adjustment suggestion data, the target user's satisfaction with the initial replica digital human is determined. The subjective perception data is determined based on the physiological state change data and the satisfaction level.

[0189] In one possible implementation, the device environment data includes device location information and device ambient light information. The updating module 403 is configured to determine first image adjustment data and first motion adjustment data for the initial replica digital human based on the adjustment suggestion data. Reference device adjustment data is determined based on the device location information and the device ambient light information, where the reference device adjustment data matches the device location information and the device ambient light information. The first adjustment data is determined based on the first image adjustment data, the first motion adjustment data, and the reference device adjustment data.

[0190] In one possible implementation, the updating module 403 is configured to determine second motion adjustment data for the initial replica digital human based on the personality description. Based on the public evaluation, the updating module 403 determines third motion adjustment data and second image adjustment data for the initial replica digital human. Based on the second motion adjustment data, the third motion adjustment data, and the second image adjustment data, the updating module 403 determines second adjustment data for the initial replica digital human.

[0191] It should be noted that the above-mentioned embodiments of the real-time interactive synthesis processing device based on a replica digital human, when performing the processing of a replica digital human, only illustrate the division of the above-mentioned functional modules. In actual applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure of the computer device can be divided into different functional modules to complete all or part of the functions described above. In addition, the real-time interactive synthesis processing device based on a replica digital human provided in the above-mentioned embodiments and the real-time interactive synthesis processing method based on a replica digital human are based on the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.

[0192] Through the technical solution provided by the embodiment of the present application, in response to a target user's request to replicate a target historical figure, the first historical figure data of the target historical figure, the first user data of the target user, and the device performance data of the device sending the replication request are obtained. Based on the first historical figure data, the first user data, and the device performance data, an initial replicated digital human is synthesized and displayed, thereby achieving the preliminary generation and display of the replicated digital human. In response to the target user's request to adjust the initial replicated digital human, the second user data of the target user, the device environment data of the sending device, and the second historical figure data of the target historical figure are obtained. Based on the second user data, the device environment data, and the second historical figure data, the initial replicated digital human is updated to obtain a target replicated digital human, thereby achieving convenient replication of the replicated digital human. The target replicated digital human can subsequently be used to interact with the target user, thereby increasing the target user's satisfaction with the target virtual digital human, thereby improving the user experience.

[0193] Figure 5 This is a schematic diagram of the structure of a server provided in an embodiment of the present application. The server 500 may have relatively large differences due to different configurations or performances, and may include one or more processors (Central Processing Units, CPU) 501 and one or more memories 502, wherein the one or more memories 502 store at least one computer program, and the at least one computer program is loaded and executed by the one or more processors 501 to implement the methods provided in the above-mentioned various method embodiments. Of course, the server 500 may also have components such as a wired or wireless network interface, a keyboard, and an input / output interface for input and output. The server 500 may also include other components for implementing device functions, which will not be described in detail here.

[0194] In an exemplary embodiment, a computer-readable storage medium is also provided, such as a memory including a computer program. The computer program can be executed by a processor to implement the real-time interactive synthesis processing method based on replicating a digital human in the above embodiment. For example, the computer-readable storage medium can be a read-only memory (ROM), a random access memory (RAM), a compact disc (CD-ROM), a magnetic tape, a floppy disk, or an optical data storage device.

[0195] In an exemplary embodiment, a computer program product or computer program is also provided, which includes a program code, which is stored in a computer-readable storage medium. The processor of a computer device reads the program code from the computer-readable storage medium, and the processor executes the program code, so that the computer device performs the above-mentioned real-time interactive synthesis processing method based on replicating digital humans.

[0196] In some embodiments, the computer program involved in the embodiments of the present application may be deployed and executed on a computer device, or on multiple computer devices located at one location, or on multiple computer devices distributed at multiple locations and interconnected through a communication network. Multiple computer devices distributed at multiple locations and interconnected through a communication network may constitute a blockchain system.

[0197] Those skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware, or may be accomplished by a program to instruct the relevant hardware, and the program may be stored in a computer-readable storage medium, and the above-mentioned storage medium may be a read-only memory, a disk or an optical disk, etc.

[0198] The above are only optional embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should be included in the scope of protection of the present application.

Claims

1. A real-time interactive synthesis processing method based on replicating digital humans, characterized in that: The method comprises: In response to a target user's request for a reproduction of a target historical figure, obtaining first historical figure data of the target historical figure, first user data of the target user, and device performance data of a device sending the reproduction request, wherein the first historical figure data includes biographical information and an image record of the target historical figure, and the first user data includes user image data and user action data of the target user; synthesizing and displaying an initial replica digital human based on the first historical figure data of the target historical figure, the first user data, and the device performance data; In response to the target user's request for adjustment of the initially replicated digital human, obtaining second user data of the target user, device environment data of the sending device, and second historical figure data of the target historical figure, wherein the second user data includes physiological data of the target user and adjustment suggestion data, and the second historical figure data includes a personality description and public evaluation of the target historical figure; Based on the second user data, the device environment data and the second historical figure data, the initial replica digital human is updated to obtain a target replica digital human.

2. The method according to claim 1, characterized in that The synthesizing and displaying the initial replica digital human based on the first historical figure data of the target historical figure, the first user data, and the device performance data includes: Determining first replica digital human image data based on the image record of the target historical figure and the user image data of the target user; Determining first replica digital human motion data based on the biographical information and the user motion data; Determining second replica digital human image data based on the device performance data, the biographical information, and the first replica digital human image data; generating the initial replica digital human based on the first replica digital human motion data and the second replica digital human image data; The initial replica digital human is displayed.

3. The method according to claim 2, characterized in that The image record includes an image image and image description text, the user image data includes a first user facial image, a second user facial image, and a text describing the desired image of the target historical figure, the second user facial image being obtained by the target user editing the first user facial image, and determining the first replica digital human image data based on the image record of the target historical figure and the user image data of the target user, including: Determining first reference image data of the target historical figure based on the image image, the image description text, and the expected image description text; Determining second reference image data of the target user based on the first user facial image and the second user facial image; The first replica digital human image data is determined based on the first reference image data and the second reference image data.

4. The method according to claim 2, characterized in that The user action data includes current user action data and historical user action data. The determining of the first replica digital human action data based on the biographical information and the user action data includes: Performing data recognition on the biographical data to obtain action description data of the target historical figure in the biographical data; Determining reference action data based on the action description material and the current user action data; Based on the historical user motion data and the reference motion data, the first replicated digital human motion data is determined.

5. The method according to claim 2, characterized in that The determining of the second replica digital human image data based on the device performance data, the biographical information, and the first replica digital human image data includes: Extracting features from the biographical information to obtain biographical information features of the biographical information; Determining third replica digital human image data based on the biographical information features and the first replica digital human image data; The second replica digital human image data is determined based on the device performance data and the third replica digital human image data.

6. The method according to claim 1, characterized in that The updating of the initial replica digital human based on the second user data, the device environment data, and the second historical figure data to obtain a target replica digital human includes: Determining the target user's subjective feeling data on the initial replica digital human based on the physiological data and the adjustment suggestion data; Determining first adjustment data for the initial replica digital human based on the adjustment suggestion data and the device environment data, wherein the device environment data is used to describe the environment in which the sending device is located; Determining second adjustment data for the initial replica digital human based on the personality description and the public evaluation; Determining update data of the replica digital human based on the subjective feeling data, the first adjustment data, and the second adjustment data; The initial replica digital human is updated using the replica digital human update data to obtain the target replica digital human.

7. The method according to claim 6, characterized in that The physiological data includes first physiological data and second physiological data, wherein the first physiological data is physiological data collected before displaying the initial replica digital human, and the second physiological data is physiological data collected after displaying the initial replica digital human. Determining the target user's subjective feeling data of the initial replica digital human based on the physiological data and the adjustment suggestion data includes: determining physiological state change data of the target user based on the first physiological data and the second physiological data; Determining the target user's satisfaction with the initial replica digital human based on the adjustment suggestion data; The subjective feeling data is determined based on the physiological state change data and the satisfaction level.

8. The method according to claim 6, characterized in that The device environment data includes device location information and device ambient light information. The determining of first adjustment data for the initial replica digital human based on the adjustment suggestion data and the device environment data includes: Determining first image adjustment data and first action adjustment data of the initial replica digital human based on the adjustment suggestion data; determining reference device adjustment data based on the device location information and the device ambient light information, where the reference device adjustment data matches the device location information and the device ambient light information; The first adjustment data is determined based on the first image adjustment data, the first action adjustment data, and the reference device adjustment data.

9. The method according to claim 6, characterized in that The determining of second adjustment data for the initial replica digital human based on the personality description and the public evaluation includes: Determining second action adjustment data of the initial replica digital human based on the personality description; Determining third action adjustment data and second image adjustment data of the initial replica digital human based on the public evaluation; Based on the second action adjustment data, the third action adjustment data and the second image adjustment data, second adjustment data of the initial replica digital human is determined.

10. A real-time interactive synthesis processing system based on replicating digital humans, characterized in that: The system comprises: an acquisition module, configured to, in response to a target user's request for replicating a target historical figure, acquire first historical figure data of the target historical figure, first user data of the target user, and device performance data of a device sending the replica request, wherein the first historical figure data includes biographical information and an image record of the target historical figure, and the first user data includes user image data and user action data of the target user; a synthesis and display module, configured to synthesize and display an initial replica digital human based on the first historical figure data of the target historical figure, the first user data, and the device performance data; The acquisition module is further configured to, in response to the target user's request for adjustment of the initially replicated digital human, acquire second user data of the target user, device environment data of the sending device, and second historical figure data of the target historical figure, wherein the second user data includes physiological data of the target user and adjustment suggestion data, and the second historical figure data includes a personality description and public evaluation of the target historical figure; An updating module is used to update the initial replica digital human based on the second user data, the device environment data and the second historical figure data to obtain a target replica digital human.

Citation Information

Patent Citations

  • Virtual digital human construction method and system

    CN117828320A

  • Human-computer interaction method and device, electronic equipment and computer storage medium

    CN118567602A