Distributed node re-evolution system for virtual organism multi-modal data
The distributed node re-evolution system solves the problems of high production cost, difficulty in updating and iteration, and poor ownership security of digital humans, and realizes flexible customization and secure digital human evolution, reducing costs and ensuring ownership security.
Patent Information
- Application Number
- CN202411251650.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-08
- Publication Date
- 2026-03-10
AI Technical Summary
The existing digital human production model suffers from high costs, difficulty in updating and iterating, and difficulty in guaranteeing ownership and security. The centralized development model has drawbacks such as high resource consumption, high update costs, and difficulty in protecting rights.
This paper proposes a distributed node re-evolution system for multimodal data of virtual organisms. The system consists of a data source supply and issuance end, distributed reorganization and secondary training nodes, a server end, and an electronic ledger of ownership. It realizes distributed iterative training and ownership allocation, supports secondary modification and re-issuance of digital humans within the ownership, and combines blockchain technology to ensure security and rights.
It enables flexible customization and secure evolution of digital humans on distributed terminals, reduces production costs, improves update efficiency, ensures the security and sustainable development of ownership, and supports dynamic adjustment of weights and ownership.
Smart Images

Figure CN121639872A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the interdisciplinary field of biotechnology and information technology models and architectures, particularly the development and fabrication methods and continuous evolution systems of digital human or digital organism models. Background Technology
[0002] With the rapid development of the information society and the continuous improvement of economic and technological levels, talent has become an extremely valuable resource in society. However, individual capabilities are limited, and it is impossible to operate efficiently in multiple locations simultaneously. For example, a celebrity cannot hold concerts in different venues at the same time. Similarly, professionals such as lawyers often lose their ability to continue practicing as they age, even though these older individuals typically possess the sharpest working minds. How to preserve the appearance, demeanor, and thought patterns of these professionals, and even further develop their wisdom and enhance their appeal, has become a significant technological challenge. This need has given rise to the concepts of "digital human" and "virtual human." With the continuous evolution of artificial intelligence technology, the emerging technology of digital human has thus come into being. In February 2024, the story of musician Bao Xiaobai using AI to "revive" his daughter circulated on the internet. The daughter recreated using digital technology could have simple conversations with her parents and even sing a birthday song with them. The father, grieving the loss of his beloved daughter, returned to school to pursue a doctoral degree, studied and initially mastered some AI technology, and after more than two years of production, finally achieved his goal.
[0003] It is undeniable that the development of digital humans currently requires a fairly high level of technical expertise, thus it is mostly concentrated in specialized institutions, such as Microsoft's Xiaoice and miHoYo's Luming. Moreover, in addition to the technical hurdles, this centralized development model also has many drawbacks.
[0004] First, the current process of creating digital humans is cumbersome and resource-intensive. Traditional digital human creation involves multiple stages such as complex modeling, rigging, and animation, meaning that production costs are high and the cycle is long. For example, a silicon-based intelligence company claims that its "AI resurrection" business falls under the categories of "life cloning" and "digital immortality," offering various pricing plans ranging from "photo-based" to "AI clones on monthly or annual payments." For high-end customized services, the cost to users typically exceeds 150,000 yuan. If users are willing to invest even more, although they may get a digital life experience closer to reality, maintaining and upgrading the technical support for this digital "life" also requires continuous and high resource investment.
[0005] Secondly, the difficulties in updating and iterating digital humans are equally obvious. Once an existing digital human model is completed, its appearance design is fixed and it is difficult to add functions and movements. Due to the centralized development nature of the project, subsequent updates and feature expansions face enormous cost challenges. Most development projects rarely undergo further improvements after the initial development.
[0006] Finally, as a virtual asset directly belonging to the original entity, the ownership and security of digital humans are extremely difficult to guarantee. Once a digital human is created, its rights and data are easily misappropriated by others or third-party companies. In many cases, the company that created the digital human owns the entire source code, and once the original entity dies, the digital human may be abused regardless of its wishes. This has already sparked protests from Coco Lee's family and widespread public attention in the recent "resurrection" incident involving singer Coco Lee.
[0007] Therefore, this invention aims to propose a distributed, semi-decentralized, simpler, faster, safer, and sustainably evolving system—a distributed node re-evolution system for multimodal data of virtual organisms. This system does not require a specialized manufacturing process but instead utilizes a large number of distributed terminals to achieve distributed evolutionary iteration. The system possesses continuous iterative training capabilities and ensures the security of biological data, maintaining the rights of owners and possessing a unique weighting and ownership confirmation mechanism, creating a more reliable protection system for digital humans. Furthermore, in certain situations, a method proposed in this invention can also be used for digital pets and digital animals. Summary of the Invention
[0008] To address the aforementioned issues, this invention proposes a distributed node re-evolution system for virtual organism multimodal data. This system comprises a data source supply and distribution end (A-end), distributed recombination and secondary training nodes (B-end), a server end (S-end), a database, and an electronic ledger of ownership. The A-end is responsible for collecting user multimodal biological data and setting distributed node modes and ownership allocations; the B-end supports secondary customization, use, and training; and the S-end processes data, trains models, and manages resources. The system begins with primary users on the A-end creating projects, uploading biological multimodal data to the server, and setting initial ownership ratios. The server records information, reconstructs 3D models, and generates key public content such as voice clones and dialogue models. Secondary users perform customization, recombination, and reprocessing through the B-end to generate digital human forks. This system supports secondary modifications and redistribution of digital humans within ownership, enabling the inheritance, trading, and continuous evolution of biological data through ownership and weighting. Through deconstruction, ownership allocation, reconstruction, distributed re-customization, and retraining, the system promotes the continuous evolution and diversification of digital humans (or digital organisms).
[0009] The system has the following characteristics: a. In terms of system composition, this system consists of multiple components, including: A-side: Data source supply and initial distribution end, supporting the collection of users' multimodal biological data, setting distributed node rule modes, and setting weights and ownership allocation; B-end: Distributed reorganization and secondary training nodes, supporting personalized customization, use, and training based on the initial model; S-side: Server-side, responsible for data processing, key content generation, model training, and resource management; Database: Stores user data and model parameters; A weighted and ownership-based electronic ledger (which can be a blockchain ledger or integrated with a database) records the weights and ownership of virtual humans (or digital organisms). As these virtual humans (or digital organisms) evolve or upgrade through algorithms on distributed terminals, adjustments to the weight allocation can directly influence the direction and outcome of this evolution. For example, the adoption and implementation speed of a new skill or design may need to be determined based on the weight allocation.
[0010] b. The specific functions and operation process of the system are as follows: b-1.1.) First, the data source supply and distribution end (whose users are referred to as Level 1 users) initiates the creation of a project, collects and uploads multimodal data of the organism. Taking the human body as an example, the deconstructed organism data includes, but is not limited to: photos or videos: users can take facial images and videos through a camera, or select files to upload; audio recording: users record sound through a microphone for subsequent sound synthesis and speech model training; video interaction: motion capture and facial expression collection for subsequent enrichment of the digital organism's expressions and movements; the uploaded multimodal data will be transmitted to the server for subsequent processing.
[0011] (b-1.2.) During project creation, users need to set an initial distributed ownership (or weight) issuance ratio, allocate weights or ownership to first-level and second-level users, and record them in the ownership (or weight) electronic ledger. For example, the ownership (or weight) of a virtual person can be divided into N parts and allocated according to specific rules. If a user or distributed node obtains M blocks, the ratio is M / N. These ownership (or weight) can be changed subsequently according to certain rules, such as the rule example given in the specific implementation case below. Ownership (or weight) can also be inherited or traded and stored on a dedicated property rights blockchain to realize the inheritance and trading of ownership (or weight) of virtual people and digital organisms. The ownership confirmation model can adopt either the QA or QB model: The QA model combines symbolic mapping coding and blockchain technology. By integrating blockchain with symbolic mapping coding (SMC), this model allows all data sharing participants to generate a unique symbolic mapping table (SMT) and embed its digital identity code into the byte sequence of the shared data. This makes the claim of rights no longer linked to the data content and type, allowing all participants to jointly monitor the delivery and access of shared data, thereby achieving public verification of data rights. In contrast, the QB model is based on secret sharing and symbolic mapping coding. Addressing the pressure of storing symbolic mapping tables (SMT) for large-scale data and the issue of user privacy leakage, it uses a fuzzy C-means clustering algorithm to cluster the data and uses SMT to generate an asymmetric fingerprint to achieve data ownership confirmation. It combines secret sharing technology to complete data distribution and acquisition, uses a randomly verifiable algorithm to provide identity proof for users, and ensures data traceability and repudiation by recording every step of the sharing process in the blockchain. Of course, with technological advancements, more blockchain ownership confirmation models will be available in the future.
[0012] b-1.3.) The server creates the core part of the virtual human or digital organism for public use (such as the head and face reconstructed by the server without hair and body in 3D). Users train and learn the generated language model through dialogue and interaction. In particular, when a first-level user creates himself as a virtual human, the user interacts through dialogue, video communication, and providing documents for machine learning, and transmits his own characteristics to the server for training and learning.
[0013] (b-2.1.) Subsequently, the key functions of the server include recording user-generated information, managing the operation of the entire system, and reconstructing 3D images. The server uses algorithmic models to automatically generate 3D head models or partial or complete body models.
[0014] (b-2.2.) The server will also receive audio data uploaded by users and generate a voice clone model with specific speaker characteristics and timbre.
[0015] (b-2.3.) Furthermore, the server builds an independent dialogue model for each virtual digital human based on the content uploaded by the user. To enhance security and ensure user data safety, sandboxing, virtual containers, or virtual machine technologies (such as Docker or Kubernetes) are installed on the server side to create independent user spaces, providing a dedicated management environment for each digital human model. Sandboxed storage ensures independence and isolation between models, and multiple servers are connected to an IPFS node network. Data is split and stored across all servers, ensuring that no single server can form a complete dataset.
[0016] (b-3.1.) Any device with the distributed recombination and secondary training node software installed (which can be packaged into the same app as the data source supply and distribution software) will, after user authentication, confirm the user's ownership and usage permissions. Based on the allocation information provided by the server, the user will download the corresponding biological data or a semi-finished model generated by the server based on the relevant data, and then customize and recombine it on the distributed terminals. These distributed nodes are not equal in weight; as different secondary users log in to the distributed recombination and secondary training nodes, these nodes also acquire different weights.
[0017] (b-3.2.) At the distributed reorganization and secondary training nodes, secondary users can reprocess the digital human, including but not limited to the production of movements and dances, the modification and upgrading of clothing and makeup, and further secondary training and development through dialogue between users and the digital human. Thus, the distributed terminal reorganizes or replaces the content in digital form.
[0018] For example, on a distributed client, users can train a prototype dialogue language model for the digital human through dialogue and manual uploading of relevant document data. The server uses deep learning algorithms to generate a preliminary dialogue language model prototype on the server side and assigns a voice timbre to the digital human model based on the reference audio uploaded by the user. The overall architecture of the voice timbre simulation model consists of a posterior encoder, a prior encoder, a decoder, a discriminator, and a random duration predictor. Then, the audio output of the dialogue model is generated through text-to-speech technology. The text that the dialogue model replies to the user is processed and converted into audio to achieve natural speech interaction. During the optimization process, an adversarial training framework is constructed to optimize the conditional variational autoencoder to improve the quality and realism of the generated speech. During training, a joint loss function is used, including reconstruction loss, KL divergence loss, and adversarial loss. The reconstruction loss and KL divergence loss are used to optimize the conditional variational autoencoder, and the adversarial loss is used to optimize the discriminator network.
[0019] For example, after the virtual human head is generated, different distributed clients can assign different facial expressions and different animations to the virtual human. One specific method is as follows: 1. Design a video face feature point tracking algorithm. First, use a camera to capture images of facial expressions. Then, use a cascaded convolutional neural network algorithm with multi-task learning and a cascaded gradient-enhanced regression tree algorithm to detect faces and extract feature points; 2. Set facial expression animation control parameters. In the animation control parameters, use the EPnP algorithm to solve for head pose and generate the corresponding head rotation matrix and translation vector. Use a support vector machine to build an AU detector and an intensity value regressor. By establishing a mapping relationship between AU and expression basis, the required facial expression parameters are obtained; 3. Design a virtual facial expression driver. Transmit the facial expression control parameters to Unity3D software through a plugin to drive the animation model to generate the corresponding facial expression animation.
[0020] After the above process, and further processed by different distributed nodes, these digital humans will generate different branches, similar to a mutation of species in nature.
[0021] c. This system uses the above process to deconstruct, assign ownership to, reconstruct, and distribute the re-customization and retraining of biological data. See the appendix in the instruction manual. Figure 1 After the reprocessed digital human generates different forks through the aforementioned distributed terminal, secondary users of the distributed terminal can further modify their own digital human forks (including but not limited to dialogue model training, action and expression assignment, and appearance optimization) and redistribute them. They then redistribute all M blocks belonging to that secondary user to tertiary users. If a new tertiary user acquires Q blocks, the tertiary user's ownership (or weight) ratio for that fork is Q / M, and its ownership (or weight) ratio relative to the entire digital human is Q / N. After the fork redistribution, the child node becomes the new master node. The reprocessed digital human can then generate different forks, repeating this cycle to achieve continuous evolution of biological data after leaving the body. The aforementioned weights can also be used as parameters in algorithms for selecting evolutionary paths within the system. We know that in the process of species evolution in nature, changes in the proportion of different genotypes within a population are the basis of evolution. When the proportion of a certain genotype in a population increases or decreases, it means that the frequency of that gene in the population has changed. This is a direct manifestation of evolution. For example, an increase in the proportion of blue-winged individuals in a butterfly population indicates that the frequency of the blue-wing gene has increased in the population, thus indicating evolution. Therefore, how to assign weights can play different roles in digital human evolution models after different secondary bifurcations.
[0022] For example, a weighted mechanism can be used to achieve the following automatic distributed evolution: After a primary user creates the virtual avatar of the parent avatar according to the above process, these virtual avatars on different distributed nodes are inspired by AI rules with different boundary conditions and random parameters, and are given different secondary modifications (including but not limited to different styles of clothing, different route styles, different language models, and different dynamic expressions). These nodes are then used to interact with different audiences on relevant platforms, and different modifications are subsequently executed based on the feedback from their respective audiences. The branches with better feedback (such as receiving more likes and collections) will receive greater weight and will be executed according to certain pre-set rules. At a certain stage in this process, some of the secondary branches that have gained greater weight will become new master nodes according to certain rules, thus forming an automatic evolution and elimination mechanism. From these new secondary master nodes, the next level of branching and distributed evolution will begin again, and the cycle will repeat.
[0023] When the target organism is a human, after receiving multimodal biological data (including facial images, videos, or audio), the server may not reconstruct the complete organism data. Instead, it uses a generative adversarial network-based algorithm to generate only a 3D head model without hair (the generated 3D head model does not include hair and the area below the neck; this part will be customized and added autonomously by distributed clients). The client then downloads the generated head model, combines the pre-bonded 3D human body with a pre-adapted hairstyle, and uses accessories such as necklaces, headbands, scarves, and ties to connect the neck area. Appropriate clothing is added from the client's local storage or the server. The advantages of this approach are that it speeds up processing and gives distributed terminals more flexibility.
[0024] When the organism is an animal, such as a pet, the system can also disperse the organism's entity elements through the following deconstruction and reconstruction process, followed by distributed reconstruction and weight allocation between nodes: First, the organism's multimodal information is deconstructed and broken down into facial information (including but not limited to torso shape information, hair information, voice information, personality habits, and reaction response characteristics). After deconstruction, weight segmentation as described in the first step is performed to achieve flexible weight allocation. Subsequently, during the 3D reconstruction process, a facial feature domain transformation network model is established by combining an autoencoder network and a transformation network to first convert the animal face into a virtual human face, thus facilitating the direct application of face recognition and reconstruction models. Finally, the 3D animal face is reconstructed based on the inverse transformation network model. Similarly, for many other types of information, transformation is also performed to adapt them to relatively mature algorithms for human processing, and then the animal form is reconstructed.
[0025] The inventiveness of this invention The inventiveness of this invention is mainly reflected in the following aspects.
[0026] 1. A distributed node recombination and evolution system for multimodal biological data: Innovation: This system enables flexible customization, recombination, and retraining of multimodal biological data through a distributed node architecture. This distributed architecture not only improves data processing efficiency but also enhances system scalability and flexibility, increases data and ownership security, and allows different users to personalize their data on their respective nodes, promoting the diversified development of digital organisms.
[0027] 2. Dynamic changes in ownership and weighting and distributed ledger systems: Innovation: The system introduces an electronic ledger of ownership and a weight allocation mechanism, allowing users to precisely divide and dynamically adjust the ownership of virtual organisms (such as digital humans). Simultaneously, combining blockchain technology, a rights confirmation mechanism algorithm is proposed to ensure the accuracy and security of ownership and weights. When virtual humans freely evolve on distributed terminals according to the set algorithm rules, these weights will serve as key parameters input to the algorithm, guiding the direction and speed of evolution. For example, the learning efficiency of a new skill or the adoption speed of a new appearance design may be directly related to the magnitude of the corresponding attribute weights. If a virtual human has a higher innovation weight, it may more quickly adopt and efficiently implement new design concepts or learn new skills, thereby gaining a competitive advantage in the direction of evolution.
[0028] 3. The continuous evolution and diversification of virtual organisms: Innovation: Through deconstruction, ownership allocation, reconstruction, and distributed re-customization and retraining, the system supports the continuous evolution and diversified development of virtual organisms. This mechanism enables digital organisms to continuously adapt to external feedback, audience feedback, and environmental changes, generating new characteristics and forms.
[0029] 4. High-degree-of-freedom deconstruction and reconstruction of biological data: Innovations: A method for deconstructing, disassembling, assigning ownership to, reconstructing, and distributing the customization and retraining of biological data was proposed, enabling the continuous evolution of biological data after it leaves the body. Furthermore, when processing human data, a 3D head model without hair was generated, and other parts were autonomously customized and added on a distributed terminal, accelerating processing speed and increasing the degree of freedom.
[0030] 5. Cross-species algorithmic adaptability: Innovation: This system is not only suitable for processing human data, but also, through algorithmic transformation and processing, for the deconstruction and reconstruction of multimodal information from other types of organisms such as animals. The inspiration comes from the Fourier time-domain to frequency-domain transform in physics. This cross-species applicability demonstrates the system's powerful algorithmic adaptability and scalability, providing possibilities for the creation and evolution of more types of digital organisms in the future.
[0031] Specific Implementation Examples of the Invention This invention has been successfully implemented in a specific case. As a specific implementation case, based on the above basic architecture, this implementation case also makes the following settings: In terms of the composition of the instance system, it includes A-end (data source supply and initial issuance end, including cross-platform software that supports Windows, Mac, iPhone, Android and other system platforms), B-end (distributed reorganization and secondary training nodes, also including cross-platform software that supports Windows, Mac, iPhone, Android and other system platforms), S-end (server end), as well as database, storage bucket, ownership ledger blockchain, etc.
[0032] Among them: A-end: Data source supply and initial issuance end, used to collect users' multimodal biological data and set initiation rules and weight attributes and other parameters.
[0033] B-end: Distributed reorganization and secondary training nodes, supporting personalized customization and training based on the initial model.
[0034] S-side: Server-side, responsible for data processing, key content generation, model training, and resource management.
[0035] Database: Stores user data, model information, and training records.
[0036] Storage buckets: These store 3D data of digital humans or digital organisms, virtual body components generated from models, and resources such as clothing for digital humans. Storage can also be set up on the server side, utilizing server-side storage space. This example used a dedicated Amazon Cloud Storage bucket solution during internal testing, primarily for cost-effectiveness.
[0037] Ownership and Weighting Ledger Blockchain: Records the ownership and weighting of digital humans or digital organisms, ensuring transparency in ledger recording and security in transactions. This example uses a data ownership confirmation model based on symbolic mapping encoding and blockchain.
[0038] To address the issues of record tampering, data leakage, and applicability only to specific data types in traditional data ownership schemes, this model combines blockchain with Symbol Mapping Coding (SMC). Through SMC technology, each party sharing data can generate a unique Symbol Mapping Table (SMT) and encode its digital identity into the byte sequence of the shared data, thus making the declaration of permissions independent of the data content and type.
[0039] Part 1. Data Source Supply and Distribution Data source supply and distribution software is packaged in software and applications (APP). For example, when using an iPhone 15 as the implementation device, after installing the data source supply and distribution software, users can upload biological multimodal data in the following ways: - Taking photos and videos: Users can take facial images and videos with the iPhone 15 camera to capture detailed facial expression information, and extract actions from the video or video through the Mediapipe algorithm. They can take selfies or take pictures of others. The front and rear cameras of the phone can meet this collection needs.
[0040] - Audio Recording: Users record clear audio through a microphone for use in subsequent sound synthesis and speech model training.
[0041] The multimodal information of biological data in this specific implementation case may also include: - Facial expression information: including facial expressions, lip shapes, etc., which can be used to generate digital facial expression models.
[0042] - Personality traits: Used to generate user personality profiles through user interaction and behavior analysis.
[0043] - Mindset: Based on users' decision-making and problem-solving approaches, used to train intelligent decision-making models.
[0044] - Transaction experience: The user's past experience in handling transactions.
[0045] - Game Theory Patterns: Analyzing user behavior patterns in games or competitions.
[0046] - Knowledge reserves: The user's accumulated knowledge in a specific field.
[0047] - Memory information: The user's interaction history with the system, used to optimize the system's personalized features.
[0048] Through the deconstruction process, we disperse the various elements of an organism, making it more suitable for subsequent model recombination and iterative evolution at various distributed terminals.
[0049] The uploaded multimodal data will be transmitted to the server. The server uses a specialized four-step algorithm model based on generative adversarial networks (see below) to automatically generate a 3D head model, supporting rich facial expression features. After the model is generated, the user can confirm it. If dissatisfied, they can request to change the content and regenerate it through the system's feedback mechanism.
[0050] In addition, users can manually upload document data for training, such as manuscripts, speeches, diaries, etc. provided by the user, so as to generate a preliminary prototype of the dialogue language model on the server side using deep learning algorithms.
[0051] In the subsequent interaction phase, users can conduct periodic interaction training using the initial dialogue language model to refine the model and thus improve its personalization.
[0052] In this implementation case, during project creation, users need to set an initial distributed weight and ownership issuance ratio, which is then recorded on the property rights blockchain. For example, the weight and ownership of the virtual human can be divided into N parts, allocated according to specific rules, with a user receiving M blocks in a ratio of M / N. This ownership block not only supports inheritance but also allows for trading, greatly enhancing the economic value of the digital human and assigning different weights to different secondary evolutionary forks. This provides significant freedom for subsequent evolutionary rules, allowing for different functions to be used to apply different weights and achieve different evolutionary paths for the entire digital human or digital organism.
[0053] Part Two. Server-Side Functions In this example, the key functions of the server side include: 1. Recording user-generated information, managing the operation of the system, and providing services such as user system, economic system, weight adjustment system, distributed support, evolution rule setting, and multi-user online.
[0054] 2. Reconstructing 3D images into face and head models. In a specific example, after receiving biological data submitted by the data source supply and distribution software, the server creates a virtual digital human head as follows: First, a high-quality face dataset containing a large number of samples is constructed; then, the Dlib library is used to detect faces and crop out regions containing only faces. These processed images are used to train the model.
[0055] This example utilizes a custom-trained StyleGAN2 model to generate face images. The training process is based on the PyTorch framework. Before training begins, necessary Python libraries and functions are imported, including PyTorch and its submodules, data loading and processing tools, etc. Weights & Biases enables experimental tracking and visualization; the loss and other metrics during training are recorded on the W&B platform. Furthermore, image samples are periodically generated and saved to visualize the training progress. For data loading, a function called `data_sampler` is defined to handle data sampling. Since distributed training is not used, this function will default to using `data.RandomSampler` (for randomly shuffled data) or `data.SequentialSampler` (for sequential data), using the `sample_data` generator function to infinitely loop through batches of data from the data loader. Then, based on command-line arguments, the generator and discriminator models, along with their optimizers, are initialized. In this scenario, `channel_multiplier` is set to 2, increasing the number of channels in each convolutional layer of the model. During model training, a batch of real images is first loaded from the dataset. Then, the discriminator is trained while the generator's parameters are fixed. The loss of the discriminator for real and generated images is calculated, and backpropagation and iterative parameter updates are performed.
[0056] However, it's important to emphasize that in this specific example, a newly received face image is not processed directly. Instead, the HairMapper model is used to remove hair from the face image, and the rembg algorithm is used to remove the image background. See the attached manual. Figure 2 .
[0057] Hair Removal Method: First, pixel values are normalized. Next, a pre-trained pSp (PyramidScene Parsing Network) deep learning model is used to encode them into a latent space representation. This step transforms the image into a form that is easier to manipulate and analyze in a computational model. Subsequently, the StyleGAN2-ADA generator, a face parsing network, and image fusion techniques are used to remove hair while preserving facial features. The StyleGAN2-ADA generator is initialized, and a level mapper is loaded to adjust the latent codes to achieve the desired hair removal effect. Simultaneously, the face parsing network is loaded to generate a hair mask. This step indicates the regions in the image containing hair. Next, images and their corresponding latent codes are read from a specified data directory. For each image, its format (PNG or JPG) is checked, and the corresponding latent code is loaded accordingly. For each image, based on its latent code, a level mapper is applied to generate an edited latent code. Then, the face parsing network is used to generate a hair mask, further identifying which regions in the image contain hair. In the hair removal and image editing steps, setting the `--remain_ear` parameter preserves the ear portion within the generated hair mask. Then, the StyleGAN2-ADA generator is used to convert the edited latent code into the edited image. Image dilation and blurring techniques are applied to adjust the hair mask for a more natural hair removal effect. (In the post-hair removal image compositing stage, enabling the `--diffuse` parameter performs an additional diffusion step, optimizing the hair removal result through blending with the original image.)
[0058] Background Removal Method: This specific example uses the rembg algorithm to implement background removal. This algorithm can identify and separate foreground objects from the background in an image, thus automatically removing the background. First, the remove function from the rembg library is imported. Next, two variables, input_path and output_path, are defined to specify the path of the image to be processed and the path where the processed image will be stored, respectively. Then, the image file to be processed is opened in binary read mode, and its contents are read. Following this, the read image data is passed to the remove function to perform the background removal operation. The remove function relies on a pre-trained deep learning model, which can automatically identify and separate foreground objects from complex backgrounds. After background removal, the image data is returned by the remove function, and this data is written to the previously defined output file path.
[0059] After completing the above steps, before face reconstruction, the trained model is needed to map the image into a latent space to generate the image and its latent representation. This process involves finding a set of latent vectors such that when these vectors are forward-propagated through a pre-trained generative model, they can reconstruct a specific given image. First, initialization and parameter settings are performed. PyTorch, torchvision, the image processing library PIL, the progress bar library tqdm, and the perceptual loss calculation library LPIPS are imported. Command-line arguments are parsed, including the type checkpoint path, image size, learning rate, and noise parameters. Then, the pre-trained generative and discriminative models are loaded and set to evaluation mode, and the input image is loaded and transformed through the preprocessing steps described above. Subsequently, in the optimization process, for each input image, a set of latent vectors and learnable noise parameters are initialized. The optimizer Adam is set for the latent vectors and noise parameters, and the latent vectors and noise parameters are iteratively updated by calculating the loss between the generated and target images. Adding random noise to the latent vectors increases the diversity of the generated images. This method introduces slight variations that help explore different regions of the latent space and prevent the optimization process from getting trapped in local minima, thus improving the accuracy and quality of the projection. Finally, after each iteration, the noise parameters are normalized. The optimized latent vectors, generated images, and any relevant noise parameters are saved to a .pt file for subsequent 3D face reconstruction operations.
[0060] 3. The 3D face reconstruction (excluding hair and the area below the neck) uses the four-step method proposed in this example. A four-step model based on Generative Adversarial Networks (GANs) is used for 3D face reconstruction, and the black background is cropped using Unity's Mesh Cutter plugin to obtain an optimized 3D hairless face model. It is called a four-step method because this part of the 3D face reconstruction process in this example mainly includes the following four stages.
[0061] (1) Initial shape setting An ellipsoid is chosen as the prior for the object's shape. As a fundamental geometric shape, the ellipsoid provides a reasonable starting point for subsequent steps without unduly constraining the design.
[0062] (2) Shape and texture restoration stage This stage aims to recover the 3D shape (depth map) and surface texture from a given 2D image. Using a differentiable renderer, pseudo-samples are rendered under multiple viewpoints and lighting conditions based on the initial shape (ellipsoid). While these pseudo-samples are geometrically similar to real objects, they may deviate from the real image in terms of texture, detail, and lighting. The netD (depth prediction network) predicts the corresponding depth map from the input image, providing a foundation for 3D reconstruction based on the distances between objects in the scene and the observer. The netV (viewpoint estimation network) estimates the camera viewpoint from the input image. This includes the camera's position, orientation, and relative angle to the object, enabling the model to observe and reconstruct 3D objects from different angles. The netL (lighting estimation network) estimates the scene's lighting conditions, including the direction, intensity, and color of the light source, directly affecting the model's shadows, highlights, and overall lighting effects, which is crucial for generating realistic 3D reconstructed images. The netA (texture prediction network) is responsible for recovering the texture information of the object's surface from the input image, including color and texture, which helps to improve the visual realism and detail richness of the reconstructed model.
[0063] (3) Potential space projection stage This stage aims to learn the mapping from 2D images to a latent space defined by a pre-trained GAN model. The 2D image is projected into the latent space using netEnc (encoder network), and the corresponding 2D image is generated inversely by the GAN. The loss function for this process includes L1 loss, reconstruction loss, and regularization loss for the latent vector. The latent space projection stage enhances the model's expressive power within the latent space, laying the foundation for detailed recovery in subsequent steps.
[0064] (4) Comprehensive learning stage This stage combines the results of the previous two stages, refining information such as viewpoint, lighting, and texture through a comprehensive optimization process. In particular, it adjusts viewpoint and lighting estimates to more accurately simulate the rendering effect of 3D scenes. Finally, the loss function of the generator trained by this generative adversarial network is based on reconstruction loss, perceptual loss, and depth smoothing loss, while also considering the consistency between the reconstructed image and the projected image.
[0065] In this implementation, the generated 3D head model does not include the hair and the area below the neck; these parts will be customized and added by the distributed recombination and secondary training nodes. After downloading, the client combines the 3D head with the pre-processed body model bound to the skeletal structure. The effect can be seen in the attached manual. Figure 3 .
[0066] 4. Furthermore, the server constructs an independent dialogue model for each virtual digital human based on the content uploaded by the user. In this specific implementation example, the server also receives audio data uploaded by the user and trains the model using deep learning technology to generate a voice clone model with specific speaker characteristics. The server will mimic the timbre and intonation of the user-uploaded reference audio and, through natural language processing (NLP) technology, deduce the corresponding speech segments of the text to be generated. This enables the virtual digital human to engage in natural speech dialogue with others. The effects can be seen in the attached instruction manual. Figure 4 .
[0067] This specific implementation example utilizes a high-performance A800 AI computing server array (each server has four A800 GPUs, a 256GB RAM pool, and a 16TB hard drive for data storage, totaling eight servers). A server-side kernel sandbox is installed to create independent user spaces, providing a dedicated management environment for each digital human model and ensuring independence and isolation between models. Sandbox isolation can also be implemented using virtual machine technologies (such as Docker or Kubernetes). The system is equipped with a content management system to support administrators' independent uploading, management, and distribution functions. This case study leverages sandbox isolation technology and decentralized blockchain storage. The three servers are connected to form an IPFS node network, and data is split and stored across all servers, ensuring high availability and security. No single server can form a complete dataset, enhancing the system's resistance to attacks and fault tolerance.
[0068] Part 3. Distributed Reorganization and Secondary Training Nodes Devices equipped with distributed reassembly and secondary training node software (which can be packaged into the same software or app as the data source supply and distribution software) will first confirm the ownership and usage permissions of secondary users after user authentication. Based on the allocation information provided by the server, the client will download the corresponding data blocks from the server and storage bucket for subsequent reassembly on the client side, thereby achieving distributed reconstruction of the deconstructed biological data.
[0069] For example, taking the reconstruction of a virtual human body as an example, the server only generates a 3D facial and head model without hair. The generation and stitching of body parts and hair are then implemented on the client side. This adaptation is based on pre-developed, selectable hairstyles and body models. To ensure a natural fit of the body model, we have configured a complete skeletal structure for each pre-made body model and established a rich adaptation library covering various clothing and accessories, allowing different distributed terminals to make various choices during the stitching process. For example, distributed terminals can select different reconstruction schemes for the virtual human based on their respective user feedback, evolution paths, or human intervention. Through this flexible stitching processing method, the system can quickly respond to user change requests to meet customized needs in different scenarios.
[0070] This reorganization is not merely a physical reconstruction. In the distributed reorganization and secondary training nodes, the digital human undergoes reprocessing by artificial or automated models, which is multi-dimensional. This includes the creation of movement and dance (the effects can be found in the instruction manual). Figure 5 This includes not only direct improvements to clothing and makeup, but also more nuanced aspects such as thought models, dialogue models, behavioral habits, favorability ratings, and emotional feedback. For example, this implementation case includes a favorability rating system. Furthermore, when secondary users on distributed terminals engage in dialogue with the digital human, they also submit feedback information, which is used for further secondary training and development, allowing users to reorganize or replace the digital content.
[0071] The reprocessed digital humans will generate different forks, and their economic value will be calculated according to the proportion determined by the corresponding ownership blockchain, which will help optimize various economic models and improve user experience. More importantly, their evolution direction can also be autonomously optimized by relying on weighting mechanisms and distributed solutions.
[0072] For example, in this specific implementation case, distributed nodes can achieve automatic distributed evolution in the following way. In the initial test of the implementation case, after a user created a virtual human based on the above-described process of this invention, distributed nodes were installed and deployed on thousands of smartphones. These virtual humans on different distributed nodes were inspired by AI rules with different boundary conditions and random parameters, and were given different body shapes, different clothing styles, different language models, different dynamic expressions, and different dance and singing styles. These nodes were then used to broadcast live to different audiences on an internal testing platform, and different modifications were subsequently made based on the feedback from each audience. Those with better feedback (such as more likes and favorites) would receive greater weight and be executed according to certain pre-set rules. These weights serve as key parameters for algorithm input, guiding the direction and speed of the evolution of the virtual human (or digital organism) on the distributed nodes. At a certain stage in this process, a portion of the secondary branches that receive greater weight (such as the first 10 or the first 20, where different settings will have different effects on the results) will be set as new master nodes according to certain rules, thus forming an elimination mechanism. Based on these new secondary master nodes, the next level of branching and distributed evolution begins, repeating the cycle. Through this mechanism, the virtual human project will gain a wider reach and better audience feedback.
[0073] Furthermore, in this specific implementation case, the data source supply and distribution software, as well as the distributed reorganization and secondary training nodes, all support multiple terminals, including Android, iOS, Windows, Linux, and Mac systems. The iOS version can also be installed on the Apple Vision headset. In this case, the user terminal can utilize the motion tracking sensors on the MR / VR headset to perceive the user's body movements and provide real-time feedback to the virtual environment, interacting with the virtual human through gestures and movements to train the 3D motion feedback model. Similar new devices that continue to emerge in the future will further enhance the ability to collect multimodal data and obtain feedback from multimodal audiences. Attached Figure Description
[0074] Appendix Figure 1 The architecture and flowchart of this distributed system.
[0075] Appendix Figure 2 A technical flowchart for creating a separate 3D head, excluding hair and the area below the neck.
[0076] Appendix Figure 3 Image showing the assembled head section.
[0077] Appendix Figure 4Dialogue training results diagram.
[0078] Appendix Figure 5 The screenshot shows the effect of running the test after assigning actions to the action library.
Claims
1. A distributed node re-evolution system of virtual organism multi-modal data, the system has the following characteristics: a. The composition of the system is composed of several main components, including: A end: data source supply and initial issuance end, supporting the collection of users' multi-modal biological data, the setting of distributed node parameters and weight ownership allocation; B end: distributed recombination and secondary training node, supporting recombination, customization and retraining based on the initial model; S end: server end, responsible for data processing, key content generation, model training and resource management; Database: store user data, model parameters, and other system data; Weight ownership electronic ledger (can be merged with the database): record the weight ownership allocation of digital people or digital organisms; b. The specific functions and operation processes of the system are as follows: First, the data source supply and issuance end (its user is called the first-level user) initiates the creation of a project, collects and uploads multi-modal data of organisms, and disassembles the biological data of the human body, including but not limited to: Photos or videos: users can take facial images and videos through the camera, or select file upload; Audio recording or conversation: users record sounds or conversations through microphones, or select audio upload for subsequent sound synthesis or speech model training; The uploaded multi-modal data will be transmitted to the server end for creating part of the key content of the virtual digital person (or digital organism); The first-level user needs to set an initial distributed weight (or ownership issuance proportion allocation), allocate the weight or ownership of the first-level user and the second-level user, and record it in the ownership electronic ledger; For example, the weight or ownership of the virtual person can be divided into N parts, and if a user or distributed node obtains M parts of the block, the proportion is M / N; Server end functions: The key functions of the server end include recording user-generated information, and reconstructing three-dimensional images according to the received data, using algorithm models to generate three-dimensional head models or part or whole body models; The server will also receive user-uploaded audio data and generate speech cloning models with specific speaker voice characteristics; In addition, the server constructs a dialogue model with independent parameters or characteristics for each virtual digital person (or digital organism) based on user-uploaded content and learning user dialogue interaction; The functions of the distributed recombination and secondary training node (its user is called the second-level user): After the device installed with the distributed recombination and secondary training node software (which can be packaged into the same software or APP as the data source supply and issuance end software) is authenticated, the weight ownership and permissions of the corresponding second-level user or distributed node will be confirmed, the user will download the corresponding biological data or the semi-finished model generated by the server based on the relevant data, and the customization and recombination will be performed on the distributed terminal to form different digital people (or digital organisms) on different terminals of the distributed node; The content corresponding to different second-level users or nodes is not equal. Subsequently, on the distributed recombination and secondary training nodes, secondary users or nodes can recombine, re-edit the digital person (or digital organism), including but not limited to action choreography, costume and makeup matching and upgrading, expression and gesture addition, further secondary training through user dialogue with the digital person (or digital organism), etc., thereby reorganizing or replacing the digital form content on the distributed terminal; After the above process, the reprocessed digital person (or digital organism) will generate different forks; c. The system deconstructs, allocates weights and ownership, reconstructs and distributes the evolution of the organism data through the above process.
2. The system of any preceding claim, wherein, The system realizes automatic distributed evolution in the following mode through distributed nodes: after the first-level user creates the virtual person of the mother body according to the process of claim 1, these virtual persons on different distributed nodes are edited with different boundary conditions, different random parameters and AI rule inspiration, and given different secondary changes (including but not limited to different styles of costumes, different route styles, different language models, and different dynamic expressions), then these nodes are used to interact with different audiences on related platforms, and different modifications are performed according to the feedback of each audience; Wherein the feedback is better (such as getting more likes, more collections, or faster spread of forks), it will get a larger weight, and the specific algorithm is executed according to certain rules. At a certain stage, a part of the secondary forks (such as the top X, which can be given different settings to affect the results) with larger weights obtained will become new master nodes according to certain rules, and the next level of forks and distributed evolution will start from these new secondary master nodes, and the cycle will continue, thereby forming automatic optimization and elimination.
3. The system of any preceding claim, wherein, After the distributed recombination of the digital person, the secondary users of the distributed terminal can reissue the secondary changes (including but not limited to dialogue model training, action expression reassignment, and appearance form optimization) of the digital person forks in their ownership, and give the third-level users M copies of the reissue of all the secondary users' blocks. If the new third-level user obtains Q copies, the third-level user's ownership ratio of the fork is Q / M, and the entire digital person's ownership ratio is Q / N. After the fork reissue, the child node becomes a new master node, and the digital person evolved again can generate different forks, and the cycle continues, thereby realizing the continuous evolution of the organism data after being removed from the body; The system selects the evolution direction by setting corresponding rules, including the weight and ownership of different forks as parameters.
4. The system of any preceding claim, wherein, The target organism is a human body, the data source provides and issues the terminal software and the distributed reorganization and secondary training nodes are encapsulated in the application APP, the user installs the APP through the computer or the mobile device, uploads the multi-modal biological data including the face image, and the deconstruction and reconstruction of the biological data are as follows: after the server receives the data, the complete biological data is not reconstructed, but a three-dimensional head model without hair is generated based on the generated adversarial network algorithm (the generated three-dimensional head model does not contain hair and the part below the neck, which will be customized and added by the distributed client), and then the distributed client downloads the generated head model, splices and combines the three-dimensional human body with the pre-bound skeleton, the pre-adapted hairstyle and the head, and adds the pre-adapted three-dimensional clothing from the local client or the server.
5. The system of any preceding claim, wherein, The user on the distributed client trains the dialogue language model prototype of the digital human through dialogue communication and manual uploading of related document data, the server generates a preliminary dialogue language model prototype on the server side by using a deep learning algorithm, and gives the digital human model a voice timbre according to the reference audio uploaded by the user, the overall architecture of the voice timbre model is composed of a posteriori encoder, a priori encoder, a decoder, a discriminator and a random duration predictor, and then the audio output of the dialogue model is generated through text-to-speech technology, and the text given by the dialogue model to reply to the user is converted into audio after processing, so as to realize natural voice interaction; In the optimization process, an adversarial training framework is constructed to optimize the conditional variational autoencoder to improve the quality and authenticity of the generated voice; in the training process, a joint loss function is used, including reconstruction loss, KL divergence loss and adversarial loss, wherein the reconstruction loss and the KL divergence loss are used to optimize the conditional variational autoencoder, and the adversarial loss is used to optimize the discriminator network.
6. The system of any preceding claim, wherein, The system records the ownership issuance ratio set by the primary user as the weight or interest of different distributed users, and stores it on a special property blockchain to realize the inheritance and transaction of the virtual human property. The right confirmation model adopts the following QA model or QB model: the QA model combines symbolic mapping coding and blockchain technology. By combining blockchain with symbolic mapping coding, the model allows all data sharing participants to generate a unique symbolic mapping table and embed their digital identity coding into the byte sequence of shared data, so that the right declaration is no longer linked to the data content and type, and all participants can jointly supervise the delivery and access of shared data, thereby realizing the public verification of data rights. In contrast, the QB model is based on secret sharing and symbolic mapping coding. To address the pressure on storing symbolic mapping tables and the problem of user privacy leakage when dealing with large-scale data, the model uses fuzzy C-means clustering algorithm to cluster the data, and generates an asymmetric fingerprint based on the symbolic mapping table to realize data right confirmation. The model combines secret sharing technology to complete data distribution and acquisition, uses a random verification algorithm to provide identity proof for users, and records each link of the sharing process in the blockchain to ensure the traceability and denial of data.
7. The system of any preceding claim, wherein, After receiving the biological data submitted by the data source supply and issuance end software, the server end creates a virtual digital human head in the following steps: first, a high-quality face dataset containing a large number of samples is constructed; then, the Dlib library is used to detect the face and crop the area containing only the face. These processed images are used to train the StyleGan2 model; for a new face test image, the HairMapper model is used to perform hair removal operation on the face image, and the rembg algorithm is used to remove the image background; Subsequently, the trained model is used to map the image to the latent space to generate the image and its latent representation; the three-dimensional face reconstruction model based on the generative adversarial network is used to perform three-dimensional face reconstruction, and the Mesh Cutter plug-in of Unity is used to perform black background cutting of the model, and finally the optimized three-dimensional hairless face model is obtained.
8. The system of any preceding claim, wherein, The system also gives the virtual human expression capture and animation generation, which real-time assigns the expression changes of the real user to the virtual human, as follows:
1. Design a video face feature point tracking algorithm. First, use a camera to capture images of facial expressions, then use a multi-task learning cascade convolutional neural network algorithm and a cascade gradient boosting regression tree algorithm to detect faces and extract feature points; 2. Set the face expression animation control parameters, solve the head posture, and generate the corresponding head rotation matrix and translation vector. Use support vector machines to build an AU detector and an intensity value regressor, establish a mapping relationship between AU and expression bases, and obtain the required face expression parameters; 3. Design a virtual face expression driver. Transmit the face expression control parameters to the Unity software through a plug-in to drive the animation model to generate the corresponding face expression animation.
9. The system of any preceding claim, wherein, The target organism is a human body, and the data source, the distribution end software, and the distributed reconstruction and secondary training node are all encapsulated in a mobile MR / VR helmet, such as an Apple Vision APP. The user uploads the depth photos and audio captured by the helmet camera through the MR / VR mobile helmet APP, and then the server generates a three-dimensional head model without hair based on a generative adversarial network algorithm (the generated three-dimensional head model does not include hair and the part below the neck, which will be customized and added by the distributed client), and then the client downloads the head model, combines the three-dimensional human body with the head model, and adds different types of hairstyles according to the user's preferences, as well as clothes that fit the body type from the client or server. Then, the terminal uses the motion tracking sensor equipped with the MR / VR virtual reality helmet to perceive the user's body movements and feed them back to the virtual environment in real time, allowing the virtual person to interact with the user through gestures and postures for training.
10. The system of any preceding claim, wherein, The biological organism is an animal, and the system disperses the entity elements of the organism through the following deconstruction and reconstruction process, and then performs distributed reconstruction: first, the multi-modal information of the organism is deconstructed and disassembled into face information and other information (which can include but is not limited to torso shape information, hair information, sound line information, personality habits, and response characteristics), and then the weight ownership segmentation is performed, and then the face feature domain conversion network model is established by combining the autoencoder network and the conversion network in the three-dimensional reconstruction process, so as to convert the animal face into a virtual face, thereby facilitating the direct application of face recognition and reconstruction model, and finally the three-dimensional animal face is reconstructed again according to the inverse conversion network model.