Multi-view figure image reconstruction method and device based on twin diffusion model
Through the multi-view character image reconstruction method based on the twin diffusion model, combined with the character analytical mask and CLIP guidance, the problem of fine-grained clothing pattern information retention and pose gap in the existing technology is solved, and high consistency and high-quality multi-view character image generation is achieved.
Patent Information
- Application Number
- CN202510594067.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-05-09
AI Technical Summary
In the prior art, when generating multi-view characters images, it is difficult to effectively retain fine-grained and complex clothes pattern information, and it is impossible to reasonably generate character images in scenes with large posture gaps.
A multi-perspective character image reconstruction method based on the twin diffusion model is adopted, and a fine-grained feature fusion and image generation are achieved through the twin diffusion model framework combined with the attention fusion mechanism guided by character analytical mask and the local alignment mechanism guided by CLIP.
It significantly improves the consistency and picture quality of the character appearance from multiple perspectives, and overcomes the shortcomings of traditional methods in scenarios with large gaps in complex patterns and postures.
Smart Images

Figure CN120125473A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of image data generation and artificial intelligence, and particularly relates to a multi-view human image reconstruction method and device based on a twin diffusion model. Background Art
[0002] In recent years, the technology of AI-generated content (AIGC) has developed vigorously and has extensive social value in fields such as e-commerce, short video special effects, and fashion. The task of generating multi-view human images from a single RGB image guided by pose mainly uses a generative model to process a single RGB human image and target pose features, such as 2D human key points, to generate a human image that conforms to the target pose. The high appearance consistency of human images in different poses is the most important evaluation index for this task, which determines the upper limit of the application value of the task.
[0003] In order to achieve high-quality image generation, most of the existing technologies adopt the currently popular method based on the diffusion model. For example, CFLD (Coarse-to-fine latent diffusion for pose-guided person image synthesis) uses a hybrid strength attention and a perceptual decoder for coarse-to-fine image generation, and PCDM (Advancing pose-guided image synthesis with progressive conditional diffusion models) uses a multi-stage diffusion model to gradually fuse and repair features to generate high-fidelity images. However, the above two types of methods are lacking in the quality of generating fine-grained complex clothing pattern information and cannot effectively restore the clothing pattern of the input single RGB human image in a 1:1 manner. In addition, in the scenario where there is a large gap between the pose of the input single RGB human image and the target pose, a reasonable human image cannot be generated. Therefore, how to design a generation method that can maximize the retention of high consistency of human multi-views remains a problem to be solved. Summary of the Invention
[0004] The purpose of the present invention is to provide a multi-view human image reconstruction method and device based on a twin diffusion model in view of the deficiencies of the prior art, which can greatly improve the appearance consistency and image quality of a human in multiple views. Compared with the previous methods, this method can effectively solve the problems of unreasonable generation and complex pattern generation when there is a large gap between the pose of the input RGB human image and the target pose.
[0005] To achieve the above purpose, the present invention adopts the following technical solutions: A multi-view human image reconstruction method based on a twin diffusion model, the method includes two parts: model training and model inference:
[0006] (1) Model training
[0007] Step 1.1: Obtain a two-dimensional RGB reference human image, a two-dimensional RGB target human image, and a two-dimensional RGB target pose image;
[0008] Step 1.2: Obtain an overall human semantic segmentation mask map and a local area segmentation map based on the two-dimensional RGB reference human image and the two-dimensional RGB target pose image;
[0009] Step 1.3: Input the two-dimensional RGB reference human image and the two-dimensional RGB target human image into the encoder to obtain a latent reference human image feature and a target human image feature, and add noise to the target human image feature to obtain a latent variable; input the two-dimensional RGB target pose image into the pose multi-layer convolutional network to obtain a pose feature;
[0010] Step 1.4: Based on the twin diffusion model, fuse the results of Step 2 and Step 3, and remove the noise through the reverse diffusion process to restore and generate the target human image feature;
[0011] (2) Model inference
[0012] Step 2.1: Obtain the two-dimensional RGB reference human image and the two-dimensional RGB target pose image that need to be reconstructed, obtain the overall human semantic segmentation mask map and the local area segmentation map based on the model training process, and obtain the latent reference human image feature and the pose feature based on the trained encoder and the pose multi-layer convolutional network; and add random noise to the latent reference human image feature;
[0013] Step 2.2: Based on the trained twin diffusion model, perform feature fusion, and remove the noise through the reverse diffusion process to restore and generate a complete latent image feature, and input it into the decoder to obtain the final target human image.
[0014] Furthermore, in Step 1.1, the two-dimensional RGB reference human image should have clear clothing textures and contours and complete portrait features, and the human figure should account for more than 70% of the entire picture. The two-dimensional RGB reference human images and two-dimensional RGB target human images of the same person in different poses or perspectives are used as a data pair for training.
[0015] Furthermore, in Step 1.2, extract the human parsing information of the two-dimensional RGB reference human image and the RGB image of the key points of the reference human body pose; combine the human parsing information, the RGB image of the key points of the human body pose, and the two-dimensional RGB target pose image to obtain the overall human semantic segmentation mask map and the local area segmentation map of the overlapping area.
[0016] Further, perform estimation and inference on the two-dimensional RGB reference human image to obtain the masks of each body part to obtain human parsing information. Perform inference on the two-dimensional RGB reference human image through a pose estimation model to obtain the RGB image of the key points of the reference human body pose; based on key point matching, obtain the key points of the overlapping area between the RGB image of the key points of the reference human body pose and the two-dimensional RGB target pose image, and then screen the human parsing information to obtain the local masks of the overlapping areas, combine them into a mask image, and obtain the overall human semantic segmentation map; use the local masks to segment the two-dimensional RGB reference human image to obtain the segmentation maps of each local area.
[0017] Further, in step 1.4, the twin diffusion model uses two UNet networks, namely a feature extraction network and a generation network, and realizes fine-grained feature fusion in two stages based on the attention fusion mechanism guided by the human parsing mask and the local alignment mechanism guided by CLIP.
[0018] Further, in the first stage, the attention fusion mechanism guided by the human parsing mask uses an interpolation method to reshape the overall human semantic segmentation map into the shape of the feature variables of each layer in the generation network. In the second stage, use CLIP to extract the features of the overlapping local areas and fuse them with the latent variables to achieve feature alignment enhancement of the local areas.
[0019] Further, in step 2.1, the random noise is random Gaussian noise, which is obtained through random initialization, and its shape and size are consistent with the features of the latent reference human image.
[0020] In a second aspect, the present invention also provides a multi-view human image reconstruction device based on a twin diffusion model, including a memory and one or more processors. An executable code is stored in the memory. When the processor executes the executable code, it implements the multi-view human image reconstruction method based on a twin diffusion model described above.
[0021] In a third aspect, the present invention also provides a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, it implements the multi-view human image reconstruction method based on a twin diffusion model described above.
[0022] In a fourth aspect, the present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the multi-view human image reconstruction method based on a twin diffusion model described above.
[0023] The advantages of the present invention compared with the prior art are as follows: The present invention adopts a twin diffusion model framework to overcome the inherent drawback of the traditional method of having weak ability to retain fine-grained information, effectively improving the appearance consistency of the person in clothes with complex patterns, overcoming the problem that the traditional method cannot generate reasonably when the pose of the reference person and the target pose change greatly, and significantly improving the quality of the generated pictures. Brief Description of the Drawings
[0024] Figure 1 It is a flowchart of the implementation of the method of the present invention.
[0025] Figure 2 It is a general block diagram of the method framework of the present invention.
[0026] Figure 3 It is a schematic diagram of the segmentation and estimation of the overlapping area between the pose of the reference person image and the pose of the target person image provided by the present invention.
[0027] Figure 4 It is a schematic diagram of the mask-guided attention fusion mechanism proposed by the present invention.
[0028] Figure 5 It is a schematic diagram of the comparison of the application effect of the method provided by the example of the present invention and the effect of the current method.
[0029] Figure 6 It is a schematic diagram of the structure of a multi-view human image reconstruction device based on a twin diffusion model provided by the present invention. Detailed Description of the Embodiment
[0030] The present invention will be further described in detail below with reference to the drawings and specific embodiments. First, several concepts in this article are introduced here.
[0031] The variational autoencoder belongs to a generative model and includes two main parts: an encoder and a decoder . Among them, the encoder can encode the input high-dimensional variable into the latent space to obtain a low-dimensional latent variable, enabling the neural network to learn the features with the maximum information density and reducing the computational consumption. The decoder then maps the processed low-dimensional latent variable back to the original high-dimensional space.
[0032] The diffusion model belongs to a generative model, and its training process includes a forward noise addition process and a backward denoising process; among them, the forward noise addition process is to gradually add random noise to the input variable according to the time step t to gradually turn it into random Gaussian noise. The backward denoising process is to train a denoising model , input random Gaussian noise to predict the noise added at each step and then gradually remove it to realize generating the target data from the noise.
[0033] A multi-view human image reconstruction method based on the twin diffusion model provided by the present invention, as Figure 1 shown, includes the following steps:
[0034] Model training part:
[0035] Step 1: Obtain the two-dimensional RGB reference human image , and the two-dimensional RGB target human image ;
[0036] The reference human image can be sourced from online shopping websites or real photos, and can be from various perspectives such as the front, side, and back of the person. To achieve better generation quality, the person should occupy more than 70% of the entire picture, and at the same time, the image should have clear clothing textures and contours as well as complete portrait features. The present invention is trained using the In-Shop Clothes Retrieval dataset in the open-source multi-view human dataset DeepFashion, and two images of the same person in different poses or perspectives (reference image - target image) are used as a data pair for training.
[0037] Step 2: Extract the human parsing information of the above reference human image , the RGB image of the key points of the reference human body posture , and use , the two-dimensional RGB target pose image and to obtain the overall human semantic segmentation mask image of the overlapping area and the local area segmentation map ;
[0038] The human parsing information of the reference human image is estimated and inferred through the open-source parsing segmentation model for to obtain the masks of each body part, and are inferred through the open-source openpose pose estimation model for and . The present invention classifies the obtained masks into four regions: head, torso, legs, and hands , that is, . Then the present invention uses and to perform Figure 3 shown key point matching to obtain the key points of the overlapping region between the two, and then is screened to obtain the local masks of each overlapping region , which are combined into a mask image, that is, the overall human semantic segmentation map , and use the local mask for Perform segmentation to obtain pictures of each local area .
[0039] Step 3: Input the above and into the encoder of the pre-trained Variational Autoencoder to obtain the latent reference person image features and the target person image features . Input the target pose image into the pose multi-layer convolutional network to obtain the pose features . Gradually add random Gaussian noise to the target person image features according to the preset time step t to obtain the noisy latent variable ;
[0040] Combine the appendix Figure 2 , and use the encoder to encode into the latent space to obtain , which can be expressed as . The pose multi-layer convolutional network consists of four layers of networks. The input channels of each layer of network are (16, 32, 96, 256) respectively. Each layer contains two standard convolutional layers, with the stride being 1 and 2 respectively, and the kernel size being 3×3. Gradually add random Gaussian noise to according to the preset time step t to obtain the noisy latent variable .
[0041] Step 4: Input the said , , , and into the twin diffusion model framework for feature fusion, and gradually remove the noise through the reverse diffusion process to restore and generate the complete latent image features ;
[0042] Combine the appendix Figure 2 , the twin diffusion model proposed by the present invention uses two UNet networks, namely the feature extraction network and the generation network; among them, the feature extraction network is responsible for extracting the fine-grained features in. The weights in the initial two UNet networks are both loaded with the UNet weights in the pre-trained Stable Diffusion v1.5. The CLIP model loads the pre-trained model to extract the features of the local area. Only the parameters in the feature extraction network, the generation network, and the pose encoder participate in the training in the whole framework, and the remaining model parameters are fixed. The training objective function of the model is:
[0043]
[0044] in For each step Random noise added through training To predict the noise added at each step to achieve denoising. Figure 4 ,This paper proposes a two-stage strategy to achieve fine-grained feature fusion,including the attention fusion mechanism guided by the character parsing mask and the local alignment mechanism guided by CLIP. The attention fusion mechanism guided by the character parsing mask first uses the interpolation method to transform the overall human semantic segmentation map Reshaped into the shape of the feature variables of each layer in the generative network, the calculation formula for one stage in the generative network is:
[0045]
[0046] in To generate the query, key and value matrices in the attention modules of each layer of the network, d is the dimension of the key matrix. This step allows the model to focus on the feature extraction of the overlapping area, thereby avoiding the influence of irrelevant areas on the generation quality, and can effectively improve the reconstruction effect of complex pattern clothes. The calculation formula for the second stage in the generation network is:
[0047]
[0048] In the second stage, CLIP is used to extract the features of the overlapping local areas and fuse them with the latent variables to achieve feature alignment enhancement of the local areas, which can effectively retain the fine-grained features of the local clothes and improve the generation quality. Figure 4 The present invention visualizes the attention distribution map, and it can be found that the fusion mechanism proposed in the present invention can make the generation process pay more attention to the features of the overlapping areas and thus better retain the fine-grained information.
[0049] Based on the above loss function, the trainable parameters are back-propagated and calculated, and the optimizer and stochastic gradient descent method are used to optimize and update the trainable parameters until the model converges. After the model training is completed, the model can be used for inference to realize image synthesis.
[0050] Model reasoning part:
[0051] Step 1: Get the 2D RGB reference person image input by the user and a 2D RGB target pose image of the target human pose key points ;
[0052] The reference human image can be sourced from online shopping websites or real-life photos, and can be from various perspectives such as the front, side, and back of the person. To achieve better generation quality, the person should occupy more than 70% of the entire image area. Meanwhile, the image should have clear clothing textures and contours, as well as complete human portrait features. It should be consistent with the openpose data format, such as the number of key points.
[0053] Step 2: Extract the human parsing information , refer to the two-dimensional RGB target pose image of the key points of the reference human body pose , and utilize , and or directly input by the user to obtain the overall human semantic segmentation mask map of the overlapping area and the local area segmentation map ;
[0054] There are two ways to obtain and . The first way is to estimate and infer the masks of each body part through the open-source parsing segmentation model for , and infer through the open-source openpose pose estimation model for . Then, utilize and to perform key point matching to obtain the key points of the overlapping area between the two, and then screen to obtain the local masks of each overlapping area , combine them into a single mask image, which is the overall human semantic segmentation map . Use the local mask to segment to obtain the local area images . The second way is for the user to directly input the overall human semantic segmentation map and .
[0055] Step 3: Input the above into the encoder of the pre-trained Variational Autoencoder to obtain the latent reference human image features , and input the target pose image into the pose multi-layer convolutional network to obtain the pose features . Randomly initialize to obtain random Gaussian noise , whose shape and size are consistent with ;
[0056] Combine Appendix Figure 2 , and adopt the encoder Encode into the latent space to obtain , which can be expressed as . The pose multi-layer convolutional network consists of four layers of networks. The input channels of each layer of network are (16, 32, 96, 256) respectively. Each layer contains two standard convolutional layers, where the stride is 1 and 2 respectively, and the convolutional kernel size is 3×3. Randomly initialize to obtain Gaussian noise with the same shape as . .
[0057] Step 4: Input the , , , and into the twin diffusion model framework for feature fusion, and gradually remove the noise through the reverse diffusion process to generate the target latent image feature ;
[0058] Combine with Appendix Figure 2 . In the inference process, the trained twin diffusion model is used, and the fine-grained appearance features are extracted by the feature extraction network. The noise at each time step is gradually predicted through the proposed two-stage fusion mechanism, and then the noise in is gradually removed to obtain the target latent image feature . The present invention adopts the Classifier-Free (Classifier-free diffusion guidance) inference strategy to achieve controllable generation, and the formula is as follows .
[0059]
[0060] where represents a vector with the same shape as the input but with a value of 0. Through this strategy, the diversity and accuracy of controllable generation can be improved, is the control condition coefficient (usually set to 3.5).
[0061] Step 5: Input the into to map it into the high-dimensional picture space to obtain the final target person image.
[0062] Combine with Appendix Figure 5 shows the comparison chart of the generation effect of the method of the present invention and the previous method. The previous method is shown in Table 1.
[0063] Table 1
[0064]
[0065] Compared with the prior methods, the present invention adopts a twin diffusion model framework to overcome the inherent drawback of the prior methods of having weak ability to retain fine-grained information, effectively improving the appearance consistency of the person with clothes having complex patterns, overcoming the problem that the prior methods cannot generate reasonably when the pose of the reference person varies greatly from that of the target pose, and significantly enhancing the quality of the generated pictures.
[0066] Corresponding to the embodiment of the multi-view human image reconstruction method based on the twin diffusion model described above, the present invention also provides an embodiment of a multi-view human image reconstruction device based on the twin diffusion model.
[0067] See Figure 6 , an embodiment of a multi-view human image reconstruction device based on the twin diffusion model provided by the embodiment of the present invention includes a memory and one or more processors. Executable code is stored in the memory. When the processor executes the executable code, it is used to implement the multi-view human image reconstruction method in the above embodiment.
[0068] The embodiment of the multi-view human image reconstruction device based on the twin diffusion model provided by the present invention can be applied to any device with data processing capabilities. The any device with data processing capabilities can be a device or apparatus such as a computer. The device embodiment can be implemented by software, or by hardware or a combination of software and hardware. Taking software implementation as an example, as a logically meaningful device, it is formed by the processor of any device with data processing capabilities where it is located reading the corresponding computer program instructions in the non-volatile memory into the memory for running. From the hardware level, as Figure 6 shown, it is a hardware structure diagram of any device with data processing capabilities where the multi-view human image reconstruction device based on the twin diffusion model provided by the present invention is located. Except for Figure 6 the processor, memory, network interface, and non-volatile memory shown, generally according to the actual functions of the any device with data processing capabilities where the device in the embodiment is located, other hardware may also be included, which will not be elaborated here.
[0069] The implementation processes of the functions and roles of each unit in the above device are specifically detailed in the implementation processes of the corresponding steps in the above method, which will not be elaborated here.
[0070] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to the descriptions in the method embodiments. The device embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of the present invention. A person of ordinary skill in the art can understand and implement it without creative work.
[0071] An embodiment of the present invention also provides a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, a method for multi-view human image reconstruction based on a twin diffusion model in the above embodiment is implemented.
[0072] The computer-readable storage medium may be an internal storage unit of any device with data processing capabilities described in any of the foregoing embodiments, such as a hard disk or memory. The computer-readable storage medium may also be an external storage device of any device with data processing capabilities, such as a plug-in hard disk, a smart media card (SMC), an SD card, a flash card, etc. equipped on the device. Further, the computer-readable storage medium may also include both an internal storage unit and an external storage device of any device with data processing capabilities. The computer-readable storage medium is used to store the computer program and other programs and data required by any device with data processing capabilities, and can also be used to temporarily store data that has been output or will be output.
[0073] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the method for multi-view human image reconstruction based on a twin diffusion model is implemented.
[0074] The above embodiments are used to explain the present invention, rather than limit the present invention. Any modifications and changes made within the spirit and scope of the claims of the present invention fall within the protection scope of the present invention.
Claims
1. A multi-view character image reconstruction method based on a twin diffusion model, characterized in that: This method includes two parts: model training and model inference: (1) Model training Step 1.1: Obtain a 2D RGB reference person image, a 2D RGB target person image, and a 2D RGB target posture image; Step 1.2: Obtain an overall human semantic segmentation mask map and a local region segmentation map based on a two-dimensional RGB reference person image and a two-dimensional RGB target posture image; Step 1.3: Input the two-dimensional RGB reference person image and the two-dimensional RGB target person image into the encoder to obtain the latent reference person image features and the target person image features, and add noise to the target person image features to obtain the latent variables; input the two-dimensional RGB target posture image into the posture multi-layer convolutional network to obtain the posture features; Step 1.4: Based on the twin diffusion model, the results of step 2 and step 3 are fused, and the noise is removed through the reverse diffusion process to restore the target person image features; (2) Model Reasoning Step 2.1: Obtain the 2D RGB reference person image and the 2D RGB target pose image to be reconstructed, obtain the overall human semantic segmentation mask map and the local area segmentation map based on the model training process, obtain the potential reference person image features and pose features based on the trained encoder and pose multi-layer convolutional network; and add random noise to the potential reference person image features; Step 2.2: Perform feature fusion based on the trained twin diffusion model, remove noise through the reverse diffusion process, restore and generate complete latent image features, and input them into the decoder to obtain the final target person image.
2. The multi-view character image reconstruction method based on the twin diffusion model according to claim 1, characterized in that: In step 1.1, the 2D RGB reference person image should have clear clothing texture and contours as well as complete portrait features. The person should account for more than 70% of the entire image. The 2D RGB reference person image and the 2D RGB target person image of the same person in different postures or perspectives are used as a data pair for training.
3. The multi-view character image reconstruction method based on the twin diffusion model according to claim 1, characterized in that: In step 1.2, the human body parsing information of the two-dimensional RGB reference person image and the RGB image of the human body posture key points of the reference person are extracted; the human body parsing information, the RGB image of the human body posture key points and the two-dimensional RGB target posture image are combined to obtain the overall human body semantic segmentation mask map and the local area segmentation map of the overlapping area.
4. The multi-view character image reconstruction method based on the twin diffusion model according to claim 3 is characterized in that: The two-dimensional RGB reference person image is estimated and inferred to obtain the masks of each body part to obtain the human body analysis information. The two-dimensional RGB reference person image is inferred through the posture estimation model to obtain the RGB image of the reference person's human body posture key points; Based on key point matching, the key points of the overlapping area of the reference character's human body posture key point RGB image and the two-dimensional RGB target posture image are obtained, and then the human body analysis information is screened to obtain the local masks of the overlapping areas, which are combined into a mask image to obtain the overall human body semantic segmentation map; the two-dimensional RGB reference character image is segmented using the local mask to obtain the segmentation maps of each local area.
5. The multi-view character image reconstruction method based on the twin diffusion model according to claim 1, characterized in that: In step 1.4, the twin diffusion model uses two UNet networks, namely the feature extraction network and the generation network, to achieve fine-grained feature fusion in two stages based on the attention fusion mechanism guided by the character parsing mask and the local alignment mechanism guided by CLIP.
6. The multi-view character image reconstruction method based on the twin diffusion model according to claim 5, characterized in that: In the first stage, the attention fusion mechanism guided by the character parsing mask uses the interpolation method to reshape the overall human semantic segmentation map into the shape of the feature variables of each layer in the generative network. In the second stage, CLIP is used to extract the features of the overlapping local areas and fuse them with the latent variables to achieve feature alignment enhancement of the local areas.
7. The multi-view character image reconstruction method based on the twin diffusion model according to claim 1, characterized in that: In step 2.1, the random noise is random Gaussian noise, which is obtained by random initialization, and its shape and size are consistent with the characteristics of the potential reference person image.
8. A multi-view character image reconstruction device based on a twin diffusion model, comprising a memory and one or more processors, wherein the memory stores executable code, characterized in that: When the processor executes the executable code, a multi-view character image reconstruction method based on a twin diffusion model as described in any one of claims 1 to 7 is implemented.
9. A computer-readable storage medium having a program stored thereon, characterized in that: When the program is executed by a processor, a multi-view character image reconstruction method based on a twin diffusion model as described in any one of claims 1 to 7 is implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, a multi-view character image reconstruction method based on a twin diffusion model as described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Virtual fitting method based on diffusion model
CN117011207A
Semantic perception multi-view three-dimensional human body reconstruction method and device and medium
CN117292041A
Virtual model clothing display image intelligent generation method and device based on diffusion model
CN119131212A
Garment multi-view generation method and device
CN119600128A
Workpiece surface morphology generation method and apparatus based on multimodal image generation
WO2025060317A1
Cited By
Digital human image generation method and system based on 3D digital twinning dynamic interaction
CN121458845A
A Method and System for Generating Digital Human Images Based on Dynamic Interaction of 3D Digital Twins
CN121458845B