Millisecond binocular human body reconstruction method and device based on basic reconstruction model

Through the basic reconstruction model method, using the human point cloud prediction model and the 3D Gaussian parameter regression model, the rapid reconstruction of the three-dimensional human model can be achieved with only two images, solving the problems of poor reconstruction effect and difficult data acquisition in the existing technology.

CN120563718APending Publication Date: 2025-08-29HUAZHONG UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510558941.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-08-29

AI Technical Summary

Technical Problem

In the prior art, it is difficult to achieve good human reconstruction effect when only inputting binocular images, and the generalization ability is lacking, resulting in high demand for the number of input images and difficulty in data acquisition.

Method used

Using a basic reconstruction model method, the binocular images are processed by defining the human point cloud prediction model and the side image enhancement module, point clouds of four perspectives are generated, and the three-dimensional mannequin model is reconstructed using the 3D Gaussian parameter regression model. The reconstruction of the three-dimensional mannequin can be achieved by only two images.

Benefits of technology

It realizes the rapid and accurate reconstruction of the three-dimensional mannequin with only two images input, reducing the need for the number of input images and improving the efficiency and consistency of reconstruction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120563718A_ABST
    Figure CN120563718A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computer vision, and provides a millisecond binocular human body reconstruction method and device based on a basic reconstruction model. The method comprises the following steps: defining a human body point cloud prediction model, and processing a human body front image and a human body back image by using the human body point cloud prediction model to obtain a human body front point cloud, a human body back point cloud, a human body left side point cloud and a human body right side point cloud; performing color assignment on the human body left side point cloud and the human body right side point cloud by using a side image enhancement module to obtain a human body left side image and a human body right side image; and inputting the human body front image, the human body back image, the human body front point cloud, the human body back point cloud, the human body left side image, the human body right side image, the human body left side point cloud and the human body right side point cloud into the 3D Gaussian parameter regression model, and outputting to obtain a three-dimensional human body model. According to the method, the reconstruction of the three-dimensional human body model can be realized only by two images (namely the front view and the back view).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to a millisecond-level binocular human body reconstruction method and device based on a basic reconstruction model. Background Art

[0002] With the rapid development of neural network science and the significant increase in hardware computing power, the importance of 3D representation technology is growing. In the 3D field, human body reconstruction has always been a key topic and can be seen as a crucial bridge between the real and digital worlds. 3D human body reconstruction focuses on converting 2D images or scanned data of the human body into an accurate 3D digital model. This technology can extract the body's geometric shape and surface details from photographs of the body taken from different angles or from point cloud data acquired through methods such as lasers and structured light. This goes beyond simply replicating the human form and includes highly accurate simulation of subtle changes in human movement and expression, offering broad practical applications, including virtual / augmented reality, gaming, and the metaverse. Previous reconstruction methods typically required specialized and expensive equipment to capture human data, such as using synchronized cameras to capture the target person from multiple perspectives. Lowering the capture requirements allows more people to complete their own reconstructions and makes them applicable to a wider range of scenarios, which has significant practical significance.

[0003] However, existing human reconstruction methods often accept monocular or multi-view videos as input. When the number of input images is small, such as when inputting binocular images (i.e., images from both the front and back perspectives), good reconstruction results are often difficult to achieve. Furthermore, most existing human reconstruction methods lack generalization capabilities, requiring separate training for each scene. This leads to a high demand for input images and greatly increases the difficulty of data collection.

[0004] In view of this, overcoming the defects of the existing technology, reducing the demand for the number of input images, and reducing the cost of reconstruction are issues that need to be urgently addressed in this technical field. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide a millisecond-level binocular human body reconstruction method and device based on a basic reconstruction model, so as to solve the problem in the prior art that it is difficult to obtain a good reconstruction effect when only binocular images are input.

[0006] The present invention adopts the following technical solutions:

[0007] In a first aspect, the present invention provides a millisecond-level binocular human body reconstruction method based on a basic reconstruction model, comprising:

[0008] Define a human point cloud prediction model based on the basic reconstruction model, and use the human point cloud prediction model to reconstruct the human front image If and the back image of the human body I b Processing is performed to obtain the human body front point cloud P f,f , point cloud P of the back of the human body b,f , point cloud P of the left side of the human body l,f And the right side point cloud P r,f ;

[0009] Define a side image enhancement module, and use the side image enhancement module to analyze the left side point cloud P of the human body. l,f And the right side point cloud P r,f Perform color assignment to obtain the left side image of the human body I l and the right side image of the human body I r ;

[0010] Define a 3D Gaussian parameter regression model and transform the human front image I f , human back image I b , human front point cloud P f,f , point cloud P of the back of the human body b,f , human left side image I l , human right side image I r , point cloud P of the left side of the human body l,f And the right side point cloud P r,f The input is to the 3D Gaussian parameter regression model, and the output is a three-dimensional human body model.

[0011] Preferably, the human point cloud prediction model includes a binocular image encoder, a decoder, a frontal prediction head H f and back prediction head H b ;

[0012] The binocular image encoder is configured to obtain the human front image I f Extract the frontal features of the human body G f , and from the back image of the human body I b Extract the back feature G of the human body b ;

[0013] The decoder converts the human body front feature G f After mapping, it is input to the front prediction head H f , to obtain the human body front point cloud The back feature G of the human body b After mapping, it is input to the back prediction head H b , to obtain the point cloud of the back of the human body in, is a positive feature sequence, is the back feature sequence, and B is the feature sequence length.

[0014] Preferably, the human body point cloud prediction model further includes a side processing module, a left side prediction module H l and right face prediction head H r ;

[0015] The side processing module processes the front feature G of the human body f and the back feature G b Perform average processing to obtain the side features of the human body And the human body side features Input to the left face prediction head H l , get the point cloud of the left side of the human body And the human body side features Input to the right face prediction head H r , get the point cloud of the right side of the human body

[0016]

[0017] Preferably, the side image enhancement module is the point cloud P of the left side of the human body l,f The left side face point in the image is matched with the nearest positive and negative points, and the colors of the matched positive and negative points are assigned to the corresponding left side face points to obtain a colored left side face point cloud set. The colored left side face point cloud set is mapped to the pixel plane to obtain the left side face image of the human body I l ;

[0018] The side image enhancement module also generates the right side point cloud P r,f The right side face point in the image is matched with the nearest positive and negative points, and the colors of the matched positive and negative points are assigned to the corresponding right side face points to obtain a colored right side face point cloud set. The colored right side face point cloud set is mapped to the pixel plane to obtain the right side face image of the human body I r ;

[0019] Among them, the color of the left side point is n is the number of point clouds;

[0020] For the point cloud P on the left side of the human body l,f Every left side point i∈{1,2,…n}, using the nearest neighbor search function F nns Search In the positive and negative point set {P f,f ,P b,f}, the corresponding label j, that is, Use label j to select the front and back color set {C f,f ,C b,f}Index to get the color

[0021] Preferably, the 3D Gaussian parameter regression model includes a positive and negative view parameter regression module Side view parameter regression module and output modules;

[0022] The human front image I f , human back image I b , human front point cloud P f,f And the back point cloud P b,f Input to the front and back view parameter regression module To output front Gaussian points and back Gaussian points

[0023] The human body left side image I l , human right side image I r , point cloud P of the left side of the human body l,f And the right side point cloud P r,f Input to the side view parameter regression module To output the left side Gaussian points and the right side Gaussian points

[0024] The front Gaussian point θ f 、Back Gaussian point θ b , Gaussian point θ on the left side l and the right side Gaussian point θ r Input to output module to output 3D human body model Represents a merge operation.

[0025] Preferably, the Gaussian point includes a center position μ, a color c, an opacity o, and a covariance matrix Σ, that is, the Gaussian point θ=(μ, c, o, Σ).

[0026] Preferably, the human body point cloud prediction model is pre-trained, and the loss function used in the training is Among them, L reg Represents the process of calculating the Euclidean distance, P h is the network prediction output, P gt is the training label; the Euclidean distance represents the straight-line distance between the point pairs. For two n-dimensional vectors x=(x1,x2,…,x n ) and y=(y1,y2,…,y n ), the Euclidean distance between them is

[0027] Preferably, the 3D Gaussian parameter regression model is pre-trained, and the loss function used in the training is Among them, L rgb Represents the color difference between pixels. represents the calculated SSIM loss, β is the preset coefficient, Irender is the network output 3D human body model θ rendered image, I gt is the corresponding training label.

[0028] In a second aspect, the present invention further provides a millisecond-level binocular human body reconstruction device based on a basic reconstruction model, which is used to implement the millisecond-level binocular human body reconstruction method based on the basic reconstruction model described in the first aspect, and the device includes:

[0029] At least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the processor to perform the millisecond-level binocular human body reconstruction method based on the basic reconstruction model described in the first aspect.

[0030] In a third aspect, the present invention further provides a non-volatile computer storage medium, wherein the computer storage medium stores computer-executable instructions, and the computer-executable instructions are executed by one or more processors to complete the method described in the first aspect.

[0031] In a fourth aspect, a chip is provided, comprising: a processor and an interface, for calling and running a computer program stored in a memory from a memory, and executing any method of the first aspect.

[0032] In a fifth aspect, a computer program product comprising instructions is provided, which, when executed on a computer or a processor, causes the computer or the processor to execute any of the methods of the first aspect.

[0033] The present invention first uses a human body point cloud prediction model to process binocular images (i.e., human body front image and human body back image) to obtain point clouds from four perspectives, then performs color assignment to restore the image under the side view, and finally uses a 3D Gaussian parameter regression model to restore the three-dimensional human body model. Therefore, only two images (i.e., front and back views) are needed to reconstruct the three-dimensional human body model. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments of the present invention. Obviously, the drawings described below are only some embodiments of the present invention. Those skilled in the art can also derive other drawings based on these drawings without inventive effort.

[0035] Figure 1 1 is a flow chart of a method for millisecond-level binocular human body reconstruction based on a basic reconstruction model provided by an embodiment of the present invention;

[0036] Figure 21 is a flow chart of a method for millisecond-level binocular human body reconstruction based on a basic reconstruction model provided by an embodiment of the present invention;

[0037] Figure 3 1 is a flow chart of a method for millisecond-level binocular human body reconstruction based on a basic reconstruction model provided by an embodiment of the present invention;

[0038] Figure 4 1 is a flow chart of a method for millisecond-level binocular human body reconstruction based on a basic reconstruction model provided by an embodiment of the present invention;

[0039] Figure 5 is a schematic diagram of a millisecond-level binocular human body reconstruction method based on a basic reconstruction model provided by an embodiment of the present invention;

[0040] Figure 6 is a schematic diagram of a millisecond-level binocular human body reconstruction method based on a basic reconstruction model provided by an embodiment of the present invention;

[0041] Figure 7 3D Gaussian parameter regression model in a millisecond-level binocular human body reconstruction method based on a basic reconstruction model provided by an embodiment of the present invention;

[0042] Figure 8 1 is a schematic diagram of the architecture of a millisecond-level binocular human body reconstruction device based on a basic reconstruction model provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0043] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0044] Unless the context requires otherwise, throughout the specification and claims, the term "including" is to be interpreted as meaning open inclusion, that is, "including, but not limited to". In the description of the specification, the terms "one embodiment", "some embodiments", "exemplary embodiments", "example", "specific example" or "some examples" and the like are intended to indicate that the specific features, structures, materials or characteristics associated with the embodiment or example are included in at least one embodiment or example of the present disclosure. The schematic representation of the above terms does not necessarily refer to the same embodiment or example. In addition, the specific features, structures, materials or characteristics may be included in any one or more embodiments or examples in any appropriate manner, that is, although they may be carried in the embodiments or examples of the above terms due to reasons such as the order and position of appearance, it is not limited to that they can be carried in combination by one embodiment or example.

[0045] In the description of the present invention, the terms "first" and "second" are used for descriptive purposes only, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Thus, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of the embodiments of the present disclosure, unless otherwise specified, "multiple" means two or more. In addition, for example, the description may also use the method of adding "A" and "B" at the end to describe the same type of nouns as two independent individuals. In this case, the corresponding features defined as "A" and "B" are only used to distinguish the description purposes of the same type of individuals, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated.

[0046] In the description of the present invention, the expression "A and / or B" (where A and B are used to formally represent specific characteristic contents) will be involved, and the corresponding expressions include the following three combinations: only A, only B, and a combination of A and B.

[0047] As used herein, "about," "substantially," or "approximately" includes the stated value and an average value that is within an acceptable range of deviation from the particular value as determined by one of ordinary skill in the art taking into account the measurements in question and the errors associated with the measurement of the particular quantity (i.e., the limitations of the measurement system).

[0048] In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0049] In order to make the objectives, technical solutions and advantages of the present invention more clear, the terms used in this embodiment are briefly described here:

[0050] DUSt3R: DUSt3R (full name: Dense Unsupervised Stereo 3D Reconstruction) is a dense 3D reconstruction algorithm based on deep learning, which aims to efficiently restore the 3D geometric structure of the scene through multi-view 2D images. Unlike the traditional multi-view stereo matching (MVS) method that relies on precise camera parameters and complex feature matching, DUSt3R adopts an unsupervised or self-supervised learning framework to directly predict pixel-level depth information or 3D point clouds through deep neural networks, significantly reducing the dependence on prior conditions. The core of the algorithm lies in its end-to-end optimization mechanism, which can jointly optimize depth prediction and camera pose to improve global geometric consistency while maintaining robustness to illumination changes and weak texture areas. Thanks to its dense reconstruction capabilities, DUSt3R can generate high-resolution 3D models, carefully preserve the tiny details in the scene, and achieve real-time or near real-time large-scale scene processing through GPU acceleration.

[0051] Nearest Neighbor Search (NNS): Nearest neighbor search is a technology that quickly finds the data point closest to a given query point in a dataset. It is widely used in fields such as machine learning, data mining, and information retrieval. Its core goal is to locate the most similar samples from large-scale high-dimensional data through efficient distance calculation and index structure, thereby supporting the underlying needs of tasks such as classification, clustering, and recommendation. Unlike model-based prediction methods, NNS directly relies on the geometric distribution characteristics of the data itself and does not require explicit training parameters. Instead, it accelerates the search process by constructing spatial indexes (such as KD trees, Ball trees, or hash tables) or using approximate algorithms (such as LSH local sensitive hashing). The advantage of this technology lies in its intuitiveness and flexibility, and it can adapt to a variety of distance metrics (such as Euclidean distance and cosine similarity).

[0052] Embodiment 1:

[0053] Many advanced human reconstruction methods have previously introduced 3D Gaussian Splatting (3DGS) to the field of human reconstruction, achieving state-of-the-art human reconstruction results and even high-quality animation effects. This method uses Gaussian points to reconstruct the scene and renders it in a splattering manner. Most work uses human priors to initialize the point cloud and accepts video as input for scene reconstruction. Furthermore, mainstream advanced algorithms usually accept monocular video or multi-view video as input, and multi-view video input generally produces better reconstruction results.

[0054] However, existing human reconstruction techniques often accept monocular or multi-view videos as input. When the number of input images is small, such as binocular images (i.e., images from the front and back), good reconstruction results are often difficult to achieve. Furthermore, most of these works lack generalization capabilities, requiring separate training for each scene. This leads to a high demand for input images and greatly increases the difficulty of data collection. To further improve the application of human reconstruction in real life, research on human reconstruction based on sparse viewpoints is of great significance.

[0055] Reconstructing the human body from sparse viewpoints has always been a challenging topic, and many previous works have studied this. However, the sparsity of the viewpoints accepted by these works is still limited. Some methods use human body prior knowledge to complete human body reconstruction from sparse viewpoints. However, human body prior requires a long preprocessing time and has difficulty processing loose clothing. Therefore, this embodiment proposes to explore an extremely challenging task-quickly reconstructing the human body with only two images (i.e., front and back views). This task enables any user to easily reconstruct the required three-dimensional digital human without the need for professional knowledge to capture redundant images or wait for the results for a long time. The main challenge is the lack of overlap between the front and back views, which makes it almost impossible to establish consistency for the reconstruction. In addition, the information from the two views is too limited to cover all the details of the entire human body. Experiments show that existing advanced methods are difficult to adapt to binocular tasks. In order to solve this problem, Example 1 of the present invention provides a millisecond-level binocular human body reconstruction method based on a basic reconstruction model, such as Figure 1 Shown, including:

[0056] In step 201, a human point cloud prediction model based on a basic reconstruction model is defined, and the human frontal image I is reconstructed using the human point cloud prediction model. f and the back image of the human body I b Processing is performed to obtain the human body front point cloud P f,f , point cloud P of the back of the human body b,f , point cloud P of the left side of the human body l,f And the right side point cloud P r,f Based on the prior of the basic reconstruction model, high-quality prediction of pixel-by-pixel alignment of image to point cloud is achieved. b,f It represents the point cloud of the back of the human body in the front view coordinate system. The same is true for other point clouds, that is, each point cloud is described with the front view coordinate system as a reference.

[0057] In step 202, a side image enhancement module is defined, and the side image enhancement module is used to analyze the left side point cloud P of the human body. l,f And the right side point cloud P r,f Perform color assignment to obtain the left side image of the human body I l and the right side image of the human body I r Without introducing other network parameters, this module overcomes the problem of missing side color information, effectively improves the consistency of human body reconstruction, and ensures the speed of current reconstruction methods.

[0058] In step 203, a 3D Gaussian parameter regression model is defined, and the human front image I f , human back image I b , human front point cloud P f,f , point cloud P of the back of the human bodyb,f , human left side image I l , human right side image I r , point cloud P of the left side of the human body l,f And the right side point cloud P r,f The input is fed into the 3D Gaussian parameter regression model, and the output is a 3D human body model. Considering the domain differences between images, this example uses two regression models to perform Gaussian attribute regression on the front, back, and side views.

[0059] In this embodiment, the human body point cloud prediction model is first used to predict the binocular image (i.e., the human body front image I f and the back image of the human body I b ) are processed to obtain point clouds from four perspectives, and color assignment is performed to restore the image under the side view. Finally, a 3D Gaussian parameter regression model is used to restore the three-dimensional human body model. Therefore, only two images (i.e., the front and back views) are needed to reconstruct the three-dimensional human body model.

[0060] In a specific application scenario, the human point cloud prediction model includes a binocular image encoder, a decoder, a frontal prediction head H f , back prediction head H b The human body point cloud prediction model also includes a side processing module, a left side prediction head H l and right face prediction head H r .

[0061] Wherein, the human body point cloud prediction model is used to predict the human body front image I f and the back image of the human body I b Processing is performed to obtain the human body front point cloud P f,f , point cloud P of the back of the human body b,f , point cloud P of the left side of the human body l,f And the right side point cloud P r ,f ,like Figure 2 As shown, specifically including:

[0062] In step 301, the binocular image encoder obtains the human body front image I f Extract the frontal features of the human body G f , and from the back image of the human body I b Extract the back feature G of the human body b The binocular image encoder adopts a ViT structure network, which is composed of a stack of Transformer blocks and loaded with weights pre-trained on multi-scene image pairs (8.5*10^6 image pairs) to achieve feature extraction of image information.

[0063] In step 302, the decoder converts the human body front feature Gf After mapping, it is input to the front prediction head H f , to obtain the human body front point cloud The back feature G of the human body b After mapping, it is input to the back prediction head H b , to obtain the point cloud of the back of the human body

[0064] in, is a positive feature sequence, is the back feature sequence, and B is the feature sequence length. The two prediction heads use the same network structure and use stacked Transformer blocks to input human body features G f or G b Process it and use the attention module to f or G b Processing, and then use the connection layer to get the target human point cloud output. The prediction head realizes the f or G b In the process of predicting human point cloud, more accurate point cloud prediction results are obtained by adding the attention module in the prediction head.

[0065] In step 303, the side processing module processes the front feature G of the human body. f and the back feature G b Perform average processing to obtain the side features of the human body And the human body side features Input to the left face prediction head H l , get the point cloud of the left side of the human body And the human body side features Input to the right face prediction head H r , get the point cloud of the right side of the human body The two side prediction heads introduced refer to the implementation in step 302 and extract the predicted point cloud based on Transformer.

[0066] The above steps 301 to 303 can be understood as: using the binocular image I of the standard training data set f and I b As the input of the human point cloud prediction model, the encoder and decoder are used to obtain two token sets G corresponding to the binocular image f and G b ; According to the two token sets G of the binocular image f and G b Input to the point cloud prediction head H f and H b , respectively predict the point cloud P corresponding to the binocular image corresponding to the front viewf,f and P b,f ,in, The two token sets G of the binocular image f and G b After averaging, we get the side view token set C l and C r (C l and C r The same, can be understood as the above G s );Right now According to the two token sets C of the binocular image l and G r Input to the point cloud prediction head H l and H r , respectively predict the point cloud P of the left and right sides under the front view l,f and P r,f ;

[0067]

[0068] The point cloud predicted by the human body point cloud prediction model is at the pixel level and satisfies the corresponding relationship with the input image. Thus, the positive and negative point cloud sets with color information can be obtained by confidence filtering. By using the same mask on the predicted point cloud feature map and the image, the point cloud set with corresponding color information can be extracted. The side point cloud is filtered by confidence to obtain the side point cloud set without color information, that is, the left side point cloud P of the human body. l,f And the right side point cloud P r,f .

[0069] In some embodiments, the side image enhancement module (ie Figure 5 The side color enhancement in the left side of the human body is performed on the point cloud P l,f And the right side point cloud P r,f Perform color assignment, and based on the correspondence between point cloud and color, finally obtain the left side image of the human body I l and the right side image of the human body I r ,like Figure 3 As shown, specifically including:

[0070] In step 401, the side image enhancement module is a point cloud P of the left side of the human body. l,f And the right side point cloud P r,f Each side point in the image is matched with the nearest front and back points, and the distance between the point clouds is measured using the Euclidean distance. The colors of the matched front and back points are assigned to the corresponding side points to obtain a colored side point cloud set. The front and back points are the front and back point clouds (i.e., the front point cloud P of the human body). f,f Or the back point cloud P b,fIt should be noted that the side point is a point in the side point cloud, the side point cloud includes the left side point cloud P l,f The points in the right side of the human body and the point cloud P l,f Taking the color of the point cloud on the left as an example, its color is defined as n is the number of point clouds; the nearest neighbor search function F is used nns Search Before and after point cloud {P f,f ,P b,f} label j, the process is expressed as Then use the label j to point the color set of the left and right points {C f,f ,C b,f} to get the color by index, the process is expressed as

[0071] In step 402, the colored side point cloud set is mapped to the pixel plane to obtain the left side image of the human body I l and the right side image of the human body I r , the mapping relationship is extracted from the colored positive and negative point cloud sets.

[0072] The above steps 401 to 402 can be understood as follows: the side image enhancement module is the point cloud P of the left side of the human body l,f The left side face point in the image matches the nearest positive and negative point (i.e. the human body front point cloud P f,f And the back point cloud P b,f The points in the set are also called the positive and negative points set {P f,f ,P b,f}), assign the colors of the matched positive and negative points to the corresponding left side face points to obtain the colored left side face point cloud set, map the colored left side face point cloud set to the pixel plane, and obtain the left side face image of the human body I l .

[0073] The side image enhancement module also generates the right side point cloud P r,f The right side face point in the image is matched with the nearest positive and negative points, and the colors of the matched positive and negative points are assigned to the corresponding right side face points to obtain a colored right side face point cloud set. The colored right side face point cloud set is mapped to the pixel plane to obtain the right side face image of the human body I r ; Wherein, the color of the left side point is n is the number of point clouds.

[0074] For the point cloud P on the left side of the human body l,f Every left side point i∈{1,2,…n}, using the nearest neighbor search function F nns Search In the positive and negative point set {Pf,f ,P b,f}, the corresponding label j, that is, Use label j to select the front and back color set {C f,f ,C b,f}Index to get the color The processing of the point cloud of the right side of the human body and the processing of the point cloud of the left side of the human body are based on the same concept and are not described in detail here.

[0075] The human side image enhancement model is composed of pseudo-perspective construction, front and back image projection algorithm and image enhancement algorithm. Based on the front and back point clouds with color information, the side point cloud and the front and back point clouds are associated to give color information to the side point cloud, and the real side image of the human body (i.e., the left side image of the human body) is obtained. l and the right side image of the human body I r ).

[0076] This embodiment associates the side point cloud with the front and back point clouds based on the point cloud's positional information, matching each side point cloud with its nearest front and back point clouds. After determining the correspondence between the point clouds, the color information of the front and back point clouds is assigned to the side point cloud, thereby obtaining a set of side point clouds with color information. Combined with the pixel-level relationships predicted by the basic reconstruction model, the color of the side point clouds is assigned back to the pixel plane to produce an enhanced side image.

[0077] This embodiment also provides an optional implementation method, that is, the 3D Gaussian parameter regression model includes a positive and negative view parameter regression module Side view parameter regression module The two modules use the same network structure, taking into account the domain differences between front, back, and side images, and use two network modules for regression. Each regression module receives the point cloud feature map and color feature map of the corresponding viewpoint and predicts the attribute value of the Gaussian point at the corresponding point cloud location.

[0078] The human body front image I f , human back image I b , human front point cloud P f,f , point cloud P of the back of the human body b,f , human left side image I l , human right side image I r , point cloud P of the left side of the human body l,f And the right side point cloud P r,f Input into the 3D Gaussian parameter regression model, and output a three-dimensional human body model, such as Figure 4 As shown, specifically including:

[0079] In step 501, the human front image I f, human back image I b , human front point cloud P f,f And the back point cloud P b,f Input to the front and back view parameter regression module To output front Gaussian points and back Gaussian points The Gaussian regression module is designed with reference to the U-Net structure, using convolution blocks for stacking. The network structure is symmetrically U-shaped, featuring both a traditional downsampling path (contraction path, used to capture contextual information of the image) and an upsampling path (expansion path, used for precise positioning). These two paths are connected by jump connections to help pass high-resolution information from early layers to later layers, thereby helping to accurately reconstruct object boundaries. After the U-Net network, the prediction heads of different attributes are connected to obtain the final attribute value. The network structure is referenced. Figure 7 .

[0080] In step 502, the left side image of the human body I l , human right side image I r , point cloud P of the left side of the human body l,f And the right side point cloud P r,f Input to the side view parameter regression module To output the left side Gaussian points and the right side Gaussian points Its network module is similar to the regression module in step 501.

[0081] In step 503, the front Gaussian point θ f 、Back Gaussian point θ b , Gaussian point θ on the left side l and the right side Gaussian point θ r Input to output module to output 3D human body model Indicates the merging operation. Considering the relatively low overlap between the four perspectives, the regression Gaussian points of the four perspectives are directly spliced ​​together to obtain the final Gaussian representation of the human body.

[0082] The Gaussian point consists of the center position μ, color c, opacity o, and covariance matrix Σ, i.e., Gaussian point θ = (μ, c, o, Σ). The Gaussian point can be controlled by different attributes to obtain a human body representation.

[0083] The human body point cloud prediction model is pre-trained, and the loss function used in the training is: Among them, L reg Represents the process of calculating the Euclidean distance, P h is the network prediction output, P gtFor training labels, the Euclidean distance represents the straight-line distance between pairs of points. For two n-dimensional vectors x = (x1, x2, ..., x n ) and y=(y1,y2,…,y n ), the Euclidean distance between them is defined as: Among them, the point cloud P under the two perspectives is merged f,f and P b,f Point cloud P from two side perspectives l,f and P r,f , get the human point cloud P output by network prediction h In the first stage, the point cloud is directly supervised at the 3D level, which can effectively ensure the distribution of the point cloud.

[0084] The specific training process is as follows: According to the designed loss function The back propagation algorithm is combined with the gradient descent method to iteratively optimize the model in order to minimize the overall target loss function, thereby optimizing the network model and its parameters. gt is the expected output of the network to predict the label P h To predict the network output, a target loss function is designed for the network model to determine the expected output and the predicted output, achieving the optimal network model and parameters. Before training, the human point cloud prediction model is pre-loaded with pre-trained model weights for common scenarios.

[0085] The 3D Gaussian parameter regression model is pre-trained, and the loss function used in the training is Among them, L rgb Represents the color difference between pixels. represents the calculated structural similarity index measure (SSIM) loss, and β is a preset coefficient obtained by those skilled in the art based on empirical analysis. render is the network output 3D human body model θ rendered image, I gt The training labels correspond to the training labels. The 3D Gaussian parameter regression model is trained using backpropagation and gradient descent algorithms. Phase 2 supervises the model at the 2D level. By rendering a human representation, two-dimensional images from different perspectives are generated. The rendered images are then compared to the real images and the loss function is calculated, thus achieving two-stage supervision.

[0086] The method of this embodiment further includes performing multi-angle rendering of the high-precision textured mesh in the original human body dataset to obtain rendered images at different poses. To enhance model robustness, a small amount of random noise is added to the rendered poses when rendering different human bodies to obtain a standard training dataset. This standard training dataset is used to train the human body point cloud prediction model and the 3D Gaussian parameter regression model.

[0087] Example 2:

[0088] The present invention is based on the method described in Example 1, combined with specific application scenarios, and uses technical descriptions in related scenarios to illustrate the implementation process of the present invention in characteristic scenarios.

[0089] This embodiment provides a millisecond-level binocular human body reconstruction method based on a basic reconstruction model, such as Figure 5 and Figure 6 As shown, specifically including:

[0090] (1) Training an end-to-end binocular human reconstruction method based on the basic reconstruction model, including the following sub-steps:

[0091] (1.1) Perform multi-angle surround rendering on the high-precision textured mesh file in the original human body dataset to obtain rendered images at different poses. To increase the model's robustness to the input image's perspective, a small amount of random noise is added to the rendered poses of different human bodies during the rendering process to obtain the corresponding 2D information annotation dataset.

[0092] (1.2) Define a human point cloud prediction model based on the basic reconstruction model. This geometric human point cloud prediction model consists of three core components: a binocular image encoder, a decoder, and a point cloud prediction head. The binocular image encoder is responsible for converting the input binocular image (i.e., an image taken from two different perspectives) into a high-dimensional vector or feature map that can represent its features, capturing the spatial information and detail features such as color in the image. The decoder receives the feature representation output by the encoder, maps it, and inputs it into the point cloud prediction head to predict the three-dimensional point cloud coordinates of the human body surface. According to the annotated standard training data set in (1.1), the training labels are calculated, and the corresponding loss function is designed. The backpropagation and gradient descent algorithms are used to train the geometric human point cloud prediction model based on the basic reconstruction model, which specifically includes the following sub-steps:

[0093] (1.2.1) Build a human point cloud prediction model based on the basic reconstruction model and load the weights of the general scene pre-training model, hoping to adapt the powerful geometric prior knowledge in the general scene to the specific human body field.

[0094] (1.2.2) First, the binocular image I in the standard training data set f (front view) and Ib (Back view) is input into the human point cloud prediction model. This process uses the powerful ability of deep learning to process the input binocular image through the encoder. The encoder is usually composed of a series of convolutional layers that can effectively extract feature information in the image, such as edges, textures, and color distribution, and convert the original image into a higher level of abstract representation. For the binocular image I f and I b , the encoder and decoder generate corresponding feature representations, which are called token sets G f and G b The two token sets obtained are fed into the point cloud prediction head H f and H b The point cloud prediction head is a module specially designed to infer 3D point clouds from feature representations. Specifically, H f Responsible for the token set G corresponding to the front view f Predict the point P from the front view f,f , and H b Then the token set G based on the back view b Predict the point P corresponding to the back view b,f The two sets of point clouds represent the geometric structures of the human body surface observed from different angles.

[0095]

[0096] (1.2.3) In order to capture the geometric structure of the human body more comprehensively, in addition to the front view and the back view, the side view information is also crucial. First, from the binocular image I f (front view) and I b (Back view) The two token sets G extracted f and G b Represent the human body feature representation under these two perspectives respectively. In order to obtain the feature representation of the side perspective, this embodiment adopts a strategy: G f and G b Perform averaging to obtain a new token set G l and G r , which correspond to the human body feature sets of the left and right perspectives respectively. The side view token set G obtained by averaging l and G r is input into a specially designed point prediction head H l and H r Middle. Point cloud prediction head H l and H rIt is a module specially optimized for side view, which can infer the 3D point cloud P corresponding to the left and right sides under the front view based on the input feature representation. l,f and P r,f .

[0097]

[0098] (1.2.4) Specifically, P f,f represents the point representation of the front view point cloud in the front camera coordinate system, and P b,f The corresponding back view is represented by points in the same coordinate system. These two point cloud datasets provide information on the geometric structure of the human body surface observed from different perspectives. f,f and P b,f ,By analyzing the corresponding feature points in these two sets of point cloud data and their spatial position relationship, the camera intrinsic parameter K can be accurately estimated.

[0099] (1.2.5) Since the points output by the human body point cloud prediction model are at the pixel level and have a pixel-level correspondence with the input image information, corresponding color information can be assigned to the front and back points to obtain a set of front and back points with color information, as well as a set of side point clouds without color information.

[0100] (1.2.6) Since the side pseudo-color information directly obtained from the front or back images through geometric transformation may suffer from data loss, color distortion, or unclear details, an effective method is needed to repair and optimize the quality of these side images. Here, an inpainting algorithm is introduced to fill in and improve the pseudo-color information estimation of the left and right sides.

[0101] This embodiment uses NNS to match each uncolored side point cloud with the nearest colored positive and negative point set, and assigns the color of the nearest positive and negative point to the side point cloud, thereby obtaining a side point cloud set with color information. The formula is as follows:

[0102]

[0103]

[0104] Similarly, based on the pixel-level correspondence between the predicted point cloud and the image, this embodiment further maps the side point cloud set with color information back to the pixel plane to obtain an enhanced side image, thereby completing and enhancing the side color map.

[0105] (1.2.7) Based on the current advanced 3D Gaussian splashing method, this embodiment designs a feed-forward 3D Gaussian parameter regression model, whose structure is as follows: Figure 7As shown in the figure, based on the human body point cloud predicted in (1.2.3), the human body is reconstructed with high quality by combining the color information of multiple perspectives. Considering the feature differences between the front and back views and the side views, this embodiment designs two modules to perform 3D Gaussian parameter regression respectively: Figure 3 D Gaussian parameter regression module and side view of the human body Figure 3 D Gaussian parameter regression module

[0106] (1.2.8) In order to achieve accurate prediction of the geometric structure and appearance features of the human body, the module is based on the human point cloud P predicted in step (1.2.3) f,f ,P b,f ,P l,f ,P r,f , and the input human front and back views I f ,I b and enhanced side view I l ,I r To perform comprehensive attribute prediction. This process is a key link in the entire reconstruction process, aiming to generate a high-precision, detailed three-dimensional human body model through deep fusion of multi-view information. In order to achieve high-precision human body reconstruction, this embodiment adopts an advanced 3D Gaussian Splatting (3DGS) technology to represent the human body. The 3D Gaussian point consists of the center position μ, color c, opacity o and covariance matrix ∑, which is specifically expressed as:

[0107] θ=(μ,c,o,∑)

[0108] (1.2.9) Based on the human attributes obtained using the 3D Gaussian splattering (3DGS) technique in the previous step, an intuitive and realistic 3D human model can be obtained. This process not only involves accurately capturing geometric shape and appearance features, but also requires considering how to efficiently integrate this information and present it in a high-quality manner. By setting different internal and external parameters, the reconstructed human body can be projected onto different camera planes to complete the rendering of the final result.

[0109] (1.2.10) In the process of building and training a deep learning model, a key step is to iteratively optimize the model using the backpropagation algorithm combined with the gradient descent method based on the designed overall target loss function, in order to minimize the overall target loss function, thereby optimizing the network model and its parameters. To ensure the normal convergence of the model, this embodiment defines different loss functions for the two stages of model training. The loss function for stage 1 based on the basic reconstruction model training in a general scenario is as follows, where L regRepresents the calculation of Euclidean distance.

[0110]

[0111] The loss function of the feedforward 3D Gaussian parameter regression model training in stage 2 is as follows, where L rgb Represents the color difference between pixels. Represents the calculated SSIM loss. render is the network output 3D human body model θ rendered image, I gt is the corresponding training label.

[0112]

[0113] By training the two-stage models separately, the overall model is optimized to achieve the expected human body reconstruction effect.

[0114] (2) Use the trained model to reconstruct the human body: input the two images to be reconstructed into the basic reconstruction model based on the general scene, and predict the points P on the four faces of the human body. f,f ,P b,f ,P l,f ,P r,f , the point representation P of the human body is obtained by splicing h Considering the color loss of the side view point cloud, this embodiment expects to convert the front and back color information I f ,I b Furthermore, this embodiment performs minimum distance matching on the side point cloud and the front and back point clouds, finds the nearest front and back point clouds for each side point cloud, and assigns the color of the front and back point clouds to the side point cloud, and finally obtains the color information of the side point cloud I l ,I r . The positive and negative color information I f ,I b and enhanced side color information I l ,I r , the positive and negative points P predicted by the basic reconstruction model f,f ,P b,f and side point cloud P l,f ,P r,f The data are fed into a feed-forward 3D Gaussian parameter regression model to predict the Gaussian point attributes of the human body. Based on the obtained human attributes, an intuitive and realistic 3D human body model can be obtained.

[0115] The method described in this embodiment can be understood as: passing the binocular image to be reconstructed through a binocular image encoder and decoder to obtain two tokens corresponding to the binocular image. Through two point cloud prediction heads, point clouds corresponding to the two binocular perspectives are obtained. This is to average the two token sets of the binocular image to obtain a token set for the side perspective, and then pass two point cloud prediction heads to obtain point clouds corresponding to the left and right side perspectives. Based on the colored front and back point clouds, we further associate the side point cloud with the front and back point clouds, thereby giving color to the side point clouds, and finally obtaining an enhanced side image. The binocular image and the obtained side enhanced image are respectively input into the front and back view of the human body. Figure 3 D Gaussian parameter regression and human side view Figure 3 D Gaussian parameter regression model is used to obtain the final human body reconstruction result.

[0116] In general, compared with the prior art, this embodiment has the following beneficial effects:

[0117] (1) High accuracy: This embodiment addresses the problem of human body reconstruction under sparse viewing angles and can achieve high-quality human body reconstruction quickly and efficiently based on the basic reconstruction model in general scenarios.

[0118] (2) Fast speed: The basic reconstruction model and 3D Gaussian parameter regression model proposed in this embodiment take a short time to infer, and do not require the calibration of internal and external parameters and the preprocessing of human body prior information, thus achieving millisecond-level human body reconstruction.

[0119] (3) Strong robustness: This embodiment achieves advanced reconstruction quality on both the in-domain test set and the out-of-domain test set, and can reconstruct the human body based on offline real data, with good robustness.

[0120] It should be noted that, for the purpose of privacy protection, the faces of the people in the figures of this embodiment are blurred.

[0121] Example 3:

[0122] like Figure 8 , is a schematic diagram of the architecture of a millisecond binocular human body reconstruction device based on a basic reconstruction model according to an embodiment of the present invention. The millisecond binocular human body reconstruction device based on a basic reconstruction model according to this embodiment includes one or more processors 21 and a memory 22. Figure 8 A processor 21 is taken as an example.

[0123] The processor 21 and the memory 22 may be connected via a bus or other means. Figure 8 The bus connection is taken as an example.

[0124] The memory 22, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs and non-volatile computer executable programs, such as the millisecond-level binocular human body reconstruction method based on the basic reconstruction model in Example 1. The processor 21 executes the millisecond-level binocular human body reconstruction method based on the basic reconstruction model by running the non-volatile software program and instructions stored in the memory 22.

[0125] The memory 22 may include high-speed random access memory and non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state memory device. In some embodiments, the memory 22 may optionally include a memory remotely located relative to the processor 21, and such remote memory may be connected to the processor 21 via a network. Examples of such networks include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0126] The program instructions / modules are stored in the memory 22 , and when executed by the one or more processors 21 , the millisecond-level binocular human body reconstruction method based on the basic reconstruction model in the above-mentioned embodiment 1 is executed.

[0127] It is worth noting that the information interaction, execution process, etc. between the modules and units within the above-mentioned devices and systems are based on the same concept as the processing method embodiment of the present invention. The specific content can be found in the description of the method embodiment of the present invention and will not be repeated here.

[0128] Those skilled in the art will understand that all or part of the steps in the various methods of the embodiments can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium, which may include: read-only memory (ROM), random access memory (RAM), a disk or an optical disk, etc.

[0129] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A millisecond-level binocular human body reconstruction method based on a basic reconstruction model, characterized in that: include: Define a human point cloud prediction model based on the basic reconstruction model, and use the human point cloud prediction model to reconstruct the human front image I f and the back image of the human body I b Processing is performed to obtain the human body front point cloud P f,f , point cloud P of the back of the human body b,f , point cloud P of the left side of the human body l,f And the right side point cloud P r,f , where P b,f Represents the point cloud of the back of the human body in the front view coordinate system, and the same applies to other point clouds; Define a side image enhancement module, and use the side image enhancement module to analyze the left side point cloud P of the human body. l,f And the right side point cloud P r,f Perform color assignment to obtain the left side image of the human body I l and the right side image of the human body I r ; Define a 3D Gaussian parameter regression model and transform the human front image I f , human back image I b , human front point cloud P f ,f , point cloud P of the back of the human body b,f , human left side image I l , human right side image I r , point cloud P of the left side of the human body l,f And the right side point cloud P r,f The input is to the 3D Gaussian parameter regression model, and the output is a three-dimensional human body model.

2. The millisecond-level binocular human body reconstruction method based on the basic reconstruction model according to claim 1 is characterized in that: The human point cloud prediction model includes a binocular image encoder, a decoder, a frontal prediction head H f and back prediction head H b ; The binocular image encoder is configured to obtain the human front image I f Extract the frontal features of the human body G f , and from the back image of the human body I b Extract the back feature G of the human body b ; The decoder converts the human body front feature G f After mapping, it is input to the front prediction head H f , to obtain the human body front point cloud The back feature G of the human body b After mapping, it is input to the back prediction head H b , to obtain the point cloud of the back of the human body in, is a positive feature sequence, is the back feature sequence, and B is the feature sequence length.

3. The millisecond-level binocular human body reconstruction method based on the basic reconstruction model according to claim 2 is characterized in that: The human body point cloud prediction model also includes a side processing module, a left side prediction module H l and right face prediction head H r ; The side processing module processes the front feature G of the human body f and the back feature G b Perform average processing to obtain the side features of the human body And the human body side features Input to the left face prediction head H l , get the point cloud of the left side of the human body And the human body side features Input to the right face prediction head H r , get the point cloud of the right side of the human body 4. The millisecond-level binocular human body reconstruction method based on the basic reconstruction model according to claim 1 is characterized in that: Specifically include: The side image enhancement module is the point cloud P of the left side of the human body l,f The left side face point in the image is matched with the nearest positive and negative points, and the colors of the matched positive and negative points are assigned to the corresponding left side face points to obtain a colored left side face point cloud set. The colored left side face point cloud set is mapped to the pixel plane to obtain the left side face image of the human body I l ; The side image enhancement module also generates the right side point cloud P r,f The right side face point in the image is matched with the nearest positive and negative points, and the colors of the matched positive and negative points are assigned to the corresponding right side face points to obtain a colored right side face point cloud set. The colored right side face point cloud set is mapped to the pixel plane to obtain the right side face image of the human body I r ; Among them, the color of the left side point is n is the number of point clouds; for the left side point cloud P l,f Every left side point i∈{1,2,…n}, using the nearest neighbor search function F nns Search In the positive and negative point set {P f,f ,P b,f }, the corresponding label j, that is, Use label j to select the front and back color set {C f,f ,C b,f }Index to get the color 5. The millisecond-level binocular human body reconstruction method based on the basic reconstruction model according to claim 1 is characterized in that: The 3D Gaussian parameter regression model includes a front and back view parameter regression module Side view parameter regression module and output modules; The human front image I f , human back image I b , human front point cloud P f,f And the back point cloud P b,f Input to the front and back view parameter regression module To output front Gaussian points and back Gaussian points The human body left side image I l , human right side image I r , point cloud P of the left side of the human body l,f And the right side point cloud P r,f Input to the side view parameter regression module To output the left side Gaussian points and the right side Gaussian points The front Gaussian point θ f 、Back Gaussian point θ b , Gaussian point θ on the left side l and the right side Gaussian point θ r Input to output module to output 3D human body model Represents a merge operation.

6. The millisecond-level binocular human body reconstruction method based on the basic reconstruction model according to claim 5 is characterized in that: The Gaussian point includes the center position μ, color c, opacity o and covariance matrix Σ, that is, the Gaussian point θ = (μ, c, o, Σ).

7. The millisecond-level binocular human body reconstruction method based on the basic reconstruction model according to claim 1 is characterized in that: The human body point cloud prediction model is pre-trained, and the loss function used in the training is: Among them, L reg Represents the process of calculating the Euclidean distance, P h is the network prediction output, P gt is the training label; the Euclidean distance represents the straight-line distance between the point pairs. For two n-dimensional vectors x=(x1,x2,…,x n ) and y=(y1,y2,…,y n ), the Euclidean distance between them is 8. The millisecond-level binocular human body reconstruction method based on the basic reconstruction model according to claim 1 is characterized in that: The 3D Gaussian parameter regression model is pre-trained, and the loss function used in the training is Among them, L rgb Represents the color difference between pixels. represents the calculated SSIM loss, β is the preset coefficient, I render is the network output 3D human body model θ rendered image, I gt is the corresponding training label.

9. A millisecond-level binocular human body reconstruction device based on a basic reconstruction model, characterized in that: The device comprises: At least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the processor to perform the millisecond-level binocular human body reconstruction method based on the basic reconstruction model according to any one of claims 1 to 8.

10. A non-volatile computer storage medium, characterized in that The computer storage medium stores computer-executable instructions, which are executed by one or more processors to complete the millisecond-level binocular human body reconstruction method based on the basic reconstruction model described in any one of claims 1-8.