Character image reconstruction method and device, electronic equipment and storage medium

By replacing the image of the online communication subject with an image of a human model and removing facial information other than pose, the problem of emotion perception in online communication is solved, and the accuracy and immersion of the communication subject's emotion perception are improved.

CN120912779APending Publication Date: 2025-11-07AGRICULTURAL BANK OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511070885.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-31
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Existing online communication methods make it difficult to accurately determine the true emotions of the person being communicated with, making it difficult for staff to assess the level of satisfaction with the person being communicated with.

Method used

By acquiring an image of the first object and sending it to a pre-trained human model reconstruction model, the image is replaced with a human model image. Other facial features, except for pose, are removed, while pose is preserved to enhance immersion and provide a basis for analysis.

Benefits of technology

This approach achieves the goal of maintaining an immersive image service while removing sensitive information, providing a more accurate basis for emotion perception and laying an analytical foundation for future service improvements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120912779A_ABST
    Figure CN120912779A_ABST
Patent Text Reader

Abstract

The invention discloses a figure image reconstruction method and device, electronic equipment and a storage medium. The method comprises the steps that in response to a first instruction, a first image is acquired, and the first instruction is an instruction which is sent by a first object and allows acquisition of the first image; the first image is an image comprising a first object; sending the first image to a pre-trained character model reconstruction model to obtain a second image; the second image is an image obtained by replacing the first object in the first image with a human body model. According to the technical scheme, the first image is obtained in response to the first instruction; and sending the first image to the pre-trained character model reconstruction model to obtain the second image, so that the pose of the first object can be reserved while removing sensitive information such as other appearances except the pose of the first object, thereby providing an analysis basis for subsequent service improvement while the immersion of the current image service is realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of video communication, and in particular to a character image reconstruction method and device, electronic equipment and a storage medium. BACKGROUND

[0002] With the continuous growth of network transmission speed, image transmission, video telephone and the like are increasingly favored by people, and part of offline communication is gradually transferred to online processing.

[0003] Existing online communication often directly obtains appearance information and behavior information of a communication object, but the appearance information of the first object is often not provided by the first object, so it is difficult to accurately determine the satisfaction degree of the communication object to the service. How to improve the staff's perception of the real emotions of the communication object has become a difficult problem. SUMMARY

[0004] The present application provides a character image reconstruction method and device, electronic equipment and a storage medium to solve the problem of improving the staff's perception of the real emotions of the communication object.

[0005] According to an aspect of the present application, a character image reconstruction method is provided, which comprises:

[0006] In response to a first instruction, a first image is obtained, the first instruction being an instruction issued by a first object to allow the first image to be obtained; the first image comprising an image of the first object;

[0007] The first image is sent to a pre-trained character model reconstruction model to obtain a second image; the second image being an image obtained by replacing the first object in the first image with a human body model.

[0008] According to another aspect of the present application, a character image reconstruction device is provided, which comprises:

[0009] A first image determination module is configured to obtain a first image in response to a first instruction, the first instruction being an instruction issued by a first object to allow the first image to be obtained; the first image comprising an image of the first object;

[0010] A second image determination module is configured to send the first image to a pre-trained character model reconstruction model to obtain a second image; the second image being an image obtained by replacing the first object in the first image with a human body model.

[0011] According to another aspect of the present application, an electronic device is provided, which comprises:

[0012] At least one processor; and

[0013] A memory in communication connection with the at least one processor; wherein,

[0014] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to perform the person image reconstruction method of any of the embodiments of the present application.

[0015] According to another aspect of the present application, a computer readable storage medium is provided, which stores computer instructions for enabling a processor to implement the person image reconstruction method of any of the embodiments of the present application when the processor executes the computer instructions.

[0016] The technical solution of the embodiments of the present application can achieve the following effects: in response to the first instruction, the first image is acquired; the first image is sent to the pre-trained person model reconstruction model to obtain the second image, which can realize the elimination of sensitive information such as appearance of the first object in addition to the pose while retaining the pose of the first object, thereby providing an analysis basis for subsequent service improvement while improving the immersion of the current image service.

[0017] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the present application, nor is it used to limit the scope of the present application. Other features of the present application will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced as follows. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.

[0019] Figure 1 is a flowchart of a person image reconstruction method according to the first embodiment of the present application;

[0020] Figure 2 is a model down-sampling schematic diagram according to the embodiments of the present application;

[0021] Figure 3 is a schematic diagram of constructing a mixed feature vector according to the embodiments of the present application;

[0022] Figure 4 is a structural schematic diagram of a person image reconstruction device according to the third embodiment of the present application;

[0023] Figure 5 is a structural schematic diagram of an electronic device for implementing the person image reconstruction method of the present application. DETAILED DESCRIPTION

[0024] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0025] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0026] Example 1

[0027] Figure 1 This is a flowchart of a method for reconstructing a person's image, provided in Embodiment 1 of the present invention. This embodiment is applicable to scenarios such as online calls and image transmission, where the pose of a person needs to be determined. This method can be executed by a person image reconstruction device, which can be implemented in hardware and / or software and can be configured in an electronic device with data processing capabilities. Figure 1 As shown, the method includes:

[0028] S110. In response to a first instruction, acquire a first image, wherein the first instruction is issued by a first object and is an instruction that allows the acquisition of the first image; the first image includes an image of the first object.

[0029] The first instruction can be issued by the first object, allowing the acquisition of a first image containing the first object. The first instruction can be issued after prior permission or upon real-time request. The first instruction can be issued by the first object via voice, screen click, or other means. The first image can be an image containing facial information such as the first object's body posture. The first image can be an image obtained by currently capturing the first object, or an image obtained by pre-capturing the first object. The first image can be acquired by the first object uploading it, or by capturing it through a camera that the first object has pre-granted permission to capture.

[0030] S120, send the first image to the pre-trained character model reconstruction model to obtain a second image; the second image is an image obtained by replacing the first object in the first image with a character model.

[0031] The character model reconstruction model can be pre-trained to replace a character in an image with a character model that only has a human body pose.

[0032] By sending the first image to the character model reconstruction model, the first object in the first image is replaced with a character model that only includes a human body pose by the character model reconstruction model, and the second image is obtained.

[0033] By sending the first image to the pre-trained character model reconstruction model to obtain the second image, the sensitive information of the first object other than the pose can be removed while the pose of the first object is retained, thereby improving the immersion of the current image service and providing an analysis basis for subsequent service improvement.

[0034] In an optional solution, sending the first image to the pre-trained character model reconstruction model to obtain the second image can include steps A1-A3:

[0035] Step A1, determining a third image according to the first image, the third image being obtained by cropping the image region where the first object is located in the first image.

[0036] Step A2, replacing the first object in the third image with a human body model to obtain a fourth image.

[0037] Step A3, determining the second image according to the fourth image and the first image.

[0038] The third image can be obtained by cropping the first object in the first image. The cropping of the first object in the first image can be to crop a rectangular image. In order to meet the input picture size of the resnet network, the first object is cropped to determine the third image, while the first object is as complete as possible.

[0039] In order to reduce the calculation amount of the character model reconstruction model and thereby reduce the calculation pressure of the character model reconstruction model, the region in the first image that is irrelevant to the replacement of the human body model can be removed as much as possible.

[0040] To this end, the image region where the first object is located in the first image is cropped to obtain a third image with smaller data volume and containing a complete first object. The third image is input to the character model reconstruction model instead of the first image, and the first object in the third image is replaced by the human body model by the character model reconstruction model to obtain a fourth image. Finally, the fourth image and the first image are integrated to obtain the final second image.

[0041] For example, a single first image of any resolution is determined, a human body detector based on YOLOv3 or FasterR-CNN is used to locate the image region of the first object in the first image, and a rectangular region where the image region is located is intercepted as a third image.

[0042] In an alternative, the first object in the third image is replaced by a human body model to obtain a fourth image, including steps B1-B5:

[0043] Step B1, the third image is scaled to obtain a fifth image; the number of pixels in the vertical direction, the number of pixels in the horizontal direction, and the number of channels of the fifth image are all pre-set sizes.

[0044] Step B2, the fifth image is down-sampled to obtain a global feature vector.

[0045] Step B3, based on the global feature vector, a first human body model of the first object is determined; the human body model has a plurality of vertices.

[0046] Step B4, the first human body model is adjusted to obtain a second human body model.

[0047] Step B5, based on the second human body model and the third image, a fourth image is obtained.

[0048] To further improve the calculation efficiency of the human model reconstruction model, reduce the amount of data required for calculation of the human model reconstruction model, and reduce the operation pressure of the human model reconstruction model, the third image is further scaled to obtain a fifth image.

[0049] To improve the accuracy of the human body model, the fifth image is down-sampled to obtain a global feature vector, and the global feature vector is used to complete the initial construction of the first human body model. To further improve the accuracy of the human body model, the first human body model is adjusted to obtain a second human body model, and finally based on the second human body model and the third image, a fourth image is obtained.

[0050] For example, the third image is scaled to a fifth image of 224x224x3 size. The third image is down-sampled to finally obtain a global feature vector of 2048 dimensions. The global feature vector is used to determine a first human body model of the first object. The first human body model is adjusted to obtain a second human body model. Based on the second human body model and the third image, a fourth image is obtained.

[0051] In an alternative, the first human body model is adjusted to obtain a second human body model, including steps C1-C5:

[0052] Step C1, down-sampling the first human body model to obtain a third human body model.

[0053] Step C2, projecting each vertex of the third human body model to the image plane and determining a local feature vector in the projection result based on bilinear interpolation.

[0054] Step C3, concatenating the vertex and the local feature vector and the global feature vector corresponding to the vertex to obtain a hybrid feature vector.

[0055] Step C4, determining a three-dimensional offset of each vertex based on the geometric perception graph convolution network and each hybrid feature vector.

[0056] Step C5, up-sampling based on the three-dimensional offset of each vertex and the hybrid feature vector corresponding to each vertex to obtain a second human body model.

[0057] To improve the subsequent optimization efficiency, the first human body model is down-sampled to obtain a third human body model, and the human body key topological structure is maintained during the down-sampling process, such as retaining the limb joints, the trunk midline and other feature areas. Each vertex after down-sampling is projected to the image plane, and the local feature vector is extracted through bilinear interpolation. The vertex and the local feature vector and the global feature vector corresponding to the vertex are concatenated to obtain a hybrid feature vector. That is, each hybrid feature vector is composed of a local feature vector, a global feature vector and a vertex standard coordinate. Each hybrid feature vector is input into the geometric perception graph convolution network to obtain a three-dimensional offset of each vertex.

[0058] The geometric perception graph convolution network includes several layers of graph neural networks, each layer performing: ① vertex feature transformation; ② neighbor-based message aggregation; and ③ edge feature encoding (including Euclidean distance, normal angle and other geometric information).

[0059] Finally, based on the three-dimensional offset of each vertex and the hybrid feature vector corresponding to each vertex, up-sampling is performed to obtain a second human body model with the same number of vertices as the first human body model.

[0060] For example, the first human body model has 6890 vertices. The edge contraction algorithm based on the second error metric is used to down-sample the 6890 vertices by 4 times to obtain 1723 vertices, and the second stage is further down-sampled by 4 times to obtain 431 vertices. The 431 down-sampled vertices are projected to the image plane, and the local feature vector is extracted from the ResNet conv4 feature map (size 14x14x2048) through bilinear interpolation. The final feature of each vertex is composed of three parts: a local feature vector (512 dimensions), a global feature vector (2048 dimensions) and a vertex standard coordinate (3 dimensions), forming a hybrid feature vector of 2563 dimensions.

[0061] Geometric-aware graph convolution is performed on each mixed feature vector, and a 3D offset AV of each vertex is calculated. The 431 vertices are upsampled to 1723, and then to 6890 complete vertices.

[0062] In the upsampling process, a residual connection is introduced, and the position of each vertex of the first human body model is taken as a basic coordinate.

[0063] The specific loss function is as follows:

[0064]

[0065] wherein, represents the human body parameter loss; x represents the 3D coordinates of the predicted vertex; y represents the real 3D coordinates of the vertex; V represents the predicted vertex set; Y is the real 3D point cloud data of each vertex.

[0066]

[0067] wherein, represents the normal vector loss; |F| represents the number of triangular facets in the human body model; n f represents the normal vector of the predicted triangular facet f; represents the normal vector corresponding to the real scan data of the triangular facet f.

[0068]

[0069] wherein, represents the grid edge length loss; |E| represents the number of grid edges; e represents the predicted grid edge vector; e * represents the standard edge length of the SMPL template.

[0070]

[0071] represents the total loss function; μ1, μ2 and μ3 represent weights.

[0072] In an optional solution, based on the global feature vector, a first human body model of the first object is determined, including steps D1-D2:

[0073] Step D1, input the global feature vector into a pre-constructed regression structured human body parameterization model to obtain at least one vertex and the pose parameter, shape parameter and camera parameter of each vertex; the parameter prediction model is composed of a preset number of fully connected layers.

[0074] Step D2, according to the pose parameter, shape parameter and camera parameter of each vertex, a first human body model of the first object is constructed.

[0075] The regression structured human parametric model estimates the parameters of the parametric model that can structurally describe the human shape, pose, expression and other attributes by regression method (mapping from input data to continuous parameters). The regression SMPL model estimates 72-dimensional pose parameters θ (using 6D continuous rotation representation), 10-dimensional shape parameters β and 3-dimensional camera parameters π. These parameters instantly generate an initial human model M_coarse = SMPL(θ0, β0) containing 6890 vertices.

[0076] The loss function of the regression SMPL model is:

[0077]

[0078] wherein, represents the L2 regularization loss of the regression SMPL model; θ represents the pose parameters, β represents the shape parameters, and π represents the camera parameters; θ * represents the real pose parameters, β * represents the real shape parameters, and π * represents the real camera parameters.

[0079]

[0080] represents the 3D joint position loss; J k represents the predicted kth joint coordinate; represents the real kth joint coordinate; k represents the number of joints.

[0081]

[0082] wherein, represents the 2D re-projection loss; Proj(J k ) is the 3D joint J k projected to the 2D image coordinate by the camera parameter π; represents the labeled 2D joint coordinate of J k .

[0083]

[0084] represents the total loss; λ1, λ2 and λ3 represent the weights.

[0085] Optionally, the fifth image is down-sampled to obtain a global feature vector, including:

[0086] The fifth image is input into the fifty-layer residual network, and the output result of the last pooling layer in the fifty-layer residual network is taken as the global feature vector.

[0087] A fifty-layer residual network (ResNet-50) which has an overall structure that can be divided into an input layer, four residual block stages, a global pooling layer, and a fully connected layer.

[0088] The fifth image is input into the fifty-layer residual network, and after layer-by-layer down-sampling, a global feature tensor of 2048 is obtained after passing through the global pooling layer.

[0089] The training process of the character model reconstruction model includes:

[0090] At least one first sample image and the corresponding second sample image of each first sample image are determined, the first sample image is an image containing a character, and the second sample image is an image obtained by replacing the character in the first sample image with a human body model;

[0091] The character model reconstruction model is trained according to at least one first sample image and the corresponding second sample image of each first sample image.

[0092] By responding to the first instruction, the first image is obtained, the first image is sent to the pre-trained character model reconstruction model, and the second image is obtained, which can realize the elimination of sensitive information such as appearance of the first object in addition to the pose, while retaining the pose of the first object, thereby providing an analysis basis for subsequent service improvement while improving the immersion of the current image service.

[0093] Embodiment two

[0094] The embodiment of the application provides a specific example of a character image reconstruction method, which is based on the above-mentioned embodiment and provides a complete example of the process in the above-mentioned embodiment. The embodiment can be combined with each optional scheme in one or more of the above-mentioned embodiments. The example is as follows:

[0095] A single first image of any resolution is used as input, a human body detector based on YOLOv3 or Faster R-CNN is used to locate the image region where the first object is located in the image, and the image region where the human body is located is intercepted to obtain a third image. The third image obtained by cutting is scaled to a standard size of 224x224x3 to obtain a fifth image. This standardization processing ensures the size consistency of the subsequent network input.

[0096] The fifth image is input into the improved ResNet-50 backbone network, which removes the last fully connected layer of the original architecture, and finally outputs a 2048-dimensional channel feature at a 7x7 spatial resolution, forming a global feature vector with a dimension of 7x7x2048. After global average pooling, a 2048-dimensional global feature vector is obtained, while its spatial feature map is retained for subsequent local feature extraction.

[0097] The 2048-dimensional global feature vector is input into a parameter prediction head consisting of three fully connected layers, which regress the 72-dimensional pose parameters θ (using 6D continuous rotation representation), 10-dimensional shape parameters β, and 3-dimensional camera parameters π of the SMPL model. These parameters generate an initial coarse human model M_coarse = SMPL(θ0, β0) containing 6890 vertices on the fly.

[0098] The goal of the first stage is to accurately predict the pose (θ), shape (β), and camera parameters (π) of the SMPL model and ensure the accuracy of 3D joint positions and 2D projections. The specific loss functions are as follows:

[0099]

[0100] where, represents the loss of L2 regularization of the SMPL model; θ represents the pose parameters, β represents the shape parameters, and π represents the camera parameters; θ * represents the real pose parameters, β * represents the real shape parameters, and π* represents the real camera parameters.

[0101]

[0102] represents the 3D joint position loss; J k represents the predicted k-th joint coordinate; represents the real k-th joint coordinate; k represents the number of joints.

[0103]

[0104] where, represents the 2D re-projection loss; Proj(J k ) is the 3D joint J k projected to 2D image coordinates through camera parameters π; represents the labeled 2D joint coordinates of J k .

[0105]

[0106] represents the total loss; λ1, λ2, and λ3 represent the weights.

[0107] See Figure 2, the first stage reduces 6890 vertices by 4 times to 1723 vertices through edge collapse algorithm based on quadratic error metric, and the second stage further reduces 4 times to 431 vertices. Both stages of sampling maintain the key topological structure of the human body, such as preserving the joints of the limbs, the midline of the torso, and other feature areas. This stage focuses on the overall posture and body shape features of the human body, providing reliable initial estimates for subsequent optimization.

[0108] As shown in Figure 3 : For each down-sampled vertex, three types of key features are spliced: image local ROI features extracted by vertex projection and bilinear interpolation, global context feature vectors from the first stage, and normalized geometric coordinates of the vertex in the standard SMPL space. Project 431 down-sampled vertices to the image plane, and extract local visual features from the conv4 feature map (size 14x14x2048) of ResNet through bilinear interpolation. The final feature of each vertex is composed of three parts: local image features (512 dimensions), global feature vectors (2048 dimensions), and standard coordinates of the vertex (3 dimensions), forming a hybrid feature representation of 2563 dimensions.

[0109] Geometric-aware graph convolution: a graph neural network containing 6 layers is designed, each layer performing: ① vertex feature transformation; ② message aggregation based on k=8 neighbors; ③ edge feature encoding (including Euclidean distance, normal angle, etc. Geometric information). The network finally outputs a 3D offset ΔV for each vertex.

[0110] Hierarchical upsampling: first upsample 431 vertices to 1723 vertices through a precomputed interpolation matrix, and then upsample to 6890 complete vertices, preserving the original SMPL network topology. The upsampling process introduces residual connections, taking the vertex positions of the initial SMPL model as the base coordinates.

[0111] The goal of the second stage is to optimize the vertex positions to better fit the real human surface while maintaining the smoothness and reasonableness of the mesh. The specific loss function is as follows:

[0112] The specific loss function is as follows:

[0113]

[0114] where, represents the human parameter loss; x represents the 3D coordinates of the predicted vertex; y represents the real 3D coordinates of the vertex; V represents the predicted vertex set; Y is the real 3D point cloud data of each vertex.

[0115]

[0116] where, represents the normal vector loss; |F| represents the number of triangular patches in the human model; nf represents a normal vector of the predicted triangle f; represents a normal vector corresponding to the real scan data of the triangle f.

[0117]

[0118] wherein, represents a mesh edge length loss; |E| represents the number of mesh edges; e represents a predicted mesh edge vector; e * represents a standard edge length of the SMPL template.

[0119]

[0120] represents a total loss function; μ1, μ2 and μ3 represent weights.

[0121] Finally, a fixed directional parallel light is adopted, a uniform diffuse reflection color is specified, and 6890 mesh vertices are rasterized by an output rendering module.

[0122] Embodiment three

[0123] Figure 4 A structural block diagram of a person image reconstruction device is provided for the embodiments of the present application. The embodiments can be applicable to online calls, image transmission and other scenarios to determine the pose of a person. The person image reconstruction device can be realized in the form of hardware and / or software, and can be configured in an electronic device with data processing capability. As shown in the figure, the person image reconstruction device of the embodiments can include a first image determination module 310 and a third image determination module 320. Among them: Figure 4

[0124] The first image determination module 310 is configured to acquire a first image in response to a first instruction. The first instruction is an instruction issued by a first object to allow the first image to be acquired. The first image is an image including the first object.

[0125] The second image determination module 320 is configured to send the first image to a pre-trained person model reconstruction model to obtain a second image. The second image is an image obtained by replacing the first object in the first image with a human body model.

[0126] On the basis of the above-mentioned embodiments, the second image determination module 320 can optionally include:

[0127] According to the first image, a third image is determined, and the third image is obtained by intercepting the image region where the first object is located in the first image.

[0128] The fourth image is obtained by replacing the first object in the third image with a human body model. ​

[0129] determining the second image according to the fourth image and the first image.

[0130] On the basis of the above-mentioned embodiments, optionally, the first object in the third image is replaced with a human body model to obtain the fourth image, comprising:

[0131] scaling the third image to obtain a fifth image; the fifth image has a preset size in the number of pixels in the vertical direction, the number of pixels in the horizontal direction, and the number of channels;

[0132] down-sampling the fifth image to obtain a global feature vector;

[0133] determining a first human body model of the first object based on the global feature vector; the human body model has a plurality of vertices;

[0134] adjusting the first human body model to obtain a second human body model;

[0135] obtaining the fourth image based on the second human body model and the third image.

[0136] On the basis of the above-mentioned embodiments, optionally, the first human body model is adjusted to obtain the second human body model, comprising:

[0137] down-sampling the first human body model to obtain a third human body model;

[0138] projecting each vertex of the third human body model to an image plane, and determining a local feature vector in the projection result based on bilinear interpolation;

[0139] splicing the vertices, the local feature vectors corresponding to the vertices, and the global feature vector to obtain a mixed feature vector;

[0140] determining a three-dimensional offset of each vertex based on a geometric perception graph convolution network and each mixed feature vector;

[0141] up-sampling based on the three-dimensional offset of each vertex and the mixed feature vector corresponding to each vertex to obtain the second human body model.

[0142] On the basis of the above-mentioned embodiments, optionally, the first human body model of the first object is determined based on the global feature vector, comprising:

[0143] inputting the global feature vector into a pre-constructed regression structured human parameterization model to obtain at least one vertex, and pose parameters, shape parameters, and camera parameters of each vertex; the parameter prediction model is composed of a preset number of fully connected layers;

[0144] constructing the first human body model of the first object according to the pose parameters, the shape parameters, and the camera parameters of each vertex.

[0145] On the basis of the above-mentioned embodiments, optionally, the fifth image is down-sampled to obtain a global feature vector, comprising:

[0146] The fifth image is input into a fifty-layer residual network, and an output result of a last pooling layer in the fifty-layer residual network is taken as the global feature vector.

[0147] On the basis of the above-mentioned embodiments, optionally, the training process of the person model reconstruction model comprises:

[0148] At least one first sample image and a second sample image corresponding to each first sample image are determined, the first sample image is an image containing a person, and the second sample image is an image after replacing the person in the first sample image with a human body model;

[0149] The person model reconstruction model is trained according to the at least one first sample image and the second sample image corresponding to each first sample image.

[0150] The person image reconstruction device provided in the embodiments of the present application can execute the person image reconstruction method provided in any of the embodiments of the present application, and has the corresponding function modules and beneficial effects of the execution method.

[0151] Embodiment four

[0152] Figure 5 A structural schematic diagram of an electronic device 10 that can be used to implement embodiments of the present application is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular telephones, smart phones, wearable devices (e.g., headsets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions, are meant to be examples only, and are not intended to limit the implementations of the present application described and / or claimed in this document.

[0153] As Figure 5As shown, the electronic device 10 includes at least one processor 11, and a memory, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., connected to the at least one processor 11 in communication. The memory stores a computer program executable by the at least one processor 11, and the processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0154] Various components in the electronic device 10 are connected to the I / O interface 15, including an input unit 16, such as a keyboard, a mouse, etc., an output unit 17, such as various types of displays, a speaker, etc., a storage unit 18, such as a magnetic disk, an optical disk, etc., and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.

[0155] The processor 11 can be various general and / or special-purpose processing components having processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 performs various methods and processes described above, such as the person image reconstruction method.

[0156] In some embodiments, the person image reconstruction method can be implemented as a computer program tangibly embodied in a computer readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the person image reconstruction method described above can be performed. Alternatively, in other embodiments, the processor 11 can be configured to perform the person image reconstruction method by any other appropriate means, such as by means of firmware.

[0157] The various embodiments of the systems and techniques described above can be implemented in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a load programmable logic device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0158] Computer programs used to implement the processes of the application can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the computer program

[0159] In the context of the present application, a computer-readable storage medium can be a tangible medium that can contain or store computer programs for use by or in connection with an instruction execution system, apparatus, or device. Computer-readable storage media can include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium will include one or more lines of electrical connections, portable computer disks, hard disk drives, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fibers, portable compact disc read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0160] To provide for interaction with a user, the systems and techniques described here can be implemented on an electronic device having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0161] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0162] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. A server can be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system, to solve the defects of large management difficulty and weak business scalability in traditional physical host and VPS service.

[0163] It should be understood that the various forms of flow shown above can be re-ordered, added to, or deleted from without departing from the scope of the present disclosure. For example, the steps recited in the present disclosure can be executed in parallel, executed in sequence, or executed in a different order, as long as the desired results of the present disclosure are achieved, and the present disclosure is not limited herein.

[0164] The specific embodiments described above are not intended to be limiting, and persons skilled in the art will appreciate that various modifications, combinations, sub-combinations and alternatives can be made to the specific embodiments without departing from the spirit and principles of the disclosure. Accordingly, the disclosure is not limited to the specific embodiments described above, but only by the scope of the appended claims.

Claims

1. A person image reconstruction method, characterized by, The method comprises the following steps: in response to a first instruction, acquiring a first image, the first instruction being an instruction issued by a first object and allowing the first image to be acquired; the first image being an image including the first object; sending the first image to a pre-trained human model reconstruction model to obtain a second image; the second image being an image obtained by replacing the first object in the first image with a human model.

2. The method of claim 1, wherein, The method comprises the following steps: determining a third image according to the first image, the third image being obtained by cropping an image region in which the first object is located in the first image; replacing the first object in the third image with a human model to obtain a fourth image; determining the second image according to the fourth image and the first image.

3. The method of claim 2, wherein, The method comprises the following steps: scaling the third image to obtain a fifth image; the fifth image having a predetermined size in terms of the number of pixels in the vertical direction, the number of pixels in the horizontal direction, and the number of channels; performing down-sampling on the fifth image to obtain a global feature vector; determining a first human model of the first object based on the global feature vector; the human model having a plurality of vertices; adjusting the first human model to obtain a second human model; obtaining the fourth image based on the second human model and the third image.

4. The method of claim 3, wherein, The method comprises the following steps: down-sampling the first human model to obtain a third human model; projecting each vertex of the third human model onto an image plane and determining a local feature vector in the projection result based on bilinear interpolation; concatenating the vertices, the local feature vectors corresponding to the vertices, and the global feature vector to obtain a mixed feature vector; determining a three-dimensional offset of each vertex based on a geometry-aware graph convolutional network and each mixed feature vector; performing up-sampling based on the three-dimensional offset of each vertex and the mixed feature vector corresponding to each vertex to obtain the second human model.

5. The method of claim 3, wherein, The method comprises the following steps: inputting the global feature vector into a pre-constructed regression structured human parameterization model to obtain at least one vertex, as well as the pose parameters, shape parameters, and camera parameters of each vertex; the parameter prediction model being composed of a predetermined number of fully connected layers; constructing the first human model of the first object according to the pose parameters, shape parameters, and camera parameters of each vertex.

6. The method of claim 3, wherein, The method comprises the following steps: inputting the fifth image into a fifty-layer residual network, and taking the output result of the last pooling layer in the fifty-layer residual network as the global feature vector.

7. The method of claim 1, wherein, The training process of the human model reconstruction model comprises the following steps: determining at least one first sample image and the second sample image corresponding to each first sample image, the first sample image being an image containing a human, and the second sample image being an image obtained by replacing the human in the first sample image with a human model; The character model reconstruction model is trained according to the at least one first sample image and a second sample image corresponding to each first sample image.

8. A person image reconstruction apparatus characterized by comprising: Comprise: A first image determination module is configured to acquire a first image in response to a first instruction, wherein the first instruction is an instruction issued by a first object and allows the first image to be acquired; and the first image is an image including the first object; A second image determination module is configured to send the first image to a pre-trained character model reconstruction model to obtain a second image; and the second image is an image in which the first object in the first image is replaced by a human body model.

9. An electronic device, comprising: The electronic device comprises: At least one processor; and A memory connected in communication with the at least one processor; wherein The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the character image reconstruction method in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer instructions for enabling the processor to implement the character image reconstruction method in any one of claims 1-7 when executed.