Image processing method and related apparatus

By transforming image features through a point cloud feature transformation network to generate the first point cloud features, the problem of poor reconstruction effect of invisible areas in existing 3D reconstruction technology is solved, and more accurate and complete 3D model reconstruction is achieved.

WO2025241724A1PCT designated stage Publication Date: 2025-11-27TENCENT TECHNOLOGY (SHENZHEN) CO LTD +1
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/086855
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-21
Filing Date
2025-04-02
Publication Date
2025-11-27

AI Technical Summary

Technical Problem

Existing 3D reconstruction techniques are poor at reconstructing object regions that are not visible in 2D images, and are prone to serious geometric deformation, singularity, or geometric blurring problems.

Method used

Image features are transformed by a point cloud feature transformation network to generate the first point cloud feature. This feature is then used for 3D model reconstruction. The point cloud feature transformation network is trained using standard point cloud features as supervision signals and sample images of sample objects as input data to learn the transformation relationship between image features and point cloud features.

Benefits of technology

It improves the reconstruction effect of 3D models, avoids geometric deformation and singularities, ensures the integrity and accuracy of 3D models, and can accurately infer and supplement information in invisible areas.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025086855_27112025_PF_FP_ABST
    Figure CN2025086855_27112025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed in the present application are an image processing method and a related apparatus. The image processing method comprises: when it is required to perform three-dimensional reconstruction on a preset object, acquiring a two-dimensional image of a preset object; performing feature extraction on the two-dimensional image to obtain first image features of the two-dimensional image; by means of a point cloud feature conversion network, performing feature conversion on the first image features to obtain corresponding first point cloud features; and then inputting the first point cloud features into a model creation module, so that the model creation module outputs a three-dimensional model of the preset object on the basis of the first point cloud features. The first point cloud features can reflect complete three-dimensional space information of the preset object and the three-dimensional geometric shape of the surface of the preset object, and provide richer content for three-dimensional reconstruction. Even if some parts are invisible, information of invisible areas can still be accurately inferred and supplemented by means of first point cloud features of other angles, thus improving the reconstruction effect of three-dimensional models, and avoiding the problems of serious geometric deformation, singularity or geometric blurring.
Need to check novelty before this filing date? Find Prior Art

Description

Image processing method and related apparatus

[0001] The present application claims priority to the Chinese patent application No. 202410635407.5, filed on May 21, 2024, and entitled "Image processing method and related apparatus", the content of which is incorporated herein by reference in its entirety. TECHNICAL FIELD

[0002] The present disclosure relates to the field of computer technology, and more particularly, to an image processing method. BACKGROUND

[0003] Three-dimensional reconstruction is a technology that obtains three-dimensional reconstruction data of an object by processing, calculating and three-dimensional restoring two-dimensional images of the object, and finally reconstructs a three-dimensional model of the object in a computer device. The three-dimensional reconstruction technology is a key technology for establishing a virtual reality expressing the real world in a computer device, and has important applications in various fields.

[0004] At present, three-dimensional reconstruction is mainly performed by a large reconstruction model (LRM) method. In the LRM method, image features can be extracted from two-dimensional images of an object, and then three-dimensional reconstruction of the object is performed based on the image features to obtain a three-dimensional model.

[0005] However, this method can obtain a relatively good three-dimensional model in the visible object region in the two-dimensional image, but the three-dimensional model of the object in the invisible object region in the two-dimensional image is often poor, and problems such as serious geometric deformation, singularity or geometric blur may occur. SUMMARY

[0006] To solve the above technical problems, the present application provides an image processing method and related apparatus, which can improve the three-dimensional reconstruction effect of the three-dimensional model and avoid problems such as serious geometric deformation, singularity or geometric blur.

[0007] The embodiments of the present application disclose the following technical solutions:

[0008] In one aspect, the present application provides an image processing method, which comprises:

[0009] obtaining a two-dimensional image of a preset object;

[0010] performing feature extraction on the two-dimensional image to obtain first image features of the two-dimensional image;

[0011] The first image feature is converted by a point cloud feature conversion network to obtain a first point cloud feature corresponding to the first image feature, the first point cloud feature is used to reflect three-dimensional space information of the preset object, the point cloud feature conversion network is obtained by training with a standard point cloud feature as a supervision signal and a sample image of a sample object as input data, the point cloud feature conversion network obtained by training has a function of converting an image feature into a point cloud feature, and the first point cloud feature obtained by conversion has the same feature dimension as the standard point cloud feature.

[0012] The first point cloud feature is input into a model creation module, and a three-dimensional model of the preset object is output by the model creation module based on the first point cloud feature.

[0013] In one aspect, an embodiment of the present application provides an image processing device, the device comprising an acquisition unit, an extraction unit, a conversion unit and a generation unit:

[0014] The acquisition unit is configured to acquire a two-dimensional image of a preset object.

[0015] The extraction unit is configured to perform feature extraction on the two-dimensional image to obtain a first image feature of the two-dimensional image.

[0016] The conversion unit is configured to convert the first image feature by a point cloud feature conversion network to obtain a corresponding first point cloud feature, the first point cloud feature is used to reflect three-dimensional space information of the preset object, the point cloud feature conversion network is obtained by training with a standard point cloud feature as a supervision signal and a sample image of a sample object as input data, the point cloud feature conversion network obtained by training has a function of converting an image feature into a point cloud feature, and the first point cloud feature obtained by conversion has the same feature dimension as the standard point cloud feature.

[0017] The generation unit is configured to input the first point cloud feature into a model creation module, and output a three-dimensional model of the preset object by the model creation module based on the first point cloud feature.

[0018] In one aspect, an embodiment of the present application provides a computer device, the computer device comprising a processor and a memory:

[0019] The memory is configured to store a computer program and transmit the computer program to the processor.

[0020] The processor is configured to execute the method according to the instructions in the computer program.

[0021] In an aspect, an embodiment of the present application provides a computer readable storage medium for storing a computer program, the computer program, when executed by a processor, causing the processor to perform the method of any one of the preceding aspects.

[0022] In an aspect, an embodiment of the present application provides a computer program product comprising a computer program, which, when executed by a processor, implements the method of any one of the preceding aspects.

[0023] As can be seen from the above technical solution, when three-dimensional reconstruction of a preset object is needed, a two-dimensional image of the preset object can be acquired, and a first image feature of the two-dimensional image can be obtained by performing feature extraction on the two-dimensional image. However, the present application does not directly perform three-dimensional reconstruction based on the image feature, but performs feature conversion on the image feature through a point cloud feature conversion network to obtain a first point cloud feature corresponding to the first image feature. The point cloud feature conversion network is trained by taking a standard point cloud feature, i.e., a real point cloud feature, as a supervision signal and taking a sample image of a sample object as input data, and in the training process, the point cloud feature conversion network learns the conversion relationship between the image feature and the point cloud feature, so that the trained point cloud feature conversion network has the function of converting the image feature into the point cloud feature. Therefore, the first point cloud feature can be accurately obtained based on the first image feature through the point cloud feature conversion network. The first point cloud feature can reflect the complete three-dimensional spatial information of the preset object and reflect the three-dimensional geometric shape of the surface of the preset object, and can understand and reconstruct the object from multiple angles and aspects, thereby providing more abundant content for three-dimensional reconstruction. Thus, when a three-dimensional model of the preset object is output by the model creation module based on the first point cloud feature, even if some parts are invisible, the information of these invisible areas can be accurately inferred and supplemented through the first point cloud feature from other angles, thereby improving the reconstruction effect of the three-dimensional model and avoiding problems such as serious geometric deformation, singularity, or geometric blur. BRIEF DESCRIPTION OF DRAWINGS

[0024] FIG. 1 is a whole architecture diagram of three-dimensional reconstruction provided by the related art;

[0025] FIG. 2 is an application scenario architecture diagram of an image processing method provided by an embodiment of the present application;

[0026] FIG. 3 is a flowchart of an image processing method provided by an embodiment of the present application;

[0027] FIG. 4 is a network structure schematic diagram of a diffusion model provided by an embodiment of the present application;

[0028] FIG. 5a is a principle example diagram of a Stable Diffusion model provided by an embodiment of the present application;

[0029] FIG. 5b is a schematic diagram of another Stable Diffusion model according to an embodiment of the present application;

[0030] FIG. 6 is a schematic diagram of obtaining a multi-view image from a single image according to an embodiment of the present application;

[0031] FIG. 7 is a schematic diagram of reconstructing a three-dimensional model from a two-dimensional image according to an embodiment of the present application;

[0032] FIG. 8 is a schematic diagram of three-plane feature conversion according to an embodiment of the present application;

[0033] FIG. 9 is a flowchart of a training method of a point cloud feature conversion network according to an embodiment of the present application;

[0034] FIG. 10 is a schematic diagram of an overall technical framework of an image processing method according to an embodiment of the present application;

[0035] FIG. 11 is a schematic diagram of a structure of an image processing apparatus according to an embodiment of the present application;

[0036] FIG. 12 is a schematic diagram of a structure of a terminal according to an embodiment of the present application;

[0037] FIG. 13 is a schematic diagram of a structure of a server according to an embodiment of the present application. DETAILED DESCRIPTION

[0038] Embodiments of the present application will be described below with reference to the accompanying drawings.

[0039] The related art can perform three-dimensional reconstruction through an LRM method. The LRM method can be seen from FIG. 1. The LRM method takes a single two-dimensional image of an object as input, and realizes a two-dimensional image-based object three-dimensional reconstruction method. The method first inputs a two-dimensional image of an object, extracts image features through an image encoder, then encodes the image features and the camera pose corresponding to the two-dimensional image through an image-to-triplane decoder to obtain corresponding three-plane features, and finally uses the three-plane features to perform three-dimensional reconstruction of the object through a neural radiance field (NeRF) method to obtain a three-dimensional model.

[0040] However, this method takes a single two-dimensional image as input to perform three-dimensional reconstruction of the object. The object region visible in the two-dimensional image can obtain a relatively good three-dimensional model. The object region invisible in the two-dimensional image often has a poor three-dimensional model, and may have problems such as severe geometric deformation, singularity, or geometric blur.

[0041] To solve the above technical problems, the embodiment of the present application provides an image processing method, which is not directly based on image features for three-dimensional reconstruction, but converts image features through a point cloud feature conversion network to obtain first point cloud features corresponding to first image features. The first point cloud features can reflect the complete three-dimensional spatial information of the preset object and reflect the three-dimensional geometric shape of the surface of the preset object, and can understand and reconstruct the object from multiple angles and aspects, providing more abundant content for three-dimensional reconstruction. Thus, when a three-dimensional model of the preset object is output by a model creation module based on the first point cloud features, even if some parts are invisible, the information of these invisible areas can be accurately inferred and supplemented through the first point cloud features of other angles, thereby improving the reconstruction effect of the three-dimensional model and avoiding problems such as severe geometric deformation, singularity or geometric blur.

[0042] It should be noted that the image processing method provided by the embodiment of the present application can be applied to various fields that need three-dimensional reconstruction, such as the game field, the artificial intelligence generated content (AIGC) field, the extended reality (XR) field, smart home, cultural relic reconstruction, audio-visual entertainment, virtual people, digital people, digital twins, smart cities, etc. In these fields, a large number of three-dimensional model assets are often needed, so that the three-dimensional model can be obtained by three-dimensional reconstruction through the method provided by the embodiment of the present application.

[0043] The image processing method provided by the embodiment of the present application can be executed by a computer device, which can be a server or a terminal, for example. The server can be a standalone physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. The terminal includes, but is not limited to, a smart phone, a computer, a smart voice interaction device, a smart home appliance, a vehicle-mounted terminal, an aircraft, an XR device, etc.

[0044] Among them, XR is the general term of virtual reality (VR), augmented reality (AR) and mixed reality (MR), which combines the visual interaction technologies of the three to bring users the "immersion" of seamless transformation between the virtual world and the real world. The XR device involved in the present application can be a virtual reality device such as a VR device, an AR device and an MR device.

[0045] As shown in FIG. 2, FIG. 2 shows an application scenario architecture diagram of an image processing method, which can include a server 200 in the application scenario.

[0046] When three-dimensional reconstruction of the preset object is needed, the server 200 can acquire a two-dimensional image of the preset object, perform feature extraction on the two-dimensional image, and obtain first image features of the two-dimensional image. However, in the embodiment of the present application, the server 200 does not directly perform three-dimensional reconstruction based on the image features, but performs feature conversion on the first image features through a point cloud feature conversion network to obtain corresponding first point cloud features, and then inputs the first point cloud features into a model creation module to output a three-dimensional model of the preset object based on the first point cloud features through the model creation module, thereby realizing three-dimensional reconstruction of the preset object.

[0047] The point cloud feature conversion network is trained by taking standard point cloud features as a supervision signal and taking sample images of sample objects as input data, so that the point cloud feature conversion network learns the conversion relationship between image features and point cloud features in the training process, and the trained point cloud feature conversion network has the function of converting image features into point cloud features. Therefore, the server 200 can accurately obtain the first point cloud features based on the first image features through the point cloud feature conversion network. The first point cloud features can reflect the complete three-dimensional spatial information of the preset object, reflect the three-dimensional geometric shape of the surface of the preset object, and can understand and reconstruct the object from multiple angles and aspects, providing more abundant content for three-dimensional reconstruction. Therefore, when the three-dimensional model of the preset object is output based on the first point cloud features through the model creation module, even if some parts are invisible, the information of these invisible areas can be accurately inferred and supplemented through the first point cloud features of other angles, thereby improving the reconstruction effect of the three-dimensional model and avoiding problems such as severe geometric deformation, singularity, or geometric blur.

[0048] It should be noted that in the specific embodiments of the present application, user information and other related data may be involved in the entire process. When the above embodiments of the present application are applied to specific products or technologies, the individual consent or individual permission of the user needs to be obtained, and the collection, use, and processing of related data need to comply with relevant laws, regulations, and standards of the country and region.

[0049] Next, the image processing method provided by the embodiments of the present application will be introduced in combination with the accompanying drawings. The method can be executed by a computer device. Referring to FIG. 3, FIG. 3 shows a flowchart of an image processing method. The method can include S301-S304, and the details are as follows:

[0050] S301, acquiring a two-dimensional image of a preset object.

[0051] Three-dimensional reconstruction is a technology of obtaining three-dimensional reconstruction data of an object by processing, calculating and three-dimensional restoring two-dimensional images of the object, and finally reconstructing a three-dimensional model of the object in a computer device, so as to display the object in the real world as a three-dimensional model in a virtual world. When three-dimensional reconstruction of a preset object is needed, a computer device can obtain a two-dimensional image of the preset object.

[0052] The preset object can refer to an object that needs to be three-dimensionally reconstructed. The object can be an object or a scene in the real world. The object can be, for example, a car, an animal, a handheld tool, etc. The scene can be, for example, a video conference scene, a game scene, etc. The present application does not limit the object.

[0053] The two-dimensional image of the preset object can be an image obtained for the preset object. For example, if the preset object is a handheld tool, the two-dimensional image of the preset object can be an image of the handheld tool. The present application does not limit the manner of obtaining the two-dimensional image of the preset object. In one possible implementation, the manner of obtaining the two-dimensional image of the preset object can be to directly obtain a single image of the preset object. The single image can be an image directly obtained by photographing the preset object, or can be a screenshot. For example, if the preset object is a virtual object (for example, a virtual character in a game), the two-dimensional image of the preset object can be obtained by taking a screenshot of a picture including the virtual object from a game application.

[0054] In some cases, a two-dimensional image can be an image of the preset object observed from one perspective, while three-dimensional reconstruction needs to understand the three-dimensional geometry of the surface of the preset object, and needs to understand the preset object from multiple angles in order to provide more comprehensive scene information for three-dimensional reconstruction. Therefore, in another possible implementation, the manner of obtaining the two-dimensional image of the preset object can be to obtain an initial image of the preset object, the initial image being a single image of the preset object, and then to generate multiple-view images of the initial image by using an image generation model to obtain two-dimensional images of the preset object corresponding to multiple perspectives respectively, so as to determine the two-dimensional images of the preset object corresponding to the multiple perspectives respectively as the two-dimensional images of the preset object.

[0055] The single image can be an image directly obtained by photographing the preset object, or can be a screenshot. For example, the preset object is a virtual object (for example, a virtual character in a game), and the two-dimensional image of the preset object can be obtained by taking a screenshot of a screen including the virtual object from a game application. The image generation model can be a model for generating two-dimensional images in multiple perspectives based on a single image, and can be obtained by training. The multi-perspective image generation can be a technology or method for generating two-dimensional images from multiple different perspectives. The perspective can be the angle at which the user views the preset object. This generation method can provide more comprehensive and rich information, which helps to better understand and analyze the preset object. In multi-perspective image generation, specific algorithms or models are usually used to simulate the observation effect in different perspectives. These algorithms or models can capture and represent the appearance, shape, structure, and other characteristics of the preset object in different perspectives, thereby generating multiple two-dimensional images with different perspectives.

[0056] In order to obtain an image generation model capable of generating two-dimensional images corresponding to the distribution in multiple perspectives, the image generation model can be trained in advance. The training process can include obtaining a training sample, the training sample including a training sample image and a real image of the training sample image in different perspectives. The to-be-trained model outputs a predicted image of the training sample image in different perspectives based on the training sample image. Then, the model parameters of the to-be-trained model are adjusted according to the difference between the predicted image and the real image in each perspective, and the image generation model is obtained. The image generation model has the ability to generate images in multiple perspectives based on one image.

[0057] It should be noted that the network structure of the image generation model is not limited in the embodiments of the present application. In a possible implementation manner, the image generation model can be a stable diffusion (Stable Diffusion) model. The diffusion model is a model for generating images, and its main feature is to generate images through gradual diffusion and iteration. The network structure of the diffusion model can be seen from FIG. 4. Generating an image through the stable diffusion model includes two processes: forward diffusion (i.e., diffusion process) and reverse diffusion. The core part of the forward diffusion is to train a denoising network (Denoising UNet), that is, to add noise to the original image (denoted as X in FIG. 4) to label the noise, input a noise image and a time step to the UNet network, let it predict the noise, and through continuous learning, reduce the loss value of the predicted noise and the actual noise, and finally obtain a UNet network capable of identifying noise. Reverse diffusion is the restoration process. After the UNet network is trained, a noise image and a time step are input to the model to iteratively remove the noise predicted by the UNet network, and the original image can be restored as much as possible. The UNet network is a convolutional neural network specially designed for image segmentation tasks, and is named after its U-shaped structure. In the embodiments of the present application, the initial image is taken as the control condition of the Stable Diffusion model, so as to generate multi-view images based on the initial image, and output the two-dimensional images (denoted as Y in FIG. 4) corresponding to the preset object under multiple views through the Stable Diffusion model.

[0058] Specifically, the principle of the Stable Diffusion model is shown in FIGS. 5a and 5b, in which the original image x0 is obtained after multiple denoising of the random noise image x3, each denoising is a UNet network processing of the Stable Diffusion model, and the noise image x3 is obtained after multiple noise addition to the original image x0, q(x2|x1) represents the probability of adding noise to x1 to obtain x2, q(x2|x1, x0) represents the probability of adding noise to x0 to obtain x1, and adding noise to x1 to obtain x2, q(x3|x2, x0) represents the probability of adding noise to x0 to obtain x2, and adding noise to x2 to obtain x3. During training, the original image x0 can be input and different degrees of noise can be superimposed to obtain xt (as shown in formula (1)), and the UNet network is allowed to attempt to denoise (i.e., restore noise). After inputting xt into the UNet network, the predicted noise is obtained, the error between the predicted noise and the noise true value is calculated, and then the loss function is obtained, so as to train the UNet network based on the loss function to obtain the correct denoising capability, i.e., as shown in formula (2). After training the UNet network, during inference, various noises xt can be generated, and after multiple UNet network denoising, the expected generated original image x0 is obtained, which realizes the generation capability of the Stable Diffusion model.

[0059] wherein x0 represents an original image, x t represents a noise image, q(x t |x0) represents a probability of generating x t by adding noise to x0, x 1:t represents x1 to x t , x 1:t-1 represents x1 to x t-1 , represents a Gaussian distribution, a t represents a hyperparameter, and I represents a (0, 1) normal distribution.

[0060] wherein L γ (ε θ ) represents a loss function, γ t and a t represent hyperparameters, T represents the number of noise addition, represents a UNet network in the Stable Diffusion model, ε t represents a noise true value, represents a predicted noise, x0 represents an original image, and x0 ~ q(x0) represents that x0 is subject to the probability distribution q(x0), represents that ε t is subject to a Gaussian distribution, Indicate the desire.

[0061] Based on the above principle, when the image generation model is a Stable Diffusion model, the embodiment of the present application can input the initial image into the Stable Diffusion model when generating a plurality of two-dimensional images respectively corresponding to different perspectives using the image generation model, and then obtain the two-dimensional images of the preset object respectively corresponding to different perspectives through the Stable Diffusion model.

[0062] By generating two-dimensional images respectively corresponding to different perspectives of the preset object, more comprehensive and rich information can be provided, which helps to better understand and analyze the preset object, so as to improve the effect of three-dimensional reconstruction based on the two-dimensional images respectively corresponding to different perspectives. At the same time, the multiple perspectives can be different, so that three-dimensional reconstruction is performed based on different two-dimensional images, providing designers with entrepreneurial inspiration. Moreover, in multi-perspective image generation, an image generation model can be used to simulate the observation effect under different perspectives. The image generation model can capture and represent the appearance, shape, structure and other characteristics of the preset object under different perspectives, thereby generating multiple two-dimensional images with different perspectives.

[0063] It should be noted that when multi-perspective image generation is performed based on an image generation model, the image generation model can output two-dimensional images under each perspective respectively. In one possible implementation, in order to ensure the consistency of the two-dimensional images respectively corresponding to different perspectives, the image generation model can also output a single image including a view combination of the two-dimensional images respectively corresponding to different perspectives. At this time, the way of generating two-dimensional images respectively corresponding to different perspectives of the preset object through multi-perspective image generation of the initial image by the image generation model can be that the image generation model is used to generate a multi-perspective image corresponding to the initial image through multi-perspective image generation of the initial image, the multi-perspective image is a single image including a view combination, the view combination is a combination of the two-dimensional images respectively corresponding to different perspectives of the preset object, and then the two-dimensional images of the preset object under each perspective of the multiple perspectives are cropped from the multi-perspective image.

[0064] The process of obtaining a single multi-view image through the image generation model can be seen from FIG. 6. FIG. 6 takes a hand-held tool as an example of the preset object. The size of the two-dimensional image of the preset object is WxH, as shown in 601 of FIG. 6. The size of the output multi-view image is mWxnH, and m*n is the number of generated multiple views, wherein m represents the number of rows of views corresponding to the multiple views, and n represents the number of columns of views corresponding to the multiple views. The obtained multi-view image can be represented as I_multi, as shown in 602 of FIG. 6. The multi-view image shown in 602 includes a combination of two-dimensional images corresponding to multiple views. After obtaining the multi-view image shown in 602, the two-dimensional image under each view can be cropped from the entire multi-view image, and the two-dimensional image under each view can be represented as I k Thus, m*n two-dimensional images corresponding to m*n views are obtained. In FIG. 6, m is 3 and n is 2, and six two-dimensional images are laid out into a single image with a 3x2 layout.

[0065] The above-mentioned multi-view image generation method considers the mutual communication between the two-dimensional images corresponding to different views, generates multiple two-dimensional images corresponding to multiple views in a two-dimensional image, thereby ensuring the consistency of the multiple two-dimensional images corresponding to multiple views in the multi-view image, and greatly improves the downstream three-dimensional reconstruction effect.

[0066] It can be understood that in some cases, multiple views can be set based on actual needs, and it can be controlled which two-dimensional images under which views are generated by the image generation model. In this case, the way to obtain the two-dimensional images corresponding to the preset object under multiple views through the multi-view image generation of the initial image by the image generation model can be to obtain a control condition, the control condition being used to indicate multiple views, and then to control the image generation model to generate the two-dimensional images corresponding to the preset object under multiple views according to the multiple views indicated by the control condition.

[0067] The control condition can be a condition for controlling how the image generation model generates two-dimensional images corresponding to multiple views based on the initial image. In the embodiments of the present application, the control condition can be multiple views. Assuming that the view corresponding to the initial image is the current view, the multiple views can be views obtained by rotating different angles in a certain direction (clockwise direction or counterclockwise direction) from the current view, for example, rotating 30 degrees, 60 degrees, 90 degrees, 120 degrees, 150 degrees, and 180 degrees in the clockwise direction in turn, thereby obtaining multiple views as the control condition to generate two-dimensional images corresponding to the multiple views.

[0068] By specifying multiple perspectives as control conditions, the image generation model can receive explicit instructions about the desired generated image perspective and generate images accordingly, i.e., the image generation model can generate corresponding two-dimensional images in the specified multiple perspectives respectively based on the control conditions, thereby helping to ensure that the generated two-dimensional images meet the needs of users or applications.

[0069] S302, feature extraction is performed on the two-dimensional image to obtain first image features of the two-dimensional image.

[0070] After obtaining the two-dimensional image of the preset object, feature extraction can be performed on the two-dimensional image to obtain image features of the two-dimensional image, i.e., first image features. Feature extraction can be a key step in computer vision and machine learning, which refers to a process of extracting meaningful, representative and computationally convenient information (i.e., "features") from raw data. These features are usually used to describe the attributes or characteristics of data for further analysis or model training. The purpose of feature extraction is to convert raw data into more representative and interpretable features to better describe and distinguish data. In the embodiments of the present application, the raw data can refer to the two-dimensional image of the preset object. The first image features can be representative and interpretable features extracted from the two-dimensional image of the preset object, which are used for subsequent steps of three-dimensional reconstruction.

[0071] In the embodiments of the present application, feature extraction can be performed by an image encoder. The image encoder used in the embodiments of the present application can be obtained by training, and the network structure of the image encoder is not limited in the embodiments of the present application. The image encoder can be any encoder capable of feature extraction, such as a contrastive language-image pre-training (CLIP) model and a Dino feature model. The CLIP model is a pre-trained model, which is used to extract features of matching images and texts. For a two-dimensional image and a corresponding text description, the extracted image features are usually one-dimensional feature vectors, and the text features are also one-dimensional vectors with dimensions of 512, 768 or 1024, etc. The present application can use the CLIP model to extract features from the two-dimensional image to obtain one-dimensional feature vectors as image features. The Dino feature model is an image feature extraction model, which generally uses a self-supervised vision transformer (ViT) to extract image features, has good generalization, and can achieve high accuracy in multiple computer vision tasks.

[0072] The expression of feature extraction by the image encoder can be as follows: f_image = Image_encoder(I)

[0073] Wherein, f_image represents the first image feature of the two-dimensional image, Image_encoder represents the corresponding image encoder, and I represents the two-dimensional image of the preset object.

[0074] It can be understood that when the two-dimensional image of the preset object is a two-dimensional image corresponding to each view angle respectively, the way of performing feature extraction on the two-dimensional image to obtain the first image feature of the two-dimensional image can be to perform feature extraction on the two-dimensional image under each view angle respectively to obtain the first sub-image feature of the two-dimensional image under each view angle. The first sub-image feature of the two-dimensional image under each view angle constitutes the first image feature. The first sub-image feature of the two-dimensional image under each view angle can be expressed as f image_k , k is 1, 2, 3, …, the number of view angles, for example, the number of view angles is m x n, then k is 1, 2, 3, …, m x n.

[0075] When feature extraction is performed by the image encoder, the expression of performing feature extraction on the two-dimensional image under each view angle respectively can be as follows: f_image_k = Image_encoder(I_k)

[0076] Wherein, f_image_k represents the first sub-image feature of the two-dimensional image under the kth view angle, Image_encoder represents the corresponding image encoder, and I_k represents the two-dimensional image under the kth view angle.

[0077] S303, performing feature conversion on the first image feature by a point cloud feature conversion network to obtain a corresponding first point cloud feature.

[0078] After obtaining the first image feature, the first image feature can be input into the point cloud feature conversion network, so as to convert the first image feature into corresponding first point cloud feature through the point cloud feature conversion network. The point cloud feature conversion network is obtained by training with a standard point cloud feature as a supervision signal and a sample image of a sample object as input data. The standard point cloud feature is a real and accurate point cloud feature, which can be obtained by directly extracting features based on point cloud data, or by extracting features through an accurate and reliable feature extraction method, or by artificial labeling, etc. The point cloud feature conversion network is trained with the real and accurate point cloud feature (i.e., the standard point cloud feature), and in the training process, the point cloud feature conversion network learns the conversion relationship between the image feature and the point cloud feature, so that the point cloud feature conversion network has the ability to convert the image feature into the point cloud feature. Therefore, the first point cloud feature can be accurately obtained based on the first image feature through the point cloud feature conversion network, and the first point cloud feature obtained by feature conversion has the same feature dimension as the standard point cloud feature. For example, the standard point cloud feature includes features of four dimensions A, B, C and D, and the converted first point cloud feature should also include features of four dimensions A, B, C and D. In the embodiments of the present application, the same feature dimension can mean that the feature size and the channel number of the converted first point cloud feature are respectively the same as those of the standard point cloud feature. The first point cloud feature is used to represent the three-dimensional spatial information of the preset object, can represent the complete three-dimensional spatial information of the preset object, reflect the three-dimensional geometric shape of the surface of the preset object, and can understand the preset object from multiple angles and aspects, to provide more abundant content for three-dimensional reconstruction.

[0079] The point cloud feature conversion network is a deep learning model and can be obtained by training. The point cloud feature conversion network aims to convert image features in an image into point cloud features to realize feature alignment between the image and the point cloud data. That is, through feature alignment, the image features (such as edges, corners, textures, etc.) detected in the image can be one-to-one corresponding to the point cloud features, and this corresponding relationship can better understand the structure and layout of the preset object in the three-dimensional space. In the embodiments of the present application, the image feature processed by the point cloud feature conversion network is the first image feature in a two-dimensional image, and the converted point cloud feature is the first point cloud feature, which is used for subsequent three-dimensional reconstruction.

[0080] The network structure of the point cloud feature conversion network is not limited in the embodiments of the present application, which can be a network structure based on Transformer. The Transformer can be a neural network model based on a self-attention mechanism.

[0081] When the image features are the first sub-image features of the two-dimensional images under each view angle, the manner of performing feature conversion on the first image features by the point cloud feature conversion network to obtain the first point cloud features can be performing feature conversion on the first sub-image features of the two-dimensional images under each view angle by the point cloud feature conversion network to obtain the first sub-point cloud features of the two-dimensional images under each view angle, and then summing the first point sub-cloud features of the two-dimensional images under each view angle to obtain the first point cloud features.

[0082] Taking the first image features represented by f_image_k as an example, the expression of converting the first sub-image features of the two-dimensional images under each view angle into the corresponding first point sub-cloud features by the point cloud feature conversion network is as follows: f_trans_k = Transformer(f_image_k)

[0083] Wherein, f_trans_k represents the converted first sub-point cloud features under the kth view angle, f_image_k represents the first sub-image features of the two-dimensional images under the kth view angle, and Transformer() represents the point cloud feature conversion network based on the Transformer network structure.

[0084] The first point cloud features obtained by feature conversion can be represented as:

[0085] Wherein, f_trans_all represents the first point cloud features obtained by feature conversion, f_trans_k represents the converted first sub-point cloud features under the kth view angle, and m x n represents the number of multiple view angles.

[0086] In the above manner, the respective first sub-point cloud features corresponding to multiple view angles can be captured, thereby providing more comprehensive and rich information for analyzing the spatial structure of the preset object, so as to improve the effect of three-dimensional reconstruction.

[0087] S304, input the first point cloud features into the model creation module, and output the three-dimensional model of the preset object based on the first point cloud features by the model creation module.

[0088] After the first point cloud feature is obtained through feature conversion, a three-dimensional model of the preset object can be obtained through three-dimensional reconstruction based on the obtained first point cloud feature. To this end, the obtained first point cloud feature can be input into a model creation module, so that the model creation module outputs a three-dimensional model of the preset object based on the first point cloud feature. The three-dimensional model is a model of the preset object created in a virtual three-dimensional space. The first point cloud feature can embody the complete three-dimensional space information of the preset object, reflect the three-dimensional geometric shape of the surface of the preset object, and can understand and reconstruct the object from multiple angles and aspects, providing more abundant content for three-dimensional reconstruction. Even if some parts are occluded or invisible, the information of these invisible areas can be accurately inferred and supplemented through the first point cloud feature of other angles, thereby improving the reconstruction effect of the three-dimensional model and avoiding serious geometric deformation, singularity or geometric blur and other problems.

[0089] The three-dimensional model obtained through three-dimensional reconstruction can be referred to as a three-dimensional network (mesh). In computer graphics and three-dimensional modeling, "mesh" is a commonly used term, which refers to a digital representation of a three-dimensional preset object composed of vertices, edges and faces (usually triangles or quadrilaterals). This representation method allows complex three-dimensional shapes to be stored, processed and rendered in computer devices. Based on the first point cloud feature, the three-dimensional reconstruction can first obtain a surface model of the preset object based on the first point cloud feature through a surface reconstruction algorithm, and then convert the surface model into a mesh model. The surface reconstruction algorithm can be Poisson surface reconstruction, moving least squares (MLS), triangulation, etc., and the mesh can be a triangular mesh or a quadrilateral mesh, etc.

[0090] Referring to FIG. 7, FIG. 7 takes a handheld tool as an example of a preset object. By photographing the handheld tool, a two-dimensional image of the preset object can be obtained. In FIG. 7, 701 shows the two-dimensional image of the preset object. After three-dimensional reconstruction through the steps of S301-S304, a three-dimensional model of the handheld tool can be obtained. The effect diagram of the obtained three-dimensional model can be seen in FIG. 7, 702. 702 is only an example of a three-dimensional model in a three-dimensional virtual space, and is a three-dimensional model with good effect obtained based on the two-dimensional image shown in 701. Due to the limitation of drawing, the three-dimensional effect of the three-dimensional model is limited.

[0091] The model creation module can generate a corresponding three-dimensional model based on the point cloud features, and is trained. The model architecture of the model creation module is not specifically limited in the embodiments of the present application. The model creation module can include a trained planar feature conversion network and a trained generation network, etc. The model creation module can be trained based on the initial model creation module based on the second predicted point cloud features. Please refer to the subsequent description, which will not be repeated here. Other training methods can also be used, which are not specifically limited in the present application.

[0092] It should be noted that the main surface and edge of the preset object can more accurately reflect the geometric shape and structure of the preset object, and these information is crucial for three-dimensional reconstruction, which can help to construct a more accurate three-dimensional model. The surface and edge of the preset object can be represented by the triplane feature, therefore, in a possible implementation manner, the model creation module can include a trained planar feature conversion network and a trained generation network, and the manner of inputting the first point cloud feature into the model creation module to output the three-dimensional model of the preset object based on the first point cloud feature in S304 can be inputting the first point cloud feature into the planar feature conversion network, converting the first point cloud feature into triplane features in a three-dimensional space formed by three orthogonal planes through the planar feature conversion network, the triplane features representing the voxels to which the points corresponding to the first point cloud feature belong in the three-dimensional space, and then outputting the three-dimensional model of the preset object based on the triplane features through the generation network.

[0093] Triplane feature can refer to a feature representation method in three-dimensional computer vision and image processing, which uses information on three orthogonal planes (usually X-Y, X-Z and Y-Z planes, see FIG. 8) to describe points or regions in three-dimensional space. This feature representation method plays an important role in processing point cloud data, three-dimensional reconstruction, surface analysis and rendering, etc. Voxel can be the smallest unit of three-dimensional space division, which is equivalent to a pixel in three-dimensional space. The three-dimensional space here can be a three-dimensional space formed by the three orthogonal planes, and each voxel is assigned a triplane feature through triplane feature conversion, thereby converting the first point cloud feature into triplane features. For example, in FIG. 8, the first point cloud feature is converted into the triplane feature of the voxel where the black dot is located in FIG. 8.

[0094] The conversion of the first point cloud feature into the tri-plane feature can be implemented by a plane feature conversion network. The plane feature conversion network can be obtained through training. The plane feature conversion network can be a multi-layer deep learning convolutional layer or a transformer module, and the embodiments of the present application do not limit the same. The expression for converting the first point cloud feature into the tri-plane feature can be as follows: f pf1 = Plane feature (f trans all)

[0095] wherein f pf1 represents the converted tri-plane feature, Plane feature represents a plane feature conversion network, and f trans all represents the first point cloud feature.

[0096] By converting the first point cloud feature into the tri-plane feature, the geometry and structure of the preset object can be more accurately described, so as to construct a more accurate three-dimensional model. At the same time, the tri-plane feature can organize the first point cloud feature into a more easily processed structure. By extracting and representing the tri-plane feature, the preset object in the three-dimensional scene can be more concisely described, data redundancy is reduced, and the subsequent three-dimensional reconstruction and analysis process is simplified. The amount of data to be processed is significantly reduced, thereby reducing the computational complexity and improving the processing efficiency. The three-dimensional reconstruction process is faster and more efficient, and the asset production cost is reduced.

[0097] After obtaining the three-plane features, the generation network can perform three-dimensional reconstruction on the preset object in different ways, such as a NeRF way, a sign distance function (SDF) way, and the like. Embodiments of the present application mainly introduce the SDF way. The SDF can also be referred to as an oriented distance function. The SDF is a continuous function that maps a point p = (x, y, z) in a three-dimensional space to a real number s, which can be represented as s = SDF(p). The positive and negative of s represent whether the point is inside or outside the surface of the preset object. The absolute value of s represents the distance from the point to the surface of the preset object. The point is inside the surface of the preset object, and the point is outside the surface of the preset object. The point is on the surface of the preset object. Based on the principle of the SDF way, in one possible implementation, the generation network can output a three-dimensional model of the preset object based on the three-plane features in the following way. The generation network calculates a first oriented distance function value between the vertices of the voxels to which the three-plane features belong and the surface of the preset object, then extracts the vertices with the first oriented distance function value of zero through the generation network, and generates a three-dimensional model using the vertices with the first oriented distance function value of zero. The generation network can be an SDF network, which can be obtained through training. The first oriented distance function value is the distance between the vertices of the voxels to which the three-plane features belong and the surface of the preset object, which is calculated through the SDF way or the like.

[0098] When the vertices with the first oriented distance function value of zero are extracted and the three-dimensional model is generated using the vertices with the first oriented distance function value of zero, the corresponding three-dimensional model can be calculated through a Marching Cube (MC) method or the like. The Marching Cube way is a classic algorithm in surface rendering algorithms. The Marching Cube algorithm is used to process the first oriented distance function value. Specifically, the first oriented distance function value can be divided into regular cubic units (voxels), and these voxels are processed one by one. In each voxel, the surface condition inside the voxel is determined according to the relationship between the first oriented distance function value and a certain threshold. Next, surface triangles are generated according to the surface condition, and the positions of the triangle vertices are calculated. This usually involves interpolation calculation to determine the intersection points of the isosurface and the edges of the cubic unit. Finally, the generated triangles are added to the final three-dimensional model, thereby obtaining a three-dimensional model generated by the first oriented distance function value.

[0099] The SDF is a continuous function capable of representing the distance of a given point (e.g., a vertex) to the surface of a preset object and whether the vertex is inside (negative) or outside (positive) the preset object, thereby enabling the SDF to learn a full-continuous shape function of arbitrary precision, thereby achieving high-precision three-dimensional reconstruction.

[0100] As can be seen from the above technical solution, when three-dimensional reconstruction of a preset object is needed, a two-dimensional image of the preset object can be acquired, and a first image feature of the two-dimensional image can be obtained by performing feature extraction on the two-dimensional image. However, the present application does not directly perform three-dimensional reconstruction based on the image feature, but performs feature conversion on the image feature through a point cloud feature conversion network to obtain a first point cloud feature corresponding to the first image feature. The point cloud feature conversion network is trained by taking a standard point cloud feature, i.e., a real point cloud feature, as a supervision signal and taking a sample image of a sample object as input data, and in the training process, the point cloud feature conversion network learns the conversion relationship between the image feature and the point cloud feature, so that the trained point cloud feature conversion network has the function of converting the image feature into the point cloud feature. Therefore, the first point cloud feature can be accurately obtained based on the first image feature through the point cloud feature conversion network. The first point cloud feature can reflect the complete three-dimensional spatial information of the preset object and reflect the three-dimensional geometric shape of the surface of the preset object, and can understand and reconstruct the object from multiple angles and aspects, thereby providing more abundant content for three-dimensional reconstruction. Thus, when a three-dimensional model of the preset object is output based on the first point cloud feature by the model creation module, even if some parts are not visible, the information of these invisible areas can be accurately inferred and supplemented through the first point cloud feature from other angles, thereby improving the reconstruction effect of the three-dimensional model and avoiding problems such as severe geometric distortion, singularity, or geometric blur.

[0101] Through the introduction of the foregoing embodiments, the key to improving the reconstruction effect in the embodiments of the present application is that after the two-dimensional image of the preset object is acquired, three-dimensional reconstruction is not directly performed based on the image feature of the two-dimensional image, but the image feature is converted into the first point cloud feature through the point cloud feature conversion network, and then three-dimensional reconstruction is performed based on the converted first point cloud feature. Therefore, the accuracy of the first point cloud feature can affect the reconstruction effect, and the accuracy of the first point cloud feature depends on the performance of the point cloud feature conversion network, which is determined by the model training process. Next, the training method of the point cloud feature conversion network will be introduced to obtain a point cloud feature conversion network with better performance.

[0102] Referring to FIG. 9, FIG. 9 shows a flowchart of a training method of a point cloud feature conversion network, which includes S901-S904, and is specifically as follows:

[0103] S901, acquiring a sample image of a sample object, and acquiring a standard point cloud feature.

[0104] To train the point cloud feature conversion network, a training sample for training the point cloud feature conversion network can be obtained, which can include a sample image of a sample object and a standard point cloud feature, the standard point cloud feature being a true value in the training process and can be used as a supervision signal for training the point cloud feature conversion network.

[0105] The sample object can refer to an object with a standard three-dimensional model in the training process. The object can be an object, a scene, etc. in the real world or the virtual world. The object can be a car, an animal, a handheld tool, etc. in the real world. The virtual object (e.g., a game character in a game, etc.) in the virtual world. The scene can be a video conference scene, etc. The present application does not limit this. The standard three-dimensional model can be a known and accurate three-dimensional model. The standard three-dimensional model can be obtained by three-dimensional reconstruction through an accurate and reliable three-dimensional reconstruction method, or obtained by manual annotation, etc.

[0106] The sample image of the sample object can be an image obtained for the sample object. For example, the sample object is a handheld tool, and the sample image of the sample object can be an image of the handheld tool. The present application does not limit the manner of obtaining the sample image of the sample object. In one possible implementation, the manner of obtaining the sample image of the sample object can be to directly obtain an image by photographing the sample object.

[0107] In some cases, one sample image can be an image of the sample object observed from one perspective, while three-dimensional reconstruction needs to understand the three-dimensional geometry of the surface of the sample object, and needs to understand the sample object from multiple angles in order to provide more comprehensive scene information for three-dimensional reconstruction. Therefore, in another possible implementation, the manner of obtaining the sample image of the sample object can be to obtain an initial sample image of the sample object, the initial sample image being a single image of the sample object, and then to generate multiple-view images of the initial sample image by using an image generation model to obtain sample two-dimensional images of the sample object corresponding to multiple perspectives respectively, so as to determine the sample two-dimensional images of the sample object corresponding to multiple perspectives respectively as the sample images of the sample object. The initial sample image can be a single image obtained by photographing the sample object, or can be obtained by screenshot, for example, the sample object is a virtual object (e.g., a virtual character in a game), and the initial sample image can be obtained by screenshot from a game application.

[0108] The image generation model used in the training process can refer to the introduction of the corresponding embodiment of FIG. 3, and the manner of generating the sample two-dimensional images of the sample object corresponding to multiple perspectives respectively by using the image generation model can also refer to the introduction of the corresponding embodiment of FIG. 3, which will not be described here.

[0109] It should be noted that when the multi-view image generation module generates the multi-view image, the image generation model can output a sample two-dimensional image under each view. In a possible implementation, in order to ensure the consistency of the sample two-dimensional images corresponding to the multiple views, the image generation model can also output a single sample image including a sample view combination, the sample view combination being a combination of the sample two-dimensional images corresponding to the multiple views of the sample object. At this time, the manner of generating the sample two-dimensional images corresponding to the multiple views of the sample object by the image generation model on the initial sample image can be that the image generation model generates a multi-view sample image corresponding to the initial sample image on the initial sample image, the multi-view sample image including the single sample image obtained by the sample view combination, and then the sample two-dimensional images of the sample object under each of the multiple views are cropped from the multi-view sample image. The process can be specifically referred to the introduction of the corresponding embodiment of FIG. 3, and will not be described here.

[0110] The standard point cloud feature is used to reflect the three-dimensional spatial information of the sample object and is true and accurate. In a possible implementation, the second point cloud feature generated by the trained point cloud encoder can be determined as the standard point cloud feature. The trained point cloud encoder is trained based on the point cloud data, and the trained point cloud encoder is a model capable of accurately extracting a point cloud feature. The second point cloud feature generated by the trained point cloud encoder is obtained by the trained point cloud encoder performing feature extraction on the point cloud data of the sample object. Therefore, the accurate second point cloud feature is generated by the trained point cloud encoder, and the second point cloud feature is determined as the standard point cloud feature, so as to train by using the accurate standard point cloud feature and obtain a point cloud feature conversion network with better performance.

[0111] The point cloud encoder (Pointclouds encode) can be a common neural network structure for extracting a point cloud feature. The point cloud encoder can be any network for extracting a feature from point cloud data. Common point cloud encoders include a point cloud network (pointnet), a point bidirectional encoder representation from transformers (pointbert), a point multilayer perceptron (pointmlp), pointnext, and the like. The pointnext is the next version of the pointnet. Such a neural network structure usually takes point cloud data as input and outputs a feature vector corresponding to the point cloud data, i.e., the second point cloud feature.

[0112] In one possible implementation, the expression of the second point cloud feature generated by the point cloud encoder can be as follows: f_point = Encoder_3d(point)

[0113] wherein f_point represents the second point cloud feature, Encoder_3d represents the point cloud encoder, and point represents the point cloud data. The standard point cloud feature can be represented by f_point.

[0114] It can be understood that the trained point cloud encoder is trained based on the point cloud data, and the trained point cloud encoder can be obtained in the process of training the three-dimensional generation network model based on the point cloud data of the sample object. The trained three-dimensional generation network model includes the trained point cloud encoder and the model creation module trained based on the initial creation module. The process of training the three-dimensional generation network model can be that the point cloud data of the sample object is obtained, and then the initial point cloud encoder is used to extract features from the point cloud data to obtain a second predicted point cloud feature. The second predicted point cloud feature can be a point cloud feature extracted based on the point cloud data of the sample object. Then the second predicted point cloud feature is input into the initial creation module, and the initial creation module outputs a predicted three-dimensional model of the sample object based on the second predicted point cloud feature, so as to adjust the initial point cloud encoder and the initial creation module based on the difference between the predicted three-dimensional model and the standard three-dimensional model of the sample object, and train the three-dimensional generation network model. The three-dimensional generation network model includes the trained point cloud encoder and the model creation module trained based on the initial creation module, and the trained point cloud encoder is trained based on the initial point cloud encoder, so that the trained point cloud encoder is trained.

[0115] The point cloud data of the sample object can be a discrete data set composed of a series of three-dimensional coordinate points, which is used to describe the geometric characteristics of the surface shape, spatial position, size, etc. of the object. The main characteristics of point cloud data are high precision, high resolution and high dimensional geometric information, which can directly represent the shape, surface and texture of the object in space. The generation of point cloud data is mainly through the following ways:

[0116] 1. Laser scanner: Laser scanner is a common method for obtaining point cloud data. By using the principle of laser ranging, the surface of the object (such as the surface of the sample object) is scanned to obtain a large number of coordinate points on the surface of the object, forming point cloud data.

[0117] 2. Structure light scanner: The structure light scanner projects a known light pattern onto the surface of the object, and calculates the three-dimensional coordinates of the object based on the light and shadow information obtained by the camera.

[0118] 3. Depth camera: A depth camera is a method that can directly obtain the depth information of the surface of an object. By using optical principles and combining the reflection information of the surface of the object, the depth information of the object is directly obtained, and point cloud data is generated.

[0119] 4. Robot vision navigation: A robot obtains surrounding environment images through a vision sensor, and uses a computer vision algorithm to perform image feature extraction and matching to generate point cloud data for positioning and navigation of the robot.

[0120] The feature extraction here is similar to the feature extraction performed for a two-dimensional image, except that it is performed on the point cloud data of the sample object, which will not be described here. The extracted point cloud features (such as the second predicted point cloud features) can be used to describe the structure and shape of the point cloud, thereby reducing redundant information.

[0121] The initial creation module can be a model that performs three-dimensional reconstruction based on the second predicted point cloud features. The model structure of the initial creation module is not specifically limited in the embodiments of the present application, and can be referred to the related introduction of the corresponding embodiments of FIG. 3. The predicted three-dimensional model can be a three-dimensional model output by a three-dimensional generation network model that needs to be trained during the training of the three-dimensional generation network model. The standard three-dimensional model is a real and accurate three-dimensional model of the sample object, which can be used to supervise the training of the three-dimensional generation network model. Therefore, after obtaining the predicted three-dimensional model, the initial point cloud encoder and the initial creation module can be adjusted based on the difference between the predicted three-dimensional model and the standard three-dimensional model, and the three-dimensional generation network model can be trained. For example, a reconstruction loss function can be constructed based on the predicted three-dimensional model and the standard three-dimensional model, and the initial point cloud encoder and the initial creation module can be adjusted using the reconstruction loss function.

[0122] The above method trains the three-dimensional generation network model through the point cloud data of the sample object. Since the point cloud data of the sample object can reflect the complete three-dimensional spatial information of the sample object, it can understand the sample object from multiple angles and aspects, providing more abundant content for three-dimensional reconstruction, and thus training a more accurate three-dimensional generation network model (including a completed point cloud encoder), and further ensuring that accurate standard point cloud features are obtained.

[0123] It should be noted that the main surface and edge of the sample object can more accurately reflect the geometric shape and structure of the sample object, and these information is crucial for three-dimensional reconstruction, which can help to construct a more accurate predicted three-dimensional model. The surface and edge of the sample object can be reflected by the predicted three-plane feature. Therefore, in a possible implementation, the initial creation module further includes a first initial network and a second initial network. The manner in which the initial creation module outputs the predicted three-dimensional model of the sample object based on the second predicted point cloud feature can be that the first initial network is used to perform three-plane feature conversion on the second predicted point cloud feature in a three-dimensional space formed by three orthogonal planes to obtain predicted three-plane features, the predicted three-plane features representing the voxels to which the points corresponding to the second predicted point cloud feature belong in the three-dimensional space. Then, the predicted three-plane features are input into the second initial network, and the second initial network is used to output the predicted three-dimensional model based on the predicted three-plane features. The first initial network is a basic model framework used for training a plane feature conversion network, which can be any neural network model structure, and the embodiments of the present application do not limit the same. The second initial network is a basic model framework used for training a generation network, which can be any neural network model structure, and the embodiments of the present application do not limit the same.

[0124] The manner in which the three-plane feature conversion is performed in the training process is similar to the manner in which the first point cloud feature is converted into the three-plane feature in the embodiment corresponding to FIG. 3, which will not be described herein again. The expression in which the second predicted point cloud feature is converted into the three-plane feature in the training process can be as follows: f_pf2=Plane_feature(f_point)

[0125] Wherein, f_pf2 represents the predicted three-plane feature, Plane_feature represents a plane feature conversion network, and f_point represents the second predicted point cloud feature obtained by the initial point cloud encoder by performing feature extraction on the point cloud data.

[0126] In the case where the three-dimensional generation network model further includes the first initial network and the second initial network, the manner in which the initial point cloud encoder and the initial creation module are adjusted and the three-dimensional generation network model is trained based on the difference between the predicted three-dimensional model and the standard three-dimensional model of the sample object can be that the initial point cloud encoder, the first initial network and the second initial network are adjusted and the three-dimensional generation network model is trained based on the difference between the predicted three-dimensional model and the standard three-dimensional model of the sample object. The three-dimensional generation network model includes the trained point cloud encoder, the plane feature conversion network trained based on the first initial network and the generation network trained based on the second initial network, and the trained point cloud encoder is trained based on the initial point cloud encoder.

[0127] By converting the second predicted point cloud features into predicted triplane features, the geometry and structure of the sample object can be more accurately described, so as to construct a more accurate predicted three-dimensional model. At the same time, the predicted triplane features can organize the second predicted point cloud features into a more easily processed structure. By extracting and representing the predicted triplane features, the sample object in the three-dimensional scene can be more succinctly described, data redundancy can be reduced, and subsequent three-dimensional reconstruction and analysis processes can be simplified, significantly reducing the amount of data that needs to be processed, thereby reducing the computational complexity and improving the processing efficiency, so that the three-dimensional reconstruction process is faster and more efficient, and the asset production cost is reduced.

[0128] After obtaining the predicted triplane features, the preset object can be three-dimensionally reconstructed in different ways, such as a NeRF way, a sign distance function (SDF) way, etc. The present embodiment mainly introduces the SDF way. In this case, the way in which the second initial network outputs a predicted three-dimensional model based on the predicted triplane features can be that the second initial network calculates a second directed distance function value between the vertices of the voxels to which the predicted triplane features belong and the surface of the sample object, and then extracts the vertices for which the second directed distance function value is zero by the second initial network, and generates a predicted three-dimensional model using the vertices for which the second directed distance function value is zero. The specific implementation of the three-dimensional reconstruction in the SDF way during the training process can be referred to the introduction of the corresponding embodiment of FIG. 3, which will not be described here. The second directed distance function value is the distance between the vertices of the voxels to which the predicted triplane features belong and the surface of the preset object calculated by the SDF way.

[0129] The SDF is a continuous function that can represent the distance of a given point (such as a vertex) to the surface of the sample object, and whether the vertex is inside (negative) or outside (positive) the sample object, thereby enabling the SDF to learn a full-continuous shape function of arbitrary precision, and thereby enabling a generative network capable of high-precision three-dimensional reconstruction to be trained.

[0130] S902, feature extraction is performed on the sample image to obtain sample image features of the sample image.

[0131] After obtaining the sample image, feature extraction can be performed on the sample image to obtain sample image features of the sample image. The implementation of feature extraction on the sample image can refer to the feature extraction method introduced in S302, except that the feature extraction in S902 is performed on the sample image instead of a two-dimensional image, and the obtained result is sample image features instead of image features.

[0132] S903, feature conversion is performed on the sample image features by an initial network model to obtain first predicted point cloud features.

[0133] The initial network model can be a basic model framework used for training the point cloud feature conversion network. The initial network model is used to perform feature conversion on the sample image features to obtain first predicted point cloud features. The first predicted point cloud features can be point cloud features obtained by performing feature conversion on the sample image features, and are prediction values obtained in the training process. The first predicted point cloud features can be denoted as f_trans_all.

[0134] In S904, the initial network model is trained based on the difference between the first predicted point cloud features and the standard point cloud features to obtain the point cloud feature conversion network.

[0135] The first predicted point cloud features are prediction values in the training process, and the standard point cloud features are true values in the training process. Therefore, in the process of training the point cloud feature conversion network, the initial network model can be trained based on the difference between the first predicted point cloud features and the standard point cloud features to obtain the point cloud feature conversion network.

[0136] In the training process, the extracted sample image features and the standard point cloud features can be aligned. This usually involves steps such as feature matching and transformation matrix calculation. Through a feature matching algorithm, corresponding feature points in the sample image and the point cloud data of the sample object are found, and the transformation relationship between them is calculated. Then, according to these transformation relationships, the sample image features are converted to the same coordinate system as the standard point cloud features to achieve feature alignment.

[0137] Then, the aligned standard point cloud features are used as a supervision signal to train the initial network model. In the training process, the initial network model learns how to convert the input image features into corresponding point cloud features, so as to obtain a point cloud feature conversion network that can accurately convert image features into point cloud features.

[0138] The embodiments of the present application use the standard point cloud features obtained based on the point cloud data as a supervision signal. The standard point cloud features are real and accurate point cloud features. Using the standard point cloud features as a supervision signal can enable the initial network model to continuously learn how to convert image features into accurate point cloud features, so as to train a point cloud feature conversion network that can accurately convert image features into point cloud features.

[0139] In training the point cloud feature conversion network based on the difference between the first predicted point cloud features and the standard point cloud features, a target loss function can be constructed based on the difference between the first predicted point cloud features and the standard point cloud features, and then the initial network model is trained based on the target loss function to obtain the point cloud feature conversion network. In a possible implementation manner, the calculation formula of the target loss function can be as follows: Loss l1 = ||f_trans_all-f_point||

[0140] wherein, Loss l1 represents a target loss function, f_trans_all represents the first predicted point cloud feature, and f_point represents a standard point cloud feature output by the point cloud encoder based on the sample reconstruction.

[0141] Through the above Loss l1 The f_trans_all is effectively optimized, so that the point cloud feature conversion network learns the conversion relationship between the image feature and the point cloud feature, thereby training the point cloud feature conversion network to obtain a better effect, so as to achieve a better reconstruction effect in three-dimensional reconstruction.

[0142] It should be noted that the image generation model, the image encoder, the point cloud feature conversion network, the plane feature conversion network, and the SDF network involved in the embodiments of the present application can be trained together, thereby improving the performance of each model in the three-dimensional reconstruction process, and further improving the reconstruction effect.

[0143] Based on the three-dimensional reconstruction process of the embodiment corresponding to the foregoing Fig. 3, and the training process of the point cloud feature conversion network of the embodiment corresponding to Fig. 9, the overall technical framework of the image processing method provided by the embodiments of the present application can be referred to Fig. 10, the entire technical solution includes two processes, the first process can be three-dimensional reconstruction based on point cloud data of a sample object, and the second process can be image feature-point cloud feature alignment through a point cloud feature conversion network, that is, converting the first image feature into the first point cloud feature.

[0144] The first process can be shown in Fig. 10 as 1001, in which the point cloud data of the sample object is obtained, and then the point cloud data is feature extracted through the initial point cloud encoder to obtain the point cloud feature (the point cloud feature obtained in the first process can be the second predicted point cloud feature). Then the second predicted point cloud feature is converted into a predicted three-plane feature through the first initial network. Then, based on the predicted three-plane feature, the sample object is three-dimensionally reconstructed through the second initial network to obtain a predicted three-dimensional model, so as to adjust the initial point cloud encoder, the first initial network and the second initial network based on the difference between the predicted three-dimensional model and the standard three-dimensional model of the sample object, and train a three-dimensional generation network model. The three-dimensional generation network model includes a trained point cloud encoder, a plane feature conversion network trained based on the first initial network, and a generation network trained based on the second initial network. After obtaining the trained point cloud encoder, the second point cloud feature obtained by the trained point cloud encoder can be used as a standard point cloud feature, and the standard point cloud feature can be used as a supervision signal to train a point cloud feature conversion network.

[0145] The process of training the point cloud feature conversion network can be that a sample image of a sample object is subjected to feature extraction to obtain a sample image feature of the sample image, the sample image feature is subjected to feature conversion by an initial network model to obtain a first predicted point cloud feature, and the initial network model is trained based on a difference between the first predicted point cloud feature and a standard point cloud feature to obtain the point cloud feature conversion network.

[0146] The second process can be as shown in 1002 in FIG. 10, in which only an initial image of a preset object needs to be input, and a plurality of images are generated by the image generation model to obtain a plurality of two-dimensional images of different perspectives corresponding to the initial image; for the generated multi-perspective image, the multi-perspective image is a single image obtained by combining views corresponding to a plurality of perspectives, and then a two-dimensional image of the preset object at each perspective in the plurality of perspectives is cropped from the multi-perspective image. Then, the two-dimensional image at each perspective is subjected to feature extraction by the image encoder to obtain a first image feature, and the first image feature is converted into a point cloud feature (the point cloud feature obtained in the second process can be a first point cloud feature) by using the trained point cloud feature conversion network. The first point cloud feature is subjected to three-plane feature conversion to obtain three-plane features, and the preset object is subjected to three-dimensional reconstruction in an SDF manner based on the three-plane features to obtain a three-dimensional model of the preset object.

[0147] It should be noted that the application can be further combined to provide more implementation manners on the basis of the implementation manners provided in the above aspects.

[0148] Based on the image processing method provided in the foregoing embodiments, an embodiment of the application further provides an image processing apparatus 1100. Referring to FIG. 11, the image processing apparatus 1100 includes an acquisition unit 1101, an extraction unit 1102, a conversion unit 1103, and a generation unit 1104:

[0149] The acquisition unit 1101 is configured to acquire a two-dimensional image of a preset object.

[0150] The extraction unit 1102 is configured to perform feature extraction on the two-dimensional image to obtain a first image feature of the two-dimensional image.

[0151] The conversion unit 1103 is configured to perform feature conversion on the first image feature by using a point cloud feature conversion network to obtain a corresponding first point cloud feature, the first point cloud feature is used to represent three-dimensional spatial information of the preset object, the point cloud feature conversion network is obtained by training with a standard point cloud feature as a supervision signal and a sample image of a sample object as input, the trained point cloud feature conversion network has a function of converting an image feature into a point cloud feature, and the first point cloud feature obtained by feature conversion has the same feature dimension as the standard point cloud feature.

[0152] The generation unit 1104 is configured to input the first point cloud feature into a model creation module, and output a three-dimensional model of the preset object based on the first point cloud feature through the model creation module.

[0153] In a possible implementation, the acquisition unit 1101 is configured to:

[0154] acquire an initial image of the preset object, the initial image being a single image of the preset object;

[0155] perform multi-view image generation on the initial image through an image generation model to obtain a two-dimensional image corresponding to the preset object under each of a plurality of views;

[0156] determine the two-dimensional images corresponding to the preset object under the plurality of views as two-dimensional images of the preset object.

[0157] In a possible implementation, the acquisition unit 1101 is configured to:

[0158] perform multi-view image generation on the initial image through the image generation model to obtain a multi-view image corresponding to the initial image, the multi-view image being a single image including view combinations, and the view combinations being combinations of the two-dimensional images corresponding to the preset object under the plurality of views;

[0159] crop the two-dimensional image of the preset object under each of the plurality of views from the multi-view image.

[0160] In a possible implementation, the acquisition unit 1101 is configured to:

[0161] acquire a control condition, the control condition being used to indicate the plurality of views;

[0162] control the image generation model to generate the two-dimensional images corresponding to the preset object under the plurality of views based on the initial image according to the plurality of views indicated by the control condition.

[0163] In a possible implementation, the extraction unit 1102 is configured to:

[0164] perform feature extraction on the two-dimensional image under each view respectively to obtain a first image feature of the two-dimensional image under each view;

[0165] The feature conversion on the first image feature through the point cloud feature conversion network to obtain the corresponding first point cloud feature includes:

[0166] The first sub-point cloud feature of the two-dimensional image under each view is converted by the point cloud feature conversion network to obtain a first sub-point cloud feature of the two-dimensional image under each view.

[0167] The first sub-point cloud features of the two-dimensional image under each view are summed to obtain the first point cloud feature.

[0168] In a possible implementation, the model creation module includes a trained plane feature conversion network and a trained generation network, and the generation unit 1104 is configured to:

[0169] The first point cloud feature is input into the plane feature conversion network, and the first point cloud feature is converted into a three-plane feature in a three-dimensional space formed by three orthogonal planes by the plane feature conversion network, the three-plane feature representing a voxel to which a point corresponding to the first point cloud feature belongs in the three-dimensional space.

[0170] The three-dimensional model of the preset object is output based on the three-plane feature by the generation network.

[0171] In a possible implementation, the generation unit 1104 is configured to:

[0172] A first directed distance function value between a vertex of the voxel to which the three-plane feature belongs and a surface of the preset object is calculated by the generation network.

[0173] The vertex for which the first directed distance function value is zero is extracted by the generation network, and the three-dimensional model is generated by using the vertex for which the first directed distance function value is zero.

[0174] In a possible implementation, the apparatus further includes a training unit, and the training unit is configured to:

[0175] A sample image of the sample object is obtained, and the standard point cloud feature is obtained.

[0176] A sample image feature of the sample image is obtained by performing feature extraction on the sample image.

[0177] A first predicted point cloud feature is obtained by performing feature conversion on the sample image feature by an initial network model.

[0178] The initial network model is trained based on a difference between the first predicted point cloud feature and the standard point cloud feature to obtain the point cloud feature conversion network.

[0179] In a possible implementation, the training unit is configured to:

[0180] The second point cloud feature generated by the trained point cloud encoder is determined as the standard point cloud feature, and the second point cloud feature generated by the trained point cloud encoder is obtained by feature extraction on the point cloud data of the sample object by the trained point cloud encoder.

[0181] In a possible implementation, the training unit is configured to:

[0182] obtain point cloud data of the sample object;

[0183] extract features from the point cloud data by an initial point cloud encoder to obtain a second predicted point cloud feature;

[0184] input the second predicted point cloud feature into an initial creation module, and output a predicted three-dimensional model of the sample object by the initial creation module based on the second predicted point cloud feature;

[0185] adjust the point cloud encoder and the initial creation module based on a difference between the predicted three-dimensional model and a standard three-dimensional model of the sample object, and train a three-dimensional generation network model, wherein the three-dimensional generation network model comprises the trained point cloud encoder and a model creation module trained based on the initial creation module, and the trained point cloud encoder is trained based on the initial point cloud encoder.

[0186] In a possible implementation, the initial creation module comprises a first initial network and a second initial network, and the training unit is configured to:

[0187] perform three-plane feature conversion on the second predicted point cloud feature in a three-dimensional space formed by three orthogonal planes by the first initial network to obtain a predicted three-plane feature, and the predicted three-plane feature represents a voxel to which a point corresponding to the second predicted point cloud feature belongs in the three-dimensional space;

[0188] input the predicted three-plane feature into the second initial network, and output the predicted three-dimensional model by the second initial network based on the predicted three-plane feature;

[0189] The adjusting the point cloud encoder and the initial creation module based on the difference between the predicted three-dimensional model and the standard three-dimensional model of the sample object to train the three-dimensional generation network model comprises:

[0190] The initial point cloud encoder, the first initial network and the second initial network are adjusted using a difference between the predicted three-dimensional model and a standard three-dimensional model of the sample object, and the three-dimensional generation network model is trained, the three-dimensional generation network model including the trained point cloud encoder, the plane feature conversion network trained based on the first initial network and the generation network trained based on the second initial network, and the trained point cloud encoder being trained based on the initial point cloud encoder.

[0191] In a possible implementation, the training unit is configured to:

[0192] The second directional distance function value between the vertex of the voxel to which the predicted three-plane feature belongs and the surface of the sample object is calculated by the second initial network.

[0193] The vertex with the zero second directional distance function value is extracted by the second initial network, and the predicted three-dimensional model is generated using the vertex with the zero second directional distance function value.

[0194] As can be seen from the above technical solutions, when three-dimensional reconstruction of a preset object is needed, a two-dimensional image of the preset object can be acquired, and a first image feature of the two-dimensional image can be obtained by performing feature extraction on the two-dimensional image. However, the present application does not directly perform three-dimensional reconstruction based on the image feature, but performs feature conversion on the image feature by using a point cloud feature conversion network to obtain a first point cloud feature corresponding to the first image feature. The point cloud feature conversion network is trained by using a standard point cloud feature, that is, a real point cloud feature, as a supervision signal and a sample image of a sample object as input data, and learns the conversion relationship between the image feature and the point cloud feature in the training process, so that the trained point cloud feature conversion network has the function of converting the image feature into the point cloud feature. Therefore, the first point cloud feature can be accurately obtained based on the first image feature by using the point cloud feature conversion network. The first point cloud feature can reflect the complete three-dimensional spatial information of the preset object and reflect the three-dimensional geometric shape of the surface of the preset object, and can understand and reconstruct the object from multiple angles and aspects, thereby providing more abundant content for three-dimensional reconstruction. Therefore, when a three-dimensional model of the preset object is output by the model creation module based on the first point cloud feature, even if some parts are invisible, the information of these invisible areas can be accurately inferred and supplemented by using the first point cloud feature from other angles, thereby improving the reconstruction effect of the three-dimensional model and avoiding problems such as severe geometric deformation, singularity or geometric blur.

[0195] The embodiment of the present application further provides a computer device which can execute the image processing method. The computer device can be a terminal, and FIG. 12 shows a structural diagram of a terminal according to an embodiment of the present application. In FIG. 12, the terminal is taken as a smartphone as an example:

[0196] Referring to FIG. 12, the smartphone includes a radio frequency (RF) circuit 1210, a memory 1220, an input unit 1230, a display unit 1240, a sensor 1250, an audio circuit 1260, a wireless fidelity (WiFi) module 1270, a processor 1280, and a power supply 1290, and the like. The input unit 1230 can include a touch panel 1231 and other input devices 1232, and the display unit 1240 can include a display panel 1241. The audio circuit 1260 can include a speaker 1261 and a microphone 1262. It can be understood that the structure of the smartphone shown in FIG. 12 does not constitute a limitation on the smartphone, and the smartphone can include more or fewer components than those shown in the figure, or some components can be combined, or different components can be arranged.

[0197] The memory 1220 can be used to store software programs and modules, and the processor 1280 executes various function applications and data processing of the smartphone by running the software programs and modules stored in the memory 1220. The memory 1220 can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, at least one application program required by a function (such as a sound playing function, an image playing function, etc.), and the like; and the data storage area can store data created according to the use of the smartphone (such as audio data, a phone book, etc.), and the like. In addition, the memory 1220 can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state memory device.

[0198] The processor 1280 is the control center of the smartphone, and connects all parts of the smartphone through various interfaces and lines, and executes various functions and processes data of the smartphone by running or executing the software programs and / or modules stored in the memory 1220, and calling the data stored in the memory 1220. Optionally, the processor 1280 can include one or more processing units; preferably, the processor 1280 can integrate an application processor and a modem processor, wherein the application processor mainly processes an operating system, a user interface, and an application program, and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor can also not be integrated into the processor 1280.

[0199] In this embodiment, the processor 1280 in the smart phone can perform the image processing method provided in the embodiments of the present application.

[0200] The computer device provided in the embodiments of the present application can also be a server. Referring to FIG. 13, FIG. 13 is a structural diagram of a server 1300 provided in the embodiments of the present application. The server 1300 can have great differences due to different configurations or performances. The server 1300 can include one or more processors, for example, a central processing unit (CPU) 1322, and a memory 1332, one or more storage media 1330 (for example, one or more mass storage devices) storing application programs 1342 or data 1344. The memory 1332 and the storage media 1330 can be temporary storage or persistent storage. The programs stored in the storage media 1330 can include one or more modules (not shown in the figure), and each module can include a series of instruction operations in the server. Further, the central processing unit 1322 can be configured to communicate with the storage media 1330 and execute the series of instruction operations in the storage media 1330 on the server 1300.

[0201] The server 1300 can also include one or more power supplies 1326, one or more wired or wireless network interfaces 1350, one or more input and output interfaces 1358, and / or one or more operating systems 1341, for example, Windows Server TM , Mac OS X TM , Unix TM , Linux TM , FreeBSD TM , and the like.

[0202] In this embodiment, the central processing unit 1322 in the server 1300 can perform the image processing method provided in the embodiments of the present application.

[0203] According to an aspect of the present application, a computer readable storage medium is provided, the computer readable storage medium is used to store a computer program, the computer program is used to execute the image processing method provided in the foregoing embodiments.

[0204] According to an aspect of the present application, a computer program product is provided, the computer program product includes a computer program stored in a computer readable storage medium. A processor of a computer device reads the computer program from the computer readable storage medium, and the processor executes the computer program, so that the computer device executes the image processing method provided in the various optional implementation manners of the above embodiments.

[0205] The descriptions of the processes or structures corresponding to the respective figures above each have their own emphasis, and the parts not described in detail in a certain process or structure can be referred to the relevant descriptions of other processes or structures.

[0206] The terms "first", "second", "third", "fourth" and the like in the description of the application and in the claims of the foregoing can be used for distinguishing between similar elements and not necessarily for describing a particular sequential or chronological order. It is to be understood that the use of these terms herein is to be construed as interchangeable in order to distinguish the embodiments of the application from one another. Furthermore, the terms "comprising", "having", "including" and "containing" and any variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, system, product or apparatus that comprises, has, includes or contains an item or a list of items who does not include all of the recited items can still be deemed to be encompassed by the expression.

[0207] In several embodiments provided in the present application, it should be understood that the disclosed system, device and method can be implemented in other ways. For example, the device embodiments described above are merely illustrative, for example, the division of the units is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units or components shown or discussed can be indirect coupling or communication connection through some interface, device or unit, and can be electrical, mechanical or other forms.

[0208] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, that is, they can be located in one place, or can be distributed on a plurality of network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment scheme.

[0209] In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.

[0210] The integrated unit, if implemented in the form of a software function unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or say the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a terminal, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various computer program storage media that can store computer programs.

[0211] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, and works together with other related parts to achieve a predetermined target, and can be implemented entirely or partially by using software, hardware (such as a processing circuit or a memory) or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an integral module or unit that includes the functions of the module or unit.

[0212] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacements for some technical features; and these modifications or replacements do not cause the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. An image processing method, the method being executed by a computer device, the method comprising: acquiring a two-dimensional image of a preset object; performing feature extraction on the two-dimensional image to obtain first image features of the two-dimensional image; performing feature conversion on the first image features by a point cloud feature conversion network to obtain first point cloud features corresponding to the first image features, the first point cloud features being used to reflect three-dimensional spatial information of the preset object, the point cloud feature conversion network being obtained by training with a standard point cloud feature as a supervision signal and a sample image of a sample object as input data, the point cloud feature conversion network having a function of converting image features into point cloud features, the first point cloud features obtained by feature conversion having the same feature dimension as the standard point cloud features; inputting the first point cloud features into a model creation module, and outputting a three-dimensional model of the preset object by the model creation module based on the first point cloud features.

2. The method of claim 1, wherein, The acquiring of the two-dimensional image of the preset object comprises: acquiring an initial image of the preset object, the initial image being a single image of the preset object; generating multiple-view images from the initial image by an image generation model to obtain two-dimensional images corresponding to the preset object under multiple views respectively; determining the two-dimensional images corresponding to the preset object under multiple views respectively as the two-dimensional images of the preset object.

3. The method of claim 2, wherein, The generating of the multiple-view images from the initial image by the image generation model to obtain the two-dimensional images corresponding to the preset object under multiple views respectively comprises: generating multiple-view images corresponding to the initial image by the image generation model, the multiple-view images being single images including view combinations, the view combinations being combinations of the two-dimensional images corresponding to the preset object under the multiple views respectively; cropping the two-dimensional images of the preset object under each of the multiple views from the multiple-view images.

4. The method according to claim 2 or 3, characterized in that, The generating of the multiple-view images from the initial image by the image generation model to obtain the two-dimensional images corresponding to the preset object under multiple views respectively comprises: acquiring a control condition, the control condition being used to indicate the multiple views; controlling the image generation model to generate the two-dimensional images corresponding to the preset object under the multiple views respectively based on the initial image according to the multiple views indicated by the control condition.

5. The method according to any one of claims 2 to 4, characterized in that, The performing of feature extraction on the two-dimensional image to obtain the first image features of the two-dimensional image comprises: performing feature extraction on the two-dimensional images under each view respectively to obtain first image features of the two-dimensional images under each view; The performing of feature conversion on the first image features by the point cloud feature conversion network to obtain corresponding first point cloud features comprises: performing feature conversion on the first image features of the two-dimensional images under each view by the point cloud feature conversion network to obtain first sub-point cloud features of the two-dimensional images under each view; summing the first sub-point cloud features of the two-dimensional images under each view to obtain the corresponding first point cloud features.

6. The method according to any one of claims 1 to 5, characterized in that, The model creation module comprises a plane feature conversion network completed with training and a generation network completed with training, the first point cloud feature is input into the model creation module, and a three-dimensional model of the preset object is output by the model creation module based on the first point cloud feature. The first point cloud feature is input into the plane feature conversion network, the first point cloud feature is converted into three-plane features in a three-dimensional space formed by three orthogonal planes by the plane feature conversion network, and the three-plane features represent voxels to which points corresponding to the first point cloud feature belong in the three-dimensional space. The three-dimensional model of the preset object is output by the generation network based on the three-plane features.

7. The method of claim 6, wherein, The three-dimensional model of the preset object is output by the generation network based on the three-plane features, comprising: A first directed distance function value between a vertex of a voxel to which the three-plane features belong and a surface of the preset object is calculated by the generation network; The three-dimensional model is generated by the generation network using the vertex with the first directed distance function value of zero.

8. The method according to any one of claims 1 to 7, characterized in that, The method further comprises: A sample image of the sample object is acquired, and a standard point cloud feature is acquired; Feature extraction is performed on the sample image to obtain a sample image feature of the sample image; Feature conversion is performed on the sample image feature by an initial network model to obtain a first predicted point cloud feature; The initial network model is trained based on a difference between the first predicted point cloud feature and the standard point cloud feature to obtain the point cloud feature conversion network.

9. The method of claim 8, wherein, The standard point cloud feature is acquired, comprising: A second point cloud feature generated by a trained point cloud encoder is determined as the standard point cloud feature, and the second point cloud feature generated by the trained point cloud encoder is obtained by performing feature extraction on point cloud data of the sample object by the trained point cloud encoder.

10. The method of claim 9, wherein, The method further comprises: Point cloud data of the sample object is acquired; Feature extraction is performed on the point cloud data by an initial point cloud encoder to obtain a second predicted point cloud feature; The second predicted point cloud feature is input into an initial creation module, and a predicted three-dimensional model of the sample object is output by the initial creation module based on the second predicted point cloud feature; A three-dimensional generation network model is trained by adjusting the point cloud encoder and the initial creation module based on a difference between the predicted three-dimensional model and a standard three-dimensional model of the sample object, the three-dimensional generation network model comprises the trained point cloud encoder and a model creation module trained based on the initial creation module, and the trained point cloud encoder is trained based on the initial point cloud encoder.

11. The method of claim 10, wherein, The initial creation module comprises a first initial network and a second initial network, the second predicted point cloud feature is input into the initial creation module, and a predicted three-dimensional model of the sample object is output by the initial creation module based on the second predicted point cloud feature, comprising: The second prediction point cloud feature is converted into a three-plane feature in a three-dimensional space formed by three orthogonal planes through the first initial network to obtain a prediction three-plane feature, the prediction three-plane feature representing a voxel to which a point corresponding to the second prediction point cloud feature belongs in the three-dimensional space; The prediction three-plane feature is input into the second initial network, and the prediction three-dimensional model is output by the second initial network based on the prediction three-plane feature; The point cloud encoder and the initial creation module are adjusted based on the difference between the prediction three-dimensional model and the standard three-dimensional model of the sample object, and a three-dimensional generation network model is trained, including: The initial point cloud encoder, the first initial network and the second initial network are adjusted based on the difference between the prediction three-dimensional model and the standard three-dimensional model of the sample object, and the three-dimensional generation network model is trained, the three-dimensional generation network model including the trained point cloud encoder, the plane feature conversion network trained based on the first initial network and the generation network trained based on the second initial network, and the trained point cloud encoder being trained based on the initial point cloud encoder.

12. The method of claim 11, wherein, The prediction three-dimensional model is output by the second initial network based on the prediction three-plane feature, including: A second directed distance function value between a vertex of a voxel to which the prediction three-plane feature belongs and a surface of the sample object is calculated by the second initial network; The vertex with the zero second directed distance function value is extracted by the second initial network, and the prediction three-dimensional model is generated based on the vertex with the zero second directed distance function value.

13. An image processing apparatus characterized by comprising: The device includes an acquisition unit, an extraction unit, a conversion unit and a generation unit: The acquisition unit is configured to acquire a two-dimensional image of a preset object; The extraction unit is configured to perform feature extraction on the two-dimensional image to obtain a first image feature of the two-dimensional image; The conversion unit is configured to perform feature conversion on the first image feature by a point cloud feature conversion network to obtain a corresponding first point cloud feature, the first point cloud feature being used to represent three-dimensional space information of the preset object, the point cloud feature conversion network being trained based on a standard point cloud feature as a supervision signal and a sample image of a sample object as an input, the trained point cloud feature conversion network having a function of converting an image feature into a point cloud feature, and the first point cloud feature obtained by feature conversion having the same feature dimension as the standard point cloud feature; The generation unit is configured to input the first point cloud feature into a model creation module, and output a three-dimensional model of the preset object by the model creation module based on the first point cloud feature.

14. A computer device, comprising: The computer device includes a processor and a memory: The memory is configured to store a computer program and transmit the computer program to the processor; The processor is configured to execute the method according to the instructions in the computer program.

15. A computer-readable storage medium, characterized in that, The computer-readable storage medium is configured to store a computer program, which, when executed by a processor, causes the processor to perform the method of any one of claims 1-12.

16. A computer program product comprising a computer program which, when run on a computer device, causes the computer device to perform the method of any one of claims 1-12.

Citation Information

Patent Citations

  • Scattered image based crop fruit three-dimensional point cloud extracting system

    CN108198230A

  • Method and system for reconstructing single image to three-dimensional point cloud model based on attention mechanism

    CN112258625A

  • Three-dimensional model generation method and system for generating dense point cloud based on image cascading

    CN112258626A

  • Brain structure three-dimensional reconstruction method and device and terminal equipment

    CN112598790A

  • Three-dimensional reconstruction method and device for brain structure in extreme environment and readable storage medium

    CN113920243A