Method, device and electronic equipment for creating three-dimensional model of object

Through the image processing model, the pixel correspondence relationship and point cloud information of multiple images are solved, and the problem of professional skills and time-consuming in traditional three-dimensional modeling technology is achieved, and fast and accurate three-dimensional model generation is achieved.

CN119810337BActive Publication Date: 2025-08-08TAOBAO CHINA SOFTWARE
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510280406.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-08-08
Estimated Expiration
2045-03-10

AI Technical Summary

Technical Problem

Traditional three-dimensional modeling technology requires users to have high professional skills, and the modeling process takes a long time, so the created three-dimensional model has poor accuracy, especially in the case of uneven image quality.

Method used

By acquiring multiple images of the target object, a pre-trained image processing model is used to determine the pixel correspondence between image pairs and initial point cloud information, and fine-grained processing is performed based on this information to create a three-dimensional model, including estimating the camera's internal and external parameters, determining the fine point cloud information and normal vectors, and finally generating an accurate three-dimensional model.

Benefits of technology

The user can quickly create high-accuracy three-dimensional models without the need for advanced modeling skills. The image processing model adapts to different image quality through a large number of training samples. The generated model matches the actual object with high degree of matching, adapts to low texture, blur, low resolution and other situations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119810337B_ABST
    Figure CN119810337B_ABST
Patent Text Reader

Abstract

The present application discloses a method, device, electronic device, and computer-readable storage medium for creating a three-dimensional model of an object. The method for creating a three-dimensional model of an object includes: obtaining multiple target images corresponding to the target object for which the three-dimensional model is to be created; determining multiple groups of image pairs consisting of the multiple target images; determining the pixel correspondence between two images in the image pair and the initial point cloud information corresponding to each target image based on a pre-trained image processing model, wherein the image pair includes two images from the multiple target images; determining the fine point cloud information corresponding to the target image based on the initial point cloud information and the pixel correspondence; and creating a three-dimensional model corresponding to the target object based on the fine point cloud information corresponding to each target image. The solution provided by the present application can efficiently and quickly create a three-dimensional model of an object based on its image, and the accuracy of the created three-dimensional model is higher.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computers, and specifically to a method, device, electronic device, and computer-readable storage medium for creating a three-dimensional model of an object. Background Art

[0002] With the rapid development of computer vision, 3D modeling of objects has been widely used in various fields, such as e-commerce, product image generation, product display, and personalized consumer matching. It can also be applied to virtual reality (VR), game development, cultural heritage protection, industrial design and other fields. Among them, creating a corresponding 3D model of an object based on a photo of the object is currently a relatively common 3D modeling method. It can create a corresponding 3D model of the object according to the appearance of the object shown in the photo, and then apply the created 3D model to the 3D display of the object, allowing users to observe and interact from different angles, greatly enriching the user experience and improving work efficiency.

[0003] Traditional 3D modeling technology typically involves users creating a matching 3D model of an object based on an image using 3D modeling software. However, this 3D modeling method requires users to possess advanced 3D modeling expertise, and the modeling process is often time-consuming. Furthermore, due to the large differences in technical skill levels among modelers, the accuracy of the created 3D models of objects is also relatively poor. Summary of the Invention

[0004] This application provides a method, device, electronic device, and computer-readable storage medium for creating a three-dimensional model of an object. These methods can efficiently and quickly create a three-dimensional model of an object based on an image of the object, without requiring the user to possess advanced three-dimensional modeling expertise. Furthermore, the created three-dimensional model has a higher degree of accuracy. The specific solution is as follows:

[0005] In a first aspect, the present application provides a method for creating a three-dimensional model of an object, the method comprising:

[0006] Acquire multiple target images corresponding to a target object for which a three-dimensional model is to be created, wherein the multiple target images are used to display the target object from multiple perspectives;

[0007] Determining a plurality of groups of image pairs consisting of the plurality of target images, each group of the image pairs including two different target images;

[0008] Determining a pixel correspondence between the two images in the image pair and initial point cloud information corresponding to each of the target images based on a pre-trained image processing model, wherein the pixel correspondence indicates the pixel position of the object part corresponding to the pixel position in one image of the image pair and the corresponding pixel position in the other image;

[0009] Determining fine point cloud information corresponding to the target image based on the initial point cloud information of the target image and the pixel correspondence relationship;

[0010] A three-dimensional model corresponding to the target object is created based on the fine point cloud information corresponding to each of the target images.

[0011] Optionally, determining the fine point cloud information corresponding to the target image based on the initial point cloud information of the target image and the pixel correspondence includes:

[0012] estimating initial camera intrinsic parameters corresponding to the target image based on initial point cloud information of the target image;

[0013] Based on the pixel correspondence, the initial point cloud information corresponding to the target image, and the initial camera intrinsic parameters, fine point cloud information corresponding to the target image is determined.

[0014] Optionally, estimating initial camera intrinsic parameters corresponding to the target image based on the initial point cloud information of the target image includes:

[0015] Based on the principle of minimizing the difference between the converted pixel position and the corresponding actual pixel position, the initial camera intrinsic parameters corresponding to the target image are estimated. The converted pixel position is the corresponding pixel position obtained by converting the initial point cloud information corresponding to each pixel point of the target image according to the estimated initial camera intrinsic parameters. The actual pixel position is the corresponding actual pixel position in the target image.

[0016] Optionally, the initial point cloud information corresponding to the pixel points of the target image includes x Dimensional initial point cloud information, y Dimensional initial point cloud information and z Dimensional initial point cloud information;

[0017] The converted pixel position is obtained as follows:

[0018] The target image x Dimensional initial point cloud information and y The initial point cloud information of dimensions is respectively z Normalization processing is performed on the dimensional initial point cloud information to obtain normalized initial point cloud information corresponding to each pixel point of the target image;

[0019] The normalized initial point cloud information is converted into a converted pixel position according to the estimated initial camera intrinsic parameters.

[0020] Optionally, determining the fine point cloud information corresponding to the target image based on the pixel correspondence, the initial point cloud information corresponding to the target image, and the initial camera intrinsic parameters includes:

[0021] estimating the initial camera extrinsic parameters corresponding to the target image based on the pixel correspondence relationship and the initial point cloud information corresponding to the target image;

[0022] Based on the initial camera extrinsic parameters, the pixel correspondence, the initial point cloud information corresponding to the target image, and the initial camera intrinsic parameters, fine point cloud information corresponding to the target image is determined.

[0023] Optionally, estimating the initial camera extrinsic parameters corresponding to the target image based on the pixel correspondence relationship and the initial point cloud information corresponding to the target image includes:

[0024] Based on the principle of minimizing the gap between the matched world point cloud information, the initial camera extrinsic parameters corresponding to the target image are estimated. The matched world point cloud information is the point cloud information of the initial point cloud information of the pixel points corresponding to the two images in the image pair after being converted to the world coordinate system according to the estimated initial camera extrinsic parameters.

[0025] Optionally, estimating the initial camera extrinsic parameters corresponding to the target image based on the principle of minimizing the gap between the matched world point cloud information includes:

[0026] Determining a weight corresponding to each pixel in the target image, where the weight represents the importance of the pixel in representing the target object, and the weights of corresponding pixels in the two images of the image pair are the same;

[0027] Determine the product of the gap between the matching world point cloud information and the corresponding weight as the first weighted gap for each pair of matching pixels;

[0028] The initial camera extrinsic parameters corresponding to the target image are estimated based on the principle of minimizing the sum of the first weighted differences corresponding to the image pair.

[0029] Optionally, determining the fine point cloud information corresponding to the target image based on the initial camera extrinsic parameters, the pixel correspondence, the initial point cloud information corresponding to the target image, and the initial camera intrinsic parameters includes:

[0030] determining fine point cloud information corresponding to the target image based on the principle of minimizing the difference between the actual pixel position of a pixel point in one image of the image pair and the converted pixel position corresponding to the matched pixel point, where the matched pixel point is the matched pixel point in the other image of the image pair;

[0031] The converted pixel position corresponding to the matched pixel point is determined by:

[0032] Converting the initial point cloud information corresponding to the matched pixel points into world point cloud information based on the initial camera extrinsic parameters, the pixel correspondence relationship, and the initial camera intrinsic parameters;

[0033] The world point cloud information corresponding to the matched pixel point is converted to a two-dimensional image position to obtain a converted pixel position corresponding to the matched pixel point.

[0034] Optionally, determining the fine point cloud information corresponding to the target image based on the principle of minimizing the difference between the actual pixel position of a pixel point in one image of the image pair and the converted pixel position corresponding to the matched pixel point includes:

[0035] Determining a weight corresponding to each pixel in the target image, where the weight represents the importance of the pixel in representing the target object, and the weights of corresponding pixels in the two images of the image pair are the same;

[0036] Determine the product of the difference between the actual pixel position of a pixel point in one image of the image pair and the converted pixel position corresponding to the matched pixel point and the corresponding weight as the second weighted difference for each pair of matched pixels;

[0037] The fine point cloud information corresponding to the target image is determined based on the principle of minimizing the sum of the second weighted differences corresponding to the image pair.

[0038] Optionally, the camera external parameters include at least one of camera pose and scaling factor, and the camera internal parameters include camera focal length.

[0039] Optionally, determining a plurality of image pairs consisting of the plurality of target images includes:

[0040] According to the similarities between the target images, the multiple target images are grouped into multiple groups of image pairs.

[0041] Optionally, grouping the plurality of target images into a plurality of image pairs according to the similarities between the target images comprises:

[0042] Extracting image coding features corresponding to each of the target images based on a pre-trained image coding model;

[0043] Calculating the similarity between each target image and other target images based on the image coding features;

[0044] The multiple target images are grouped into multiple image pairs according to the similarity between each target image and other target images.

[0045] Optionally, grouping the plurality of target images into a plurality of image pairs according to similarities between each target image and other target images includes:

[0046] For each target image, images ranked first in a preset number of similarities with the target image among other target images are respectively formed into image pairs with the target image, or images ranked first in a preset similarity threshold with the target image among other target images are respectively formed into image pairs with the target image.

[0047] Optionally, grouping the plurality of target images into a plurality of image pairs according to similarities between each target image and other target images includes:

[0048] Based on the principle that the total number of image pairs does not exceed a second preset number, for each target image, other target images that rank top preset number of similarities with the target image are respectively formed into image pairs.

[0049] Optionally, determining the pixel correspondence between the two images in the image pair and the initial point cloud information corresponding to each of the target images based on a pre-trained image processing model includes:

[0050] The image coding features corresponding to each of the target images and the information of the two images constituting the image pair are input into a pre-trained image processing model to obtain the pixel correspondence between the two images in the image pair and the initial point cloud information corresponding to each of the target images.

[0051] Optionally, determining the fine point cloud information corresponding to the target image based on the principle of minimizing the difference between the actual pixel position of a pixel point in one image of the image pair and the converted pixel position corresponding to the matched pixel point includes:

[0052] Determine the fine point cloud information corresponding to the target image and the precise camera intrinsic parameters corresponding to the target image based on the principle of minimizing the difference between the actual pixel position of a pixel point in one image of the image pair and the converted pixel position corresponding to the matched pixel point;

[0053] The method further comprises:

[0054] Determining a global normal vector corresponding to the target image based on the precise camera intrinsic parameters and the fine point cloud information corresponding to the target image;

[0055] Determining an accurate normal vector corresponding to the target image based on a global normal vector corresponding to the target image;

[0056] The step of creating a three-dimensional model corresponding to the target object based on the fine point cloud information corresponding to each target image includes:

[0057] A three-dimensional model corresponding to the target object is created based on the fine point cloud information corresponding to each target image and the accurate normal vector corresponding to each target image.

[0058] Optionally, the pre-trained image processing model is further used to determine a predicted normal vector corresponding to each of the image pairs during the process of determining the pixel correspondence relationship and the initial point cloud information;

[0059] Determining the accurate normal vector corresponding to the target image according to the global normal vector corresponding to the target image includes:

[0060] The accurate normal vector corresponding to the target image is determined according to the global normal vector corresponding to the target image and the corresponding predicted normal vectors.

[0061] Optionally, determining the accurate normal vector corresponding to the target image according to the global normal vector corresponding to the target image and the corresponding predicted normal vectors includes:

[0062] Calculating the error between each predicted normal vector corresponding to the target image and the corresponding global normal vector;

[0063] Eliminate pixels whose errors with the global normal vector are greater than a preset error from the predicted normal vectors corresponding to the target image, and obtain excluded predicted normal vectors corresponding to the target image after excluding the pixels;

[0064] The accurate normal vector corresponding to the target image is determined according to the mean value of each excluded predicted normal vector corresponding to the target image.

[0065] Optionally, the method further includes:

[0066] When a pixel point of the target image is excluded from each of the excluded predicted normal vectors, the normal vector of the pixel point in the global normal vector corresponding to the target image is determined as the accurate normal vector corresponding to the pixel point.

[0067] Optionally, creating a three-dimensional model corresponding to the target object according to the fine point cloud information corresponding to each of the target images includes:

[0068] The dense point cloud information corresponding to the target object is determined based on the fine point cloud information corresponding to each target image, and a three-dimensional model corresponding to the target object is created based on the dense point cloud information.

[0069] In a second aspect, the present application further provides a method for creating a three-dimensional model of an object, the method comprising:

[0070] Acquire multiple target images corresponding to a target object for which a three-dimensional model is to be created, wherein the multiple target images are used to display the target object from multiple perspectives;

[0071] Determining a plurality of groups of image pairs consisting of the plurality of target images, each group of the image pairs including two different target images;

[0072] Determining, based on a pre-trained image processing model, a pixel correspondence relationship between the two images in the image pair, initial point cloud information corresponding to each target image, and a predicted normal vector corresponding to each image pair, wherein the pixel correspondence relationship indicates the pixel position of an object part corresponding to a pixel position in one image of the image pair and the corresponding pixel position in the other image;

[0073] Determining, based on the initial point cloud information of the target image and the pixel correspondence, fine point cloud information corresponding to the target image and precise camera intrinsic parameters corresponding to the target image;

[0074] Determining a global normal vector corresponding to the target image based on the precise camera intrinsic parameters and the fine point cloud information corresponding to the target image;

[0075] Determining an accurate normal vector corresponding to the target image based on a global normal vector corresponding to the target image and corresponding predicted normal vectors;

[0076] A three-dimensional model corresponding to the target object is created based on the fine point cloud information corresponding to each target image and the accurate normal vector corresponding to each target image.

[0077] In a third aspect, the present application further provides a method for creating a three-dimensional model of an object, which is applied to a server, and the method includes:

[0078] Acquire multiple target images sent by a client for displaying a target object, wherein the multiple target images are used to display the target object from multiple perspectives;

[0079] Creating a three-dimensional model corresponding to the target object by the method for creating a three-dimensional model of an object according to any one of the first aspects;

[0080] The three-dimensional model is sent to the client, so that the client displays the three-dimensional model.

[0081] In a fourth aspect, the present application further provides a device for creating a three-dimensional model of an object, the device comprising:

[0082] an acquisition unit, configured to acquire a plurality of target images corresponding to a target object for which a three-dimensional model is to be created, wherein the plurality of target images are used to display the target object from multiple perspectives;

[0083] a pairing unit, configured to determine a plurality of image pairs consisting of the plurality of target images, each of the image pairs including two different target images;

[0084] a determination unit configured to determine, based on a pre-trained image processing model, a pixel correspondence between the two images in the image pair and initial point cloud information corresponding to each of the target images, wherein the pixel correspondence indicates the pixel position of an object part corresponding to a pixel position in one image of the image pair and the pixel position in the other image; and determine, based on the initial point cloud information of the target image and the pixel correspondence, fine point cloud information corresponding to the target image;

[0085] A modeling unit is used to create a three-dimensional model corresponding to the target object based on the fine point cloud information corresponding to each of the target images.

[0086] In a fifth aspect, the present application also provides an electronic device comprising: a processor, a memory, and computer program instructions stored on the memory and executable on the processor; when the processor executes the computer program instructions, it implements the method described in any one of the first to third aspects.

[0087] In a sixth aspect, the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by a processor, they are used to implement any one of the methods in the first to third aspects.

[0088] In a seventh aspect, an embodiment of the present application provides a computer program product, comprising a computer program, which, when executed by a processor, implements the method as described in any one of the first to third aspects.

[0089] Compared with the prior art, this application has the following advantages:

[0090] The method for creating a three-dimensional model of an object provided in an embodiment of the present application obtains multiple target images corresponding to the target object of the three-dimensional model to be created, and the multiple target images are used to display the target object from multiple perspectives, and then based on a pre-trained image processing model, the pixel correspondence between the two images in the image pair corresponding to the multiple target images and the initial point cloud information corresponding to each of the target images are determined, the image pair contains two images of the multiple target images, and the two images in a group of image pairs can display various parts of the object from different perspectives, so that each group of image pairs can better accurately reflect the information of various parts of the object from the perspective of stereo geometry, the pixel correspondence is used to indicate the pixel position of the object part corresponding to the pixel position of one image in the image pair and the corresponding pixel position in the other image, that is, the pixel correspondence can associate and correspond the pixel points belonging to the same object part in the two images, and the pixel correspondence can well reflect the position information of the same object part at different perspectives, so that the pixel correspondence can well reflect the three-dimensional geometric information of the object, and the above-mentioned image processing model is pre-trained, and the image processing model can determine the pixel correspondence between the two images in each group of the image pairs. The corresponding relationship and the initial point cloud information corresponding to the target image can be quickly determined by a pre-trained image processing model, wherein the initial point cloud information of each target image and the pixel correspondence relationship can reflect the depth image corresponding to each target image, and the initial point cloud information can reflect the geometric information of the object corresponding to a single image. Then, based on the initial point cloud information of the target image and the pixel correspondence relationship, the fine point cloud information corresponding to the target image is determined. Since the pixel correspondence relationship can well reflect the position information of the same object part at different perspectives, the initial point cloud information is refined through the pixel correspondence relationship, so that the refined point cloud information after refinement can have higher accuracy in three-dimensional space, that is, the fine point cloud information can more accurately represent each part of the target object in space. Then, based on the fine point cloud information corresponding to each target image, the three-dimensional model corresponding to the target object is determined and created. Since the fine point cloud information can more accurately represent each part of the target object in space, the three-dimensional model corresponding to the target object created based on the fine point cloud information corresponding to each target image is more accurate in three-dimensional space and more closely matches the actual target object.

[0091] The solution provided by the present application first determines the initial point cloud information corresponding to each target image efficiently and quickly through an image processing model. Since the initial point cloud information is obtained based on a single target image, the initial point cloud information may lack an accurate reflection of the target object in three-dimensional space. The present application then adjusts the initial point cloud information through the pixel correspondence corresponding to each group of images that can well reflect the three-dimensional geometric information of the object, and obtains fine point cloud information that can more accurately represent various parts of the target object in space, so that the three-dimensional model created according to the fine point cloud information corresponding to each of the target images is more matched with the actual target object and has higher accuracy.

[0092] It can be seen that the solution provided by this application can automatically generate a three-dimensional model corresponding to the target object through a computer, without the need for manual operation of the three-dimensional modeling software. Therefore, it is possible to efficiently and quickly create a three-dimensional model of an object based on the image of the object, without the need for the user to have high professional skills in three-dimensional modeling. The created three-dimensional model is more accurate and more consistent with the target object. In addition, because the image processing model is trained using a large number of training samples, the training process usually includes both good-quality images and poor-quality images. Therefore, the trained image processing model has a strong prior, which makes the image processing model predict a larger number of pixel matches. It can also perform good prediction processing for poor image quality such as low texture, blur, low resolution, and poor lighting conditions, and obtain an accurate three-dimensional model. The initial point cloud information of each target image predicted by the pre-trained image processing model itself is also relatively high in accuracy and can contain rich geometric information, creating conditions for obtaining more accurate and detailed point cloud information in the future, so that an accurate three-dimensional model can be obtained in the end. BRIEF DESCRIPTION OF THE DRAWINGS

[0093] Figure 1 This is a schematic diagram of an application scenario of the solution for creating a three-dimensional model of an object provided in this application.

[0094] Figure 2 This is a flowchart of an example of a method for creating a three-dimensional model of an object provided in an embodiment of the present application.

[0095] Figure 3 This is a flowchart of another example of the method for creating a three-dimensional model of an object provided in an embodiment of the present application.

[0096] Figure 4 It is a structural diagram of another example of a device for creating a three-dimensional model of an object provided in an embodiment of the present application.

[0097] Figure 5 This is a structural block diagram of the electronic device provided in this application. DETAILED DESCRIPTION

[0098] In order to enable those skilled in the art to better understand the technical solutions of this application, the following clearly and completely describes this application in conjunction with the drawings in the embodiments of this application. However, this application can be implemented in many other ways different from the following description. Therefore, based on the embodiments provided in this application, all other embodiments obtained by those skilled in the art without creative work should fall within the scope of protection of this application.

[0099] It should be noted that the terms "first", "source domain", "third", etc. in the claims, description and drawings of the present application are used to distinguish similar objects and are not used to describe a specific order or sequence. The data used in this way are interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including", "having" and their variations are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0100] In order to facilitate understanding of the various embodiments of the present application, the application background of the embodiments is described.

[0101] With the rapid development of computer vision, 3D object modeling has gained widespread application in various fields, such as e-commerce, product image generation, product display, and personalized product matching. It can also be applied to virtual reality (VR), game development, cultural heritage preservation, and industrial design. Creating a corresponding 3D model of an object based on its photograph is a common 3D modeling method. This method can create a 3D model of the object based on its appearance in the photograph. This 3D model can then be applied to the 3D display of the object, allowing users to observe and interact with it from different angles, greatly enriching the user experience and improving work efficiency. 3D object reconstruction is a technology that digitally represents and presents real-world objects in three dimensions. In e-commerce scenarios, 3D product reconstruction is an emerging and important technology, widely used in product design, product image generation, product display, and personalized product matching. However, traditional 3D reconstruction relies on specialized filming equipment and locations, which has a high barrier to entry and prevents large-scale application.

[0102] For small items, merchants can ship them to us, and then use professionally constructed shooting scenes and equipment to capture and model them with high precision. However, large items like furniture are inconvenient to transport, and the shooting equipment cannot support large-scale object modeling. In this context, the quality of user-captured images varies greatly, and scene lighting and other factors vary greatly, making 3D reconstruction algorithms face even greater challenges.

[0103] Traditional 3D modeling technology typically involves users creating a matching 3D model of an object based on an image using 3D modeling software. However, this 3D modeling method requires users to possess advanced 3D modeling expertise, and the modeling process is often time-consuming. Furthermore, due to the large differences in technical skill levels among modelers, the accuracy of the created 3D models of objects is also relatively poor.

[0104] In related technologies, multimodal three-dimensional model generation can also be achieved with the help of large model capabilities. However, the three-dimensional models generated by large models focus on rapid and large-scale generation, and pay attention to the integrity of the model itself, but the effect in terms of realism and precision is poor.

[0105] To address the above issues, embodiments of the present application provide a method, apparatus, electronic device, and computer-readable storage medium for creating a three-dimensional model of an object. These methods are designed to efficiently and quickly create a three-dimensional model of an object from an image, without requiring the user to possess advanced 3D modeling expertise, and with greater accuracy.

[0106] The method for creating a three-dimensional model of an object provided in this application can be used in the fields of three-dimensional display of goods, virtual try-on, object design, product image generation, etc., and this application does not specifically limit it.

[0107] In order to facilitate understanding of the method embodiment of this application, its application scenario is introduced. Figure 1 , Figure 1 The following is a schematic diagram of an application scenario of the solution provided in the embodiment of the present application. This application scenario is a schematic illustration and is not intended to be a specific description of its application scenario. Figure 1 As shown, in this application scenario, a server 102 and a client 101 are provided. In this embodiment, a connection is established between the client 101 and the server 102 via network communication, thereby performing data transmission.

[0108] The client 101 can be an electronic device with display and data processing functions, such as a mobile phone, tablet computer (pad), smart watch, desktop computer, smart TV, VR device, vehicle-mounted device, wearable device, laptop computer, etc. The client 101 is used to send a target image corresponding to the target object to the server 102, and obtain a generated three-dimensional model corresponding to the target object from the server 102.

[0109] Server 102 has high computing capabilities. Server 102 can be a server equipped with a high-speed central processing unit (CPU) computing power, long-term reliable operation, strong input / output (I / O) external data throughput, and improved scalability. Server 102 can be a single server or a server cluster. Server 102 deploys an image processing model. Based on the image processing model and target images sent by the client, server 102 can generate a three-dimensional model corresponding to the target object and send the three-dimensional model to client 101. Server 102 can also provide other specific services to client 101, such as access to user information, websites, and applications, which are not specifically limited in this application.

[0110] The client 101 and the server 102 can communicate with each other using various communication systems, such as a wired communication system or a wireless communication system. Examples of the wireless communication system include the global system for mobile communications (GSM), code division multiple access (CDMA), wideband code division multiple access (WCDMA), general packet radio service (GPRS), long term evolution (LTE), LTE frequency division duplex (FDD), LTE time division duplex (TDD), universal mobile telecommunication system (UMTS), worldwide interoperability for microwave access (WiMAX), future fifth generation (5G) or new radio (NR), satellite communication systems, and the like. Example 1

[0111] The first embodiment of the present application provides a method for creating a three-dimensional model of an object. The method is applied to an electronic device, which may be a server, a laptop, a tablet computer, a desktop computer, a mobile phone, a smart TV, or other electronic device with data processing capabilities.

[0112] like Figure 2 As shown, the method for creating a three-dimensional model of an object provided in the first embodiment of the present application includes the following steps S110 to S140.

[0113] Step S110: Acquire multiple target images corresponding to a target object for which a three-dimensional model is to be created, wherein the multiple target images are used to display the target object from multiple perspectives.

[0114] The target object can be merchandise to be displayed, such as furniture, electrical appliances, electronic products, clothing, food, or other objects. The target object can also be other objects, such as human bodies or animals. Three-dimensional modeling of human bodies and animals can be applied to animation, filmmaking, advertising, video synthesis, and other fields. Those skilled in the art can determine the specific content of the target object based on the specific object of the three-dimensional model to be created.

[0115] The target image corresponding to the target object can be a photographic image corresponding to the target object, or can be another form of image corresponding to the target object, such as a drawn image, a composite image, a video screenshot, a video frame, etc. corresponding to the target object, but is not limited thereto. The target image can only show the target object from multiple perspectives, and this application does not limit the specific type of target image. Among them, the photographic image corresponding to the target object can clearly and accurately reflect the specific appearance of the target object, and a more accurate three-dimensional model of the target object can be created based on the photographic image corresponding to the target object.

[0116] For example, the photographer can use a mobile phone, camera or other device with a shooting function to shoot the target object from multiple angles, generate multiple target images corresponding to the target object, and then send the target images to the electronic device so that the electronic device receives each target image.

[0117] The multiple perspectives may include at least one of the following: multiple display orientations, multiple display scaling ratios, multiple display angles, and multiple partial display locations. For example, the target images of the multiple perspectives may include images of the target object taken from different angles, or images of the target object taken from different distances, etc., but are not limited thereto.

[0118] It is understandable that the multiple target images can show the target object from a relatively complete perspective as much as possible, so that complete information of the target object can be obtained from the multiple images, thereby making the final three-dimensional model more accurate.

[0119] Step S120: Determine the pixel correspondence between two images in the image pair corresponding to the multiple target images and the initial point cloud information corresponding to each of the target images based on a pre-trained image processing model, wherein the pixel correspondence is used to indicate the pixel position of the object part corresponding to the pixel position of one image in the image pair and the corresponding pixel position in the other image.

[0120] The image pair includes two images from the multiple target images.

[0121] The image processing model is used to determine the initial point cloud information for each image in an image pair and the pixel correspondence between the two images in the image pair. Specifically, information about each target image included in each image pair and the two target images that make up each image pair can be input into the image processing model so that the image processing model can obtain the specific content of each target image and the images that make up the image pair. The image processing model then determines the pixel correspondence between the two images in the image pair and the initial point cloud information corresponding to each target image based on the input information.

[0122] The image processing model can be trained through a large number of training samples. The training samples may include groups of sample image pairs. The sample image pairs include two similar sample images. Each group of sample image pairs marks the pixel correspondence between the two images contained therein. Each sample image is marked with corresponding sample point cloud information. Through the training samples, the image processing model can be trained using supervised learning algorithms, unsupervised learning algorithms or other training learning algorithms.

[0123] In one embodiment, before step S120, the following step S110a may be further included:

[0124] Step S110a: determining a plurality of image pairs consisting of the plurality of target images, each of the image pairs including two different target images.

[0125] The image pairs corresponding to the multiple target images may include image pairs consisting of two different images from the multiple target images. In practical applications, two images from the multiple target images often differ significantly. For example, one image may be an image of an object taken from the front, while the other image is an image of the object taken from the back. These two images rarely share the same object part, or only a small number of the same object parts are shared between these two images. Such image pairs are of little help in understanding the structure of the same object part from different perspectives. Therefore, to reduce computational complexity and maintain the efficiency and accuracy of 3D modeling, step S110a may be implemented as step S120a.

[0126] Step S120a: Grouping the plurality of target images into a plurality of image pairs according to the similarities between the target images.

[0127] In this case, the image pair includes two similar images from the multiple target images, wherein the two target images in each group of image pairs are similar images, and the similarity of the two target images can be greater than a certain threshold. The two similar images are formed into an image pair, and the parameters such as the perspective and zoom ratio of the two similar images are usually relatively similar, so that the object parts contained in the two images in the image pair are basically the same, that is, each object part contained in one image is very likely to exist in the other similar image, so that the posture and appearance of the object under different perspectives can be reflected through the two images with different perspectives in the image pair.

[0128] In a specific embodiment, step S120a can be implemented according to the following steps S121 to S123.

[0129] Step S121: extracting image coding features corresponding to each target image based on a pre-trained image coding model.

[0130] The image coding model is used to numerically encode the target image to obtain the corresponding image coding features. Image coding features can include at least one of, but are not limited to, vector features, matrix features, tensor features, and hash coding features. When image features are coding features, it is more convenient to subsequently calculate the similarity between different target images.

[0131] The image coding model can be trained through learning algorithms such as supervised learning algorithms, unsupervised learning algorithms, and semi-supervised learning algorithms in related technologies. The image coding model can be trained simultaneously with the image processing model used in the subsequent step S130 to improve training efficiency and at the same time make the synergistic effect of the two trained models more accurate.

[0132] Step S122: Calculate the similarity between each target image and other target images based on the image coding features.

[0133] The similarity between the image coding features corresponding to each two target images can be calculated. Specifically, the vector distance, matrix distance, matrix inner product and cosine similarity, inter-matrix correlation coefficient, hash coding distance, etc. between the image coding features corresponding to each two target images can be calculated. The specific type of image coding features is different, and the measurement parameters used to represent the similarity between the two target images are also different. Those skilled in the art can flexibly set them according to actual conditions.

[0134] This step is to calculate the similarity between every two different target images.

[0135] Step S123: grouping the target images into a plurality of image pairs according to the similarity between each target image and other target images.

[0136] Specifically, for each target image, images among other target images whose similarity with the target image is greater than a preset similarity threshold can be used to form image pairs with the target image, or images among other target images whose similarity with the target image is ranked before a preset number can be used to form image pairs with the target image.

[0137] The above-mentioned similarity threshold and the above-mentioned first preset number can be flexibly set according to needs, and are not specifically limited in this application.

[0138] In this embodiment, by encoding each target image, the similarity between different target images can be conveniently and accurately calculated using the obtained image encoding features, so that the two images in the composed image pair can be similar images.

[0139] In one specific embodiment, step S123 can be implemented as follows: Based on the principle that the total number of image pairs does not exceed a second preset number, for each target image, an image pair is formed from a preset number of other target images that rank top in similarity to the target image. The second preset number is greater than the first preset number. This embodiment can control the total number of image pairs to be less than excessive, thereby reducing the time and complexity of the 3D model generation process. The second preset number should also not be set too small to better ensure algorithm accuracy, resulting in a more accurate 3D model being generated, and achieving a balance between algorithm accuracy and time consumption.

[0140] In step S123, when the number of composed image pairs exceeds a second preset number, the image pair with the lowest similarity corresponding to each target image may be deleted. The image pair with the lowest similarity is the image pair whose similarity between the two images in the group is the lowest among the image pairs composed of the target image.

[0141] In other embodiments, step S120a can also be implemented according to the following steps: calculating the color histogram of each target image, using a distance comparison method such as Bhattacharyya distance, chi-square distance, etc. to compare the similarity between the color histograms corresponding to different target images, and grouping the multiple target images into multiple groups of image pairs based on the similarity between the color histograms corresponding to different target images. Specifically, the image pairing can be performed in the manner of performing image pairing based on similarity as described above, which will not be repeated here. The similarity between the color histograms corresponding to the target images can be determined as the similarity between the target images. Those skilled in the art can also calculate the similarity between different target images by other methods and pair the target images based on the similarity, which is not specifically limited in this application.

[0142] The aforementioned pixel correspondence relationship indicates the pixel location of the object part corresponding to the pixel location in the other image of the image pair. In other words, the pixel correspondence relationship indicates the pixel locations of the same object part in both images of the image pair. For example, the pixel correspondence relationship indicates the pixel locations of a corner of a table in both images of the image pair. The pixels corresponding to the corner of the table in the two images have a pixel correspondence relationship.

[0143] A point cloud is used to represent three-dimensional spatial data. It consists of a series of discrete data points with specific locations in a three-dimensional coordinate system. Together, they represent a three-dimensional model of an object or scene. Each point contains its coordinate information in three-dimensional space (usually X, Y, and Z values). The initial point cloud information corresponding to the target image is used to represent the target object in the target image in a three-dimensional form as a point cloud.

[0144] In a specific embodiment, step S120 can be implemented by inputting the image coding features corresponding to each target image and the information of the two images constituting the image pair into a pre-trained image processing model to obtain the pixel correspondence between the two images in the image pair and the initial point cloud information corresponding to each target image. The image processing model can be a decoding model. Since the intelligent model processes various encoded data, inputting the image coding features into the image processing model can improve the efficiency of the model in outputting the initial point cloud information and pixel correspondence, thereby improving the efficiency of 3D modeling.

[0145] Step S130: Based on the initial point cloud information of the target image and the pixel correspondence, determine the fine point cloud information corresponding to the target image.

[0146] Since the initial point cloud information directly determined by the intelligent model in step S130 may have low accuracy and lack accurate representation of three-dimensional space, and the pixel correspondence relationship can reflect the structural state of the same part of the object in images from two different perspectives, the pixel correspondence relationship reflects the structural state of the object at different perspectives. This step can further optimize and adjust the initial point cloud information from a three-dimensional geometric perspective based on the pixel correspondence relationship to obtain more accurate and fine point cloud information.

[0147] In one specific embodiment, the refined point cloud information corresponding to the target image can be determined by performing multi-view point cloud registration on the initial point cloud information based on the aforementioned pixel correspondences. Specifically, the multi-view point cloud registration can be performed using an iterative closest point (ICP) algorithm or other optimization algorithm to align point clouds reflecting different viewpoints in an image pair, so that the processed initial point cloud information of the target image is aligned at two different viewpoints in the image pair. Based on the pixel correspondences and camera parameters (intrinsic and extrinsic parameters), the position of each point in the processed initial point cloud information is further refined through triangulation or other methods to obtain refined point cloud information. Alternatively, the refined point cloud information corresponding to the target image can be determined by inputting the initial point cloud information of the target image and the pixel correspondences into a pre-trained refined point cloud determination model to obtain the refined point cloud information corresponding to the target image. The refined point cloud determination model is configured to determine the refined point cloud information corresponding to the initial point cloud information of the image based on the pixel correspondences between the two images.

[0148] In another specific embodiment, step S130 can be implemented according to the following steps S131 to S132.

[0149] Step S131: estimating initial camera intrinsic parameters corresponding to the target image based on the initial point cloud information of the target image.

[0150] Optionally, the initial camera intrinsic parameters corresponding to the target image can be determined using a pre-trained camera intrinsic parameter determination model. The camera intrinsic parameter determination model is used to predict the camera intrinsic parameters corresponding to the image based on the corresponding point cloud information. The camera intrinsic parameter determination model can be trained using supervised learning algorithms, unsupervised learning algorithms, and other methods known in the relevant art. The specific training process is not described in detail in this application.

[0151] In a specific embodiment, step S131 can be implemented according to the following steps S131a.

[0152] Step S131a: Based on the principle of minimizing the difference between the converted pixel position and the corresponding actual pixel position, the initial camera intrinsic parameters corresponding to the target image are estimated, where the converted pixel position is the corresponding pixel position obtained by converting the initial point cloud information corresponding to each pixel point of the target image according to the estimated initial camera intrinsic parameters, and the actual pixel position is the corresponding actual pixel position in the target image.

[0153] Camera intrinsic parameters can include the camera's focal length, as well as other parameters. Focal length is a key parameter for converting image pixels into a point cloud. The initial camera intrinsic parameters corresponding to the target image are the camera's intrinsic parameters when the target image was captured.

[0154] In a specific embodiment, the initial point cloud information corresponding to the pixel points of the target image may include x Dimensional initial point cloud information, y Dimensional initial point cloud information and z The three-dimensional initial point cloud information represents the coordinate information of the three dimensions of the point cloud. Correspondingly, the converted pixel position can be obtained as follows: x Dimensional initial point cloud information and y The initial point cloud information of dimensions is respectively y Normalization is performed on the dimensional initial point cloud information to obtain normalized initial point cloud information corresponding to each pixel of the target image; the normalized initial point cloud information is converted to converted pixel positions according to the estimated initial camera intrinsic parameters. Normalization can make the determined camera intrinsic parameters more accurate.

[0155] Specifically, step S131a can predict the initial camera intrinsic parameters according to the following formula (1).

[0156] (1)

[0157] Among them, i, j represent the pixel coordinates of the target image, W, H represent the length and width of the target image, Represents the nth point cloud prediction result of the target image, k represents the coordinate dimension, k is 1, 2, 3 represents the x, y, z dimensions respectively, f is the conversion function that converts the point cloud to pixel positions, f The function contains the camera internal parameters, It satisfies the minimization requirement on the right side of formula (1) f Function, because f The function contains the camera internal parameters, so, The camera intrinsic parameters in are the estimated initial camera intrinsic parameters.

[0158] Optionally, normalization processing may not be performed during the determination of the conversion pixel position, but the initial point cloud information corresponding to the target image may be directly converted into the conversion pixel position according to the estimated initial camera intrinsic parameters.

[0159] When the difference between the converted pixel positions and the corresponding actual pixel positions is minimized, it indicates that the camera intrinsic parameters used to determine the converted pixel positions are relatively close to the actual camera intrinsic parameters used when capturing the target image. Based on the principle of minimizing the difference between the converted pixel positions and the corresponding actual pixel positions, this embodiment can quickly and accurately predict the initial camera intrinsic parameters corresponding to the target image, thereby improving the efficiency and accuracy of 3D modeling. Specifically, the process of determining the initial camera intrinsic parameters corresponding to the target image based on the principle of minimizing the difference between the converted pixel positions and the corresponding actual pixel positions can be gradually optimized using a gradient descent method to obtain the parameter values that minimize the difference.

[0160] Step S132: Determine the fine point cloud information corresponding to the target image based on the pixel correspondence, the initial point cloud information corresponding to the target image, and the initial camera intrinsic parameters.

[0161] Specifically, step S132 can be implemented by following steps S132a to S132b.

[0162] Step S132a: estimating initial camera extrinsic parameters corresponding to the target image based on the pixel correspondence relationship and the initial point cloud information corresponding to the target image.

[0163] The initial camera extrinsics corresponding to the target image are the extrinsic parameters of the camera used when capturing the target image. These parameters may include camera pose and zoom factor. Specifically, the initial camera extrinsics corresponding to the target image can be obtained by inputting the pixel correspondences and the initial point cloud information corresponding to the target image into a pre-trained camera extrinsic parameter determination model.

[0164] Optionally, step S132a can also estimate the initial camera extrinsic parameters corresponding to the target image according to the following steps: based on the principle of minimizing the gap between the matching world point cloud information, estimate the initial camera extrinsic parameters corresponding to the target image, and the matching world point cloud information is the initial point cloud information of the pixel points corresponding to the two images in the image pair, which is the point cloud information after being converted to the world coordinate system according to the estimated initial camera extrinsic parameters.

[0165] When the difference between the initial point cloud information of the pixel points corresponding to the two images in the image pair after being converted to the world coordinate system according to the estimated initial camera extrinsic parameters is relatively small, it means that the estimated camera extrinsic parameters are relatively close to the camera extrinsic parameters of the target image. This embodiment can quickly and accurately predict the initial camera extrinsic parameters corresponding to the target image based on the principle of minimizing the gap between the matching world point cloud information, thereby improving the efficiency and accuracy of three-dimensional modeling.

[0166] Specifically, the initial camera extrinsic parameters corresponding to the target image can be estimated according to the following steps: determining the weight corresponding to each pixel in the target image, wherein the weight is used to represent the importance of the pixel for displaying the target object, and the weights of the corresponding pixels in the two images in the image pair are the same; determining the product of the gap between the matching world point cloud information and the corresponding weight as the first weighted gap for each pair of matching pixels; estimating the initial camera extrinsic parameters corresponding to the target image based on the principle of minimizing the sum of the first weighted gaps corresponding to the image pairs. The weights of different pixels are usually different. For example, for an image of a sofa, the weights of the pixels at the edge of the sofa shape are higher than the weights of the pixels in the middle of the sofa, and the weights of the pixels in the blank area of the image are lower than the weights of the pixels with the sofa. The weights of each pixel can be determined by a pre-trained weight determination model. By setting the weights of each pixel, this embodiment can make it possible to more focus on the pixels of important parts of the reference target object when determining the camera extrinsic parameters, thereby improving computational efficiency and accuracy.

[0167] Specifically, the initial camera extrinsic parameters corresponding to the target image may be estimated according to the following formula (2).

[0168] (2)

[0169] in, 、 They are the initial camera pose and initial camera scaling coefficient corresponding to the estimated target image, that is, P and , P and are the camera pose and scaling factor corresponding to the target image, respectively. n, m ) is an image pair, is the pixel matching information of the image pair, 、 The initial point cloud corresponding to the pixel points of the two images of the image pair is converted to the point cloud of the world coordinate system according to the camera pose and camera zoom factor. q c is the weight.

[0170] Based on the principle of minimizing the gap between the matching world point cloud information, the process of estimating the initial camera extrinsic parameters corresponding to the target image can be gradually optimized by the gradient descent method to obtain the parameter value when it is minimized.

[0171] Step S132b: Determine the fine point cloud information corresponding to the target image based on the initial camera extrinsic parameters, the pixel correspondence, the initial point cloud information corresponding to the target image, and the initial camera intrinsic parameters.

[0172] This embodiment adds camera extrinsics corresponding to the target image, which can more accurately and efficiently determine detailed point cloud information, thereby more accurately generating a three-dimensional model corresponding to the target object.

[0173] Specifically, the initial camera extrinsic parameters, the pixel correspondence, the initial point cloud information corresponding to the target image, and the initial camera intrinsic parameters can be input into a pre-trained point cloud determination model to obtain fine point cloud information corresponding to the target image.

[0174] In a specific embodiment, step S132b can also be implemented according to the following steps: based on the principle of minimizing the difference between the actual pixel position of a pixel point in one image in the image pair and the converted pixel position corresponding to the matching pixel point, determine the fine point cloud information corresponding to the target image, and the matching pixel point is the matching pixel point in the other image in the image pair.

[0175] The converted pixel position corresponding to the matched pixel point is determined by following steps A to B:

[0176] Step A: Converting the initial point cloud information corresponding to the matching pixel points into world point cloud information based on the initial camera extrinsic parameters, the pixel correspondence relationship, and the initial camera intrinsic parameters.

[0177] Step B: Convert the world point cloud information corresponding to the matched pixel point to a two-dimensional image position to obtain a converted pixel position corresponding to the matched pixel point.

[0178] When the difference between the actual pixel position of a pixel point in one image in the image pair and the converted pixel position corresponding to the matching pixel point is relatively small, it means that the point cloud used to predict the converted pixel position can well reflect the actual point cloud of the object displayed in the target image. This embodiment can efficiently and accurately determine the point cloud corresponding to the target image by minimizing the difference between the actual pixel position of a pixel point in one image in the image pair and the converted pixel position corresponding to the matching pixel point.

[0179] Specifically, step S132b can be implemented according to the following steps a to c.

[0180] Step a: Determine the weight corresponding to each pixel in the target image, where the weight is used to represent the importance of the pixel for displaying the target object. The weights of the corresponding pixels in the two images in the image pair are the same.

[0181] Step b: Determine the product of the difference between the actual pixel position of a pixel point in one image of the image pair and the converted pixel position corresponding to the matched pixel point and the corresponding weight as the second weighted difference for each pair of matched pixels.

[0182] Step c: determining the fine point cloud information corresponding to the target image based on the principle of minimizing the sum of the second weighted differences corresponding to the image pair.

[0183] The specific content of the weight can be referred to above and will not be described in detail here. By setting the weight of each pixel point, this embodiment can make the pixel points of important parts of the reference target object more focused when determining detailed point cloud information, thereby improving calculation efficiency and accuracy.

[0184] Specifically, step S132b may determine the fine point cloud information corresponding to the target image according to the following formula (3).

[0185] (3)

[0186] ρ is the loss function, ( n,m ) is an image pair, M n,m is the pixel matching information of the image pair, x n c 、x m c The initial point cloud corresponding to the pixel points of the two images of the image pair is converted to the point cloud of the world coordinate system according to the camera pose and camera zoom factor. q c is the weight, π n 、π m is an image n、 The projection function of the three-dimensional point cloud of image m to the two-dimensional projection coordinates, which contains the camera intrinsic parameters corresponding to the image. y n c 、y m c are the actual pixel positions of the matching pixels in the n and m images of the image pair, Z * , K * 、 、 They are respectively the fine point cloud information corresponding to the estimated target image, the camera internal parameters, the camera pose, and the camera scaling factor, that is, Z, K, P, and , q c is the weight.

[0187] Formula (3) can be used to quickly and accurately determine the fine point cloud information, and at the same time, the accurate camera intrinsic and extrinsic parameters can also be determined by formula (3).

[0188] Based on the principle of minimizing the difference between the actual pixel position of a pixel point in one image in the image pair and the converted pixel position corresponding to the matching pixel point, the process of determining the fine point cloud information corresponding to the target image can be gradually optimized by the gradient descent method to obtain the parameter value when minimizing.

[0189] Step S140: creating a three-dimensional model corresponding to the target object according to the fine point cloud information corresponding to each of the target images.

[0190] Specifically, the fine point cloud information corresponding to each of the target images can be converted into a world coordinate system to obtain the world point cloud information corresponding to each target image, and then the world point cloud information corresponding to each target image can be spliced to obtain the dense point cloud information corresponding to the target object, and a three-dimensional model corresponding to the target object can be created based on the dense point cloud information corresponding to the target object.

[0191] In one specific embodiment, step S132b can be implemented as follows: based on the principle of minimizing the difference between the actual pixel position of a pixel point in one image of the image pair and the converted pixel position corresponding to the matching pixel point, the fine point cloud information corresponding to the target image and the precise camera intrinsic parameters corresponding to the target image are determined. In other words, in the process of minimizing the difference between the actual pixel position and the converted pixel position corresponding to the matching pixel point, the fine point cloud information and the precise camera intrinsic parameters are simultaneously determined. As described in formula (3), in the process of determining the fine point cloud information using formula (3), the precise camera intrinsic parameters can also be determined.

[0192] Correspondingly, before step S140, the following steps S140a to S140b may also be included.

[0193] Step S140a: Determine a global normal vector corresponding to the target image based on the precise camera intrinsic parameters and the fine point cloud information corresponding to the target image.

[0194] Since step S132b has obtained the depth map corresponding to each target image after global consistency constraint optimization, and camera intrinsics , where the depth map is the fine point cloud information corresponding to the target image. According to the depth map and the camera intrinsic parameters, the global normal vector with global consistency can be easily obtained. , where H represents the formula for converting the depth map to the normal vector map.

[0195] Step S140b: Determine the accurate normal vector corresponding to the target image according to the global normal vector corresponding to the target image.

[0196] Specifically, the global normal vector may be determined as the accurate normal vector corresponding to the target image, or the global normal vector may be optimized to obtain the accurate normal vector.

[0197] Because the global normal vector is converted from a depth map and lacks detail accuracy, in one embodiment, the pre-trained image processing model, while determining the pixel correspondences and the initial point cloud information, also determines a predicted normal vector corresponding to each image pair. In other words, the image processing model can also output a predicted normal vector corresponding to each image pair. Because the predicted normal vector references both images in the image pair from different perspectives, it has higher detail accuracy.

[0198] Since the predicted normal vectors corresponding to the image pairs are determined, assuming that The images involved in the composition image pairs, then the image has a total of Predicted normal vectors: .

[0199] Step S140b may be specifically implemented as follows: determining the accurate normal vector corresponding to the target image based on the global normal vector corresponding to the target image and the corresponding predicted normal vectors.

[0200] This embodiment combines the global normal vector with each predicted normal vector to obtain an accurate normal vector. This accurate normal vector achieves both global and detailed accuracy, resulting in higher accuracy. Furthermore, this embodiment determines the normal vector based on two images from different perspectives in an image pair. The determined normal vector effectively avoids perspective discrimination and contains more accurate spatial geometry information, thereby making the determined normal vector more accurate and, in turn, the generated 3D model more accurate.

[0201] For example, the vector with the highest similarity to the global normal vector among the predicted normal vectors corresponding to the target image may be determined as the accurate normal vector corresponding to the target image.

[0202] Optionally, the accurate normal vector corresponding to the target image may be determined according to the following steps 1 to 3.

[0203] Step 1: Calculate the error between each predicted normal vector corresponding to the target image and the corresponding global normal vector.

[0204] The error between the predicted normal vector and the global normal vector is the distance between the predicted normal vector and the global normal vector. The larger the error, the greater the difference between the two vectors.

[0205] Specifically, step 1 may calculate, for each pixel point of the target image, the error between each predicted normal vector corresponding to the target image and the corresponding global normal vector.

[0206] Step 2: excluding pixel points whose errors with the global normal vector are greater than a preset error from the predicted normal vectors corresponding to the target image, to obtain excluded predicted normal vectors corresponding to the target image after excluding the pixel points.

[0207] The predicted normal vector includes the predicted normal vectors of each pixel point of the target image. After excluding some pixels in the predicted normal vector, the predicted normal vector only includes the predicted normal vectors corresponding to some pixels.

[0208] Step 3: Determine the accurate normal vector corresponding to the target image according to the mean of the excluded predicted normal vectors corresponding to the target image.

[0209] The mean value of the predicted normal vectors after elimination can be determined as the accurate normal vector corresponding to each pixel point of the target image.

[0210] In this embodiment, the vectors of each predicted normal vector that have excessively large errors with the global normal vector are deleted. The deleted predicted normal vectors are vectors with poor globality, and the remaining predicted normal vectors after exclusion are vectors with relatively high accuracy and good globality. The mean of each predicted normal vector after exclusion is then determined as the accurate normal vector corresponding to the target image. That is, this application optimizes the predicted normal vector through a global consistency optimization strategy, and can obtain a normal vector with better globality and more accurate details.

[0211] Optionally, when a pixel point of the target image is excluded from each of the predicted normal vectors after exclusion, the normal vector of the pixel point in the global normal vector corresponding to the target image can be determined as the accurate normal vector corresponding to the pixel point to ensure that each pixel point can obtain the corresponding accurate normal vector.

[0212] Correspondingly, the above step S140 can be implemented according to the following step S141.

[0213] Step S141: creating a three-dimensional model corresponding to the target object according to the fine point cloud information corresponding to each target image and the accurate normal vector corresponding to each target image.

[0214] Since a normal vector is a vector defined on the surface of a three-dimensional object, its direction is perpendicular to that surface and can be used to describe the shape and orientation of the surface. In image processing and computer graphics, normal vector information is crucial for tasks such as lighting modeling, shading, texture mapping, surface restoration, and 3D reconstruction. This embodiment uses accurate normal vectors to create a more accurate three-dimensional model of the target object.

[0215] Specifically, step S141 may create a three-dimensional model corresponding to the target object based on the dense point cloud information and precise normal vector corresponding to the target object.

[0216] In one embodiment, a pre-trained 3D model generation model can be used to generate a 3D model corresponding to a target object. Specifically, dense point cloud information and precise normal vectors corresponding to the target object can be input into the 3D model generation model to obtain a 3D model corresponding to the target object. Given the dense point cloud and normal vectors of the object, the intelligent model can accurately and quickly generate the corresponding 3D model.

[0217] The following is an illustrative example of the method for creating a three-dimensional model provided by this application. Figure 3 As shown, the method for creating a 3D model provided in this example includes the following steps:

[0218] Step 1: Obtain multiple target images corresponding to the target object for which the 3D model is to be created.

[0219] The multiple target images are multiple captured images, and the multiple captured images show the target object from multiple perspectives.

[0220] Step 2: Input multiple target images into a pre-trained image coding model to obtain image coding features corresponding to each target image.

[0221] Step 3: According to the similarity between different image coding features, the multiple target images are grouped into multiple groups of image pairs.

[0222] Assume that the total number of target images is , according to the similarity between different image encoding features, each target image is obtained with other The similarity metric between target images is used to quantify the similarity between each target image. The top similarity ranking among target images The target images form an image pair, and the total To ensure the balance between the accuracy and time consumption of the algorithm, we can control Not exceeding the second preset number .

[0223] Step 4: Input the image encoding features of the image pair sequence into a pre-trained decoding model to obtain the pixel correspondence between the two images in the image pair, the initial point cloud information corresponding to each target image, and the predicted normal vector of each target image.

[0224] Step 5: Determine the dense point cloud information and camera parameters of the target object based on the pixel correspondence between the two images in the image pair and the initial point cloud information corresponding to each of the target images.

[0225] Specifically, the dense point cloud information, camera extrinsic parameters and intrinsic parameters of the target object can be determined by the gradient descent method.

[0226] Step 6: Determine the global normal vector corresponding to the target image based on the camera intrinsic parameters and the fine point cloud information corresponding to the target image, and combine the global normal vector and the predicted normal vector to obtain the accurate normal vector corresponding to each target image.

[0227] Step 7: Create a 3D model of the target object based on the accurate normal vector corresponding to each target image and the dense point cloud information of the target object.

[0228] This example is an illustrative description of an embodiment of the present invention. The detailed execution process of each step can be referred to the description above and will not be repeated here.

[0229] The method for creating a three-dimensional model of an object provided in an embodiment of the present application obtains multiple target images corresponding to the target object for which the three-dimensional model is to be created, and the multiple target images are used to display the target object from multiple perspectives. Then, based on a pre-trained image processing model, the pixel correspondence between two images in an image pair corresponding to the multiple target images and the initial point cloud information corresponding to each target image are determined. The image pair includes two images from the multiple target images. The two images in a group of image pairs can display various parts of the object from different perspectives, so that each group of image pairs can better accurately reflect the information of various parts of the object from a stereoscopic geometric perspective. The pixel correspondence is used to indicate the pixel position of the object part corresponding to the pixel position of one image in the image pair and the corresponding pixel position in the other image, that is, the pixel correspondence can associate and correspond the pixel points belonging to the same object part in the two images. The pixel correspondence can well reflect the position information of the same object part at different perspectives, so that the pixel correspondence can well reflect the stereoscopic geometric information of the object. The above-mentioned image processing model is first trained, and the image processing model can determine the pixel correspondence between the two images in each group of image pairs. The pixel correspondence relationship and the initial point cloud information corresponding to the target image are used to quickly determine the initial point cloud information of each target image and the pixel correspondence relationship through a pre-trained image processing model. The initial point cloud information reflects the depth image corresponding to each target image, and the initial point cloud information can reflect the geometric information of the object corresponding to a single image. Then, based on the initial point cloud information of the target image and the pixel correspondence relationship, the fine point cloud information corresponding to the target image is determined. Since the pixel correspondence relationship can well reflect the position information of the same object part at different perspectives, the initial point cloud information is refined through the pixel correspondence relationship, so that the refined point cloud information after refinement can have higher accuracy in three-dimensional space, that is, the fine point cloud information can more accurately represent each part of the target object in space. Then, based on the fine point cloud information corresponding to each target image, a three-dimensional model corresponding to the target object is determined. Since the fine point cloud information can more accurately represent each part of the target object in space, the three-dimensional model corresponding to the target object created based on the fine point cloud information corresponding to each target image is more accurate in three-dimensional space and more closely matches the actual target object.

[0230] The solution provided by the present application first determines the initial point cloud information corresponding to each target image efficiently and quickly through an image processing model. Since the initial point cloud information is obtained based on a single target image, the initial point cloud information may lack an accurate reflection of the target object in three-dimensional space. The present application then adjusts the initial point cloud information through the pixel correspondence corresponding to each group of images that can well reflect the three-dimensional geometric information of the object, and obtains fine point cloud information that can more accurately represent various parts of the target object in space, so that the three-dimensional model created according to the fine point cloud information corresponding to each of the target images is more matched with the actual target object and has higher accuracy.

[0231] It can be seen that the solution provided by this application can automatically generate a three-dimensional model corresponding to the target object through a computer, without the need for manual operation of the three-dimensional modeling software. Therefore, it is possible to efficiently and quickly create a three-dimensional model of an object based on the image of the object, without the need for the user to have high professional skills in three-dimensional modeling. The created three-dimensional model is more accurate and more consistent with the target object. In addition, because the image processing model is trained using a large number of training samples, the training process usually includes both good-quality images and poor-quality images. Therefore, the trained image processing model has a strong prior, which makes the image processing model predict a larger number of pixel matches. It can also perform good prediction processing for poor image quality such as low texture, blur, low resolution, and poor lighting conditions, and obtain an accurate three-dimensional model. The initial point cloud information of each target image predicted by the pre-trained image processing model itself is also relatively high in accuracy and can contain rich geometric information, creating conditions for obtaining more accurate and detailed point cloud information in the future, so that an accurate three-dimensional model can be obtained in the end. Example 2

[0232] The second embodiment of the present application further provides a method for creating a three-dimensional model of an object, which includes the following steps S210 to S270.

[0233] Step S210: Acquire multiple target images corresponding to the target object for which a three-dimensional model is to be created, wherein the multiple target images are used to display the target object from multiple perspectives;

[0234] Step S220: determining a plurality of image pairs consisting of the plurality of target images, each of the image pairs including two different target images;

[0235] Step S230: Determining, based on a pre-trained image processing model, a pixel correspondence relationship between the two images in the image pair, initial point cloud information corresponding to each target image, and a predicted normal vector corresponding to each image pair, wherein the pixel correspondence relationship indicates the pixel position of the object part corresponding to the pixel position in one image of the image pair and the corresponding pixel position in the other image;

[0236] Step S240: determining the fine point cloud information corresponding to the target image and the precise camera intrinsic parameters corresponding to the target image based on the initial point cloud information of the target image and the pixel correspondence relationship;

[0237] Step S250: determining a global normal vector corresponding to the target image based on the precise camera intrinsic parameters and the fine point cloud information corresponding to the target image;

[0238] Step S260: determining an accurate normal vector corresponding to the target image according to the global normal vector corresponding to the target image and the corresponding predicted normal vectors;

[0239] Step S270: creating a three-dimensional model corresponding to the target object according to the fine point cloud information corresponding to each of the target images and the accurate normal vector corresponding to each of the target images.

[0240] The various execution processes of this embodiment are similar to those of the first embodiment. For the details of the relevant technical features and the effects achieved, please refer to the corresponding description of the embodiment of the method for creating a three-dimensional model of an object provided in the above-mentioned first embodiment. Example 3

[0241] The third embodiment of the present application further provides a method for creating a three-dimensional model of an object, which is applied to a server and includes the following steps S310 to S330.

[0242] Step S310: Acquire multiple target images sent by the client for displaying the target object, wherein the multiple target images are used to display the target object from multiple perspectives;

[0243] Step S320: creating a three-dimensional model corresponding to the target object by using the method for creating a three-dimensional model of an object described in any one of the first embodiment or the second embodiment;

[0244] Step S330: Send the three-dimensional model to the client, so that the client displays the three-dimensional model.

[0245] This embodiment is an application scenario embodiment of the method for creating a three-dimensional model of an object provided in the above embodiment. Its specific modeling process is basically similar to that of the method embodiment, so the description is relatively simple. For the details of the relevant technical features and the effects achieved, please refer to the corresponding description of the embodiment of the method for creating a three-dimensional model of an object provided above. Example 4

[0246] The fourth embodiment of the present application also provides a device for creating a three-dimensional object model corresponding to the method embodiment for creating a three-dimensional object model provided in the first embodiment. Since the device embodiment is basically similar to the method embodiment, the description is relatively simple. For details of the relevant technical features and the effects achieved, please refer to the corresponding description of the method embodiment for creating a three-dimensional object model provided above. Figure 4 As shown, the device for creating a three-dimensional model of an object provided in this embodiment includes:

[0247] An acquisition unit 410 is configured to acquire a plurality of target images corresponding to a target object for which a three-dimensional model is to be created, wherein the plurality of target images are used to display the target object from multiple perspectives;

[0248] a pairing unit 420, configured to determine a plurality of image pairs consisting of the plurality of target images, each of the image pairs including two different target images;

[0249] a determination unit 430 configured to determine, based on a pre-trained image processing model, a pixel correspondence between the two images in the image pair and initial point cloud information corresponding to each of the target images, wherein the pixel correspondence indicates the pixel position of an object part corresponding to a pixel position in one image of the image pair and the corresponding pixel position in the other image; and determine, based on the initial point cloud information of the target image and the pixel correspondence, fine point cloud information corresponding to the target image;

[0250] The modeling unit 440 is configured to create a three-dimensional model corresponding to the target object based on the fine point cloud information corresponding to each of the target images. Example 5

[0251] The third embodiment of the present application also provides an electronic device embodiment corresponding to the method for creating a three-dimensional model of an object provided in the first embodiment. The following description of the electronic device embodiment is merely illustrative. The electronic device embodiment is as follows:

[0252] Please refer to Figure 5 Understand the above electronic devices, Figure 5 Schematic diagram of an electronic device. The electronic device provided in this embodiment includes: a processor 1001, a memory 1002, a communication bus 1003, and a communication interface 1004;

[0253] The memory 1002 is used to store computer instructions for data processing. When the computer instructions are read and executed by the processor 1001, the following steps are performed:

[0254] Acquire multiple target images corresponding to a target object for which a three-dimensional model is to be created, wherein the multiple target images are used to display the target object from multiple perspectives;

[0255] Determining a plurality of groups of image pairs consisting of the plurality of target images, each group of the image pairs including two different target images;

[0256] Determining a pixel correspondence between the two images in the image pair and initial point cloud information corresponding to each of the target images based on a pre-trained image processing model, wherein the pixel correspondence indicates the pixel position of the object part corresponding to the pixel position in one image of the image pair and the corresponding pixel position in the other image;

[0257] Determining fine point cloud information corresponding to the target image based on the initial point cloud information of the target image and the pixel correspondence relationship;

[0258] A three-dimensional model corresponding to the target object is created based on the fine point cloud information corresponding to each of the target images.

[0259] The sixth embodiment of the present application also provides a computer-readable storage medium for implementing the method described in the first embodiment. The computer-readable storage medium embodiment provided in this application is described in a relatively simple manner. For relevant parts, please refer to the corresponding description of the above method embodiment. The embodiment described below is merely illustrative.

[0260] The computer-readable storage medium provided in this embodiment stores computer instructions, which, when executed by a processor, implement the following steps:

[0261] Acquire multiple target images corresponding to a target object for which a three-dimensional model is to be created, wherein the multiple target images are used to display the target object from multiple perspectives;

[0262] Determining a plurality of groups of image pairs consisting of the plurality of target images, each group of the image pairs including two different target images;

[0263] Determining a pixel correspondence between the two images in the image pair and initial point cloud information corresponding to each of the target images based on a pre-trained image processing model, wherein the pixel correspondence indicates the pixel position of the object part corresponding to the pixel position in one image of the image pair and the corresponding pixel position in the other image;

[0264] Determining fine point cloud information corresponding to the target image based on the initial point cloud information of the target image and the pixel correspondence relationship;

[0265] A three-dimensional model corresponding to the target object is created based on the fine point cloud information corresponding to each of the target images.

[0266] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.

[0267] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.

[0268] 1. Computer-readable media, including both permanent and non-permanent, removable and non-removable media, can be implemented using any method or technology to store information. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change RAM (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include non-transitory computer-readable media, such as modulated data signals and carrier waves.

[0269] 2. Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROMs, optical storage, etc.) containing computer-usable program code.

[0270] Although the present application is disclosed as above with the preferred embodiments, it is not intended to limit the present application. Any person skilled in the art may make possible changes and modifications without departing from the spirit and scope of the present application. Therefore, the scope of protection of the present application shall be based on the scope defined by the claims of the present application.

Claims

1. A method for creating a three-dimensional model of an object, characterized in that: The method comprises: Acquire multiple target images corresponding to a target object for which a three-dimensional model is to be created, wherein the multiple target images are used to display the target object from multiple perspectives; Determining a plurality of groups of image pairs consisting of the plurality of target images, each group of the image pairs including two different target images; Determining a pixel correspondence between the two images in the image pair and initial point cloud information corresponding to each of the target images based on a pre-trained image processing model, wherein the pixel correspondence indicates the pixel position of the object part corresponding to the pixel position in one image of the image pair and the corresponding pixel position in the other image; Determining fine point cloud information corresponding to the target image based on the initial point cloud information of the target image and the pixel correspondence, including: estimating initial camera intrinsic parameters corresponding to the target image based on the initial point cloud information of the target image; determining fine point cloud information corresponding to the target image based on the pixel correspondence, the initial point cloud information corresponding to the target image, and the initial camera intrinsic parameters; A three-dimensional model corresponding to the target object is created based on the fine point cloud information corresponding to each of the target images.

2. The method for creating a three-dimensional model of an object according to claim 1, wherein: The estimating the initial camera intrinsic parameters corresponding to the target image based on the initial point cloud information of the target image includes: Based on the principle of minimizing the difference between the converted pixel position and the corresponding actual pixel position, the initial camera intrinsic parameters corresponding to the target image are estimated. The converted pixel position is the corresponding pixel position obtained by converting the initial point cloud information corresponding to each pixel point of the target image according to the estimated initial camera intrinsic parameters. The actual pixel position is the corresponding actual pixel position in the target image.

3. The method for creating a three-dimensional model of an object according to claim 2, wherein: The initial point cloud information corresponding to the pixel points of the target image includes x-dimensional initial point cloud information, y-dimensional initial point cloud information, and z-dimensional initial point cloud information; The converted pixel position is obtained as follows: Normalizing the x-dimensional initial point cloud information and the y-dimensional initial point cloud information of the target image on the z-dimensional initial point cloud information to obtain normalized initial point cloud information corresponding to each pixel of the target image; The normalized initial point cloud information is converted into a converted pixel position according to the estimated initial camera intrinsic parameters.

4. The method for creating a three-dimensional model of an object according to claim 1, wherein: The determining, based on the pixel correspondence, the initial point cloud information corresponding to the target image, and the initial camera intrinsic parameters, of fine point cloud information corresponding to the target image includes: estimating the initial camera extrinsic parameters corresponding to the target image based on the pixel correspondence relationship and the initial point cloud information corresponding to the target image; Based on the initial camera extrinsic parameters, the pixel correspondence, the initial point cloud information corresponding to the target image, and the initial camera intrinsic parameters, fine point cloud information corresponding to the target image is determined.

5. The method for creating a three-dimensional model of an object according to claim 4, wherein: The estimating the initial camera extrinsic parameters corresponding to the target image based on the pixel correspondence relationship and the initial point cloud information corresponding to the target image includes: Based on the principle of minimizing the gap between the matched world point cloud information, the initial camera extrinsic parameters corresponding to the target image are estimated. The matched world point cloud information is the point cloud information of the initial point cloud information of the pixel points corresponding to the two images in the image pair after being converted to the world coordinate system according to the estimated initial camera extrinsic parameters.

6. The method for creating a three-dimensional model of an object according to claim 5, wherein: The estimating the initial camera extrinsic parameters corresponding to the target image based on the principle of minimizing the gap between the matched world point cloud information includes: Determining a weight corresponding to each pixel in the target image, where the weight represents the importance of the pixel in representing the target object, and the weights of corresponding pixels in the two images of the image pair are the same; Determine the product of the gap between the matching world point cloud information and the corresponding weight as the first weighted gap for each pair of matching pixels; The initial camera extrinsic parameters corresponding to the target image are estimated based on the principle of minimizing the sum of the first weighted differences corresponding to the image pair.

7. The method for creating a three-dimensional model of an object according to claim 4, wherein: The determining, based on the initial camera extrinsic parameters, the pixel correspondence, the initial point cloud information corresponding to the target image, and the initial camera intrinsic parameters, of the fine point cloud information corresponding to the target image includes: determining fine point cloud information corresponding to the target image based on the principle of minimizing the difference between the actual pixel position of a pixel point in one image of the image pair and the converted pixel position corresponding to the matched pixel point, where the matched pixel point is the matched pixel point in the other image of the image pair; The converted pixel position corresponding to the matched pixel point is determined by: Converting the initial point cloud information corresponding to the matched pixel points into world point cloud information based on the initial camera extrinsic parameters, the pixel correspondence relationship, and the initial camera intrinsic parameters; The world point cloud information corresponding to the matched pixel point is converted to a two-dimensional image position to obtain a converted pixel position corresponding to the matched pixel point.

8. The method for creating a three-dimensional model of an object according to claim 7, wherein: The determining of the fine point cloud information corresponding to the target image based on the principle of minimizing the difference between the actual pixel position of a pixel point in one image of the image pair and the converted pixel position corresponding to the matched pixel point includes: Determining a weight corresponding to each pixel in the target image, where the weight represents the importance of the pixel in representing the target object, and the weights of corresponding pixels in the two images of the image pair are the same; Determine the product of the difference between the actual pixel position of a pixel point in one image of the image pair and the converted pixel position corresponding to the matched pixel point and the corresponding weight as the second weighted difference for each pair of matched pixels; The fine point cloud information corresponding to the target image is determined based on the principle of minimizing the sum of the second weighted differences corresponding to the image pair.

9. The method for creating a three-dimensional model of an object according to claim 5, wherein: The camera external parameters include at least one of a camera pose and a scaling factor, and the camera internal parameters include a camera focal length.

10. The method for creating a three-dimensional model of an object according to any one of claims 1 to 9, characterized in that: Determining a plurality of image pairs consisting of the plurality of target images includes: According to the similarities between the target images, the multiple target images are grouped into multiple groups of image pairs.

11. The method for creating a three-dimensional model of an object according to claim 10, wherein: The step of grouping the plurality of target images into a plurality of image pairs according to the similarities between the target images comprises: Extracting image coding features corresponding to each of the target images based on a pre-trained image coding model; Calculating the similarity between each target image and other target images based on the image coding features; The multiple target images are grouped into multiple image pairs according to the similarity between each target image and other target images.

12. The method for creating a three-dimensional model of an object according to claim 11, wherein: The step of grouping the plurality of target images into a plurality of image pairs according to the similarity between each target image and other target images comprises: For each target image, images ranked first in a preset number of similarities with the target image among other target images are respectively formed into image pairs with the target image, or images ranked first in a preset similarity threshold with the target image among other target images are respectively formed into image pairs with the target image.

13. The method for creating a three-dimensional model of an object according to claim 12, wherein: The step of grouping the plurality of target images into a plurality of image pairs according to the similarity between each target image and other target images comprises: Based on the principle that the total number of image pairs does not exceed a second preset number, for each target image, other target images that rank top preset number of similarities with the target image are respectively formed into image pairs.

14. The method for creating a three-dimensional model of an object according to claim 11, wherein: The determining of the pixel correspondence between the two images in the image pair and the initial point cloud information corresponding to each target image based on the pre-trained image processing model includes: The image coding features corresponding to each of the target images and the information of the two images constituting the image pair are input into a pre-trained image processing model to obtain the pixel correspondence between the two images in the image pair and the initial point cloud information corresponding to each of the target images.

15. The method for creating a three-dimensional model of an object according to claim 7, wherein: The determining of the fine point cloud information corresponding to the target image based on the principle of minimizing the difference between the actual pixel position of a pixel point in one image of the image pair and the converted pixel position corresponding to the matched pixel point includes: Determine the fine point cloud information corresponding to the target image and the precise camera intrinsic parameters corresponding to the target image based on the principle of minimizing the difference between the actual pixel position of a pixel point in one image of the image pair and the converted pixel position corresponding to the matched pixel point; The method further comprises: Determining a global normal vector corresponding to the target image based on the precise camera intrinsic parameters and the fine point cloud information corresponding to the target image; Determining an accurate normal vector corresponding to the target image based on a global normal vector corresponding to the target image; The step of creating a three-dimensional model corresponding to the target object based on the fine point cloud information corresponding to each target image includes: A three-dimensional model corresponding to the target object is created based on the fine point cloud information corresponding to each target image and the accurate normal vector corresponding to each target image.

16. The method for creating a three-dimensional model of an object according to claim 15, wherein: The pre-trained image processing model is further used to determine the predicted normal vector corresponding to each of the image pairs in the process of determining the pixel correspondence relationship and the initial point cloud information; Determining the accurate normal vector corresponding to the target image according to the global normal vector corresponding to the target image includes: The accurate normal vector corresponding to the target image is determined according to the global normal vector corresponding to the target image and the corresponding predicted normal vectors.

17. The method for creating a three-dimensional model of an object according to claim 16, wherein: Determining the accurate normal vector corresponding to the target image according to the global normal vector corresponding to the target image and the corresponding predicted normal vectors includes: Calculating the error between each predicted normal vector corresponding to the target image and the corresponding global normal vector; Eliminate pixels whose errors with the global normal vector are greater than a preset error from the predicted normal vectors corresponding to the target image, and obtain excluded predicted normal vectors corresponding to the target image after excluding the pixels; The accurate normal vector corresponding to the target image is determined according to the mean value of each excluded predicted normal vector corresponding to the target image.

18. The method for creating a three-dimensional model of an object according to claim 17, wherein: The method further comprises: When a pixel point of the target image is excluded from each of the excluded predicted normal vectors, the normal vector of the pixel point in the global normal vector corresponding to the target image is determined as the accurate normal vector corresponding to the pixel point.

19. The method for creating a three-dimensional model of an object according to any one of claims 1 to 9, characterized in that: The step of creating a three-dimensional model corresponding to the target object based on the fine point cloud information corresponding to each target image includes: The dense point cloud information corresponding to the target object is determined based on the fine point cloud information corresponding to each target image, and a three-dimensional model corresponding to the target object is created based on the dense point cloud information.

20. A method for creating a three-dimensional model of an object, characterized in that: The method comprises: Acquire multiple target images corresponding to a target object for which a three-dimensional model is to be created, wherein the multiple target images are used to display the target object from multiple perspectives; Determining a plurality of groups of image pairs consisting of the plurality of target images, each group of the image pairs including two different target images; Determining, based on a pre-trained image processing model, a pixel correspondence relationship between the two images in the image pair, initial point cloud information corresponding to each target image, and a predicted normal vector corresponding to each image pair, wherein the pixel correspondence relationship indicates the pixel position of an object part corresponding to a pixel position in one image of the image pair and the corresponding pixel position in the other image; Determining, based on the initial point cloud information of the target image and the pixel correspondence, fine point cloud information corresponding to the target image and precise camera intrinsic parameters corresponding to the target image, including: estimating the initial camera intrinsic parameters corresponding to the target image based on the initial point cloud information of the target image; determining, based on the pixel correspondence, the initial point cloud information corresponding to the target image, and the initial camera intrinsic parameters, the fine point cloud information corresponding to the target image and precise camera intrinsic parameters corresponding to the target image; Determining a global normal vector corresponding to the target image based on the precise camera intrinsic parameters and the fine point cloud information corresponding to the target image; Determining an accurate normal vector corresponding to the target image based on a global normal vector corresponding to the target image and corresponding predicted normal vectors; A three-dimensional model corresponding to the target object is created based on the fine point cloud information corresponding to each target image and the accurate normal vector corresponding to each target image.

21. A method for creating a three-dimensional model of an object, characterized in that: Applied to the server, the method includes: Acquire multiple target images sent by a client for displaying a target object, wherein the multiple target images are used to display the target object from multiple perspectives; Creating a three-dimensional model corresponding to the target object by the method for creating a three-dimensional model of an object according to any one of claims 1 to 20; The three-dimensional model is sent to the client, so that the client displays the three-dimensional model.

22. A device for creating a three-dimensional model of an object, characterized in that: The device comprises: an acquisition unit, configured to acquire a plurality of target images corresponding to a target object for which a three-dimensional model is to be created, wherein the plurality of target images are used to display the target object from multiple perspectives; a pairing unit, configured to determine a plurality of image pairs consisting of the plurality of target images, each of the image pairs including two different target images; a determination unit configured to determine, based on a pre-trained image processing model, a pixel correspondence between two images in the image pair and initial point cloud information corresponding to each of the target images, the pixel correspondence being used to indicate a pixel position in one image of the image pair corresponding to an object part in the other image; and determining, based on the initial point cloud information of the target image and the pixel correspondence, fine point cloud information corresponding to the target image, including: estimating initial camera intrinsic parameters corresponding to the target image based on the initial point cloud information of the target image; and determining, based on the pixel correspondence, the initial point cloud information corresponding to the target image, and the initial camera intrinsic parameters, fine point cloud information corresponding to the target image and precise camera intrinsic parameters corresponding to the target image. A modeling unit is used to create a three-dimensional model corresponding to the target object based on the fine point cloud information corresponding to each of the target images.

23. An electronic device, characterized in that: include: a processor, a memory, and computer program instructions stored on the memory and executable on the processor; When the processor executes the computer program instructions, the method according to any one of claims 1 to 21 is implemented.

24. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method according to any one of claims 1 to 21.

Citation Information

Patent Citations

  • Three-dimensional reconstruction method and device for power transmission line gallery based on multi-view technology

    CN114332415A

  • Three-dimensional reconstruction method, display method and electronic equipment

    CN116486008A