Object modeling method and apparatus, electronic device, storage medium, and program product
By generating a hybrid point cloud of the target object and the target plane and using the plane equation for surface reconstruction, the reconstruction quality problem caused by the missing bottom data is solved, and a high-quality three-dimensional model is generated without collecting bottom data.
Patent Information
- Application Number
- PCT/CN2024/138763
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-06
- Filing Date
- 2024-12-12
- Publication Date
- 2025-08-14
AI Technical Summary
In object modeling, the missing bottom data causes the bottom triangle mesh to be uneven, affecting the overall reconstruction quality. Especially for objects that are not easy to flip and/or are not important or do not need to collect bottom data, it is difficult for the prior art to effectively reconstruct a flat bottom.
By obtaining the image of the target object at different perspectives, a mixed point cloud of the target object and the target plane is generated, a mixed point cloud is used to generate a plane equation of the target plane, and surface reconstruction is carried out based on the plane equation, a three-dimensional model of the target object is generated, including mask extraction and the use of plane loss functions for plane constraints.
In the absence of bottom data, a flat bottom can be reconstructed, reducing the difficulty of object data acquisition, and improving the quality and accuracy of the three-dimensional model.
Smart Images

Figure CN2024138763_14082025_PF_FP_ABST
Abstract
Description
Object modeling method, device, electronic device, storage medium and program product
[0001] This application claims priority to Chinese Patent Application No. 202410172172.0 filed on February 6, 2024, and the contents of the above-mentioned Chinese patent application disclosure are hereby incorporated by reference in their entirety as a part of this application. Technical Field
[0002] The embodiments of the present disclosure relate to an object modeling method, apparatus, electronic device, storage medium, and program product. Background Art
[0003] In applications such as object modeling, users can take pictures of objects of interest from different perspectives. The application can generate a three-dimensional model of the object by algorithmically reconstructing the pictures taken by the user. Summary of the Invention
[0004] The embodiments of the present disclosure provide an object modeling method, apparatus, electronic device, storage medium, and program product to reconstruct a flat bottom for an object when bottom data is missing, thereby improving the quality of the constructed three-dimensional model.
[0005] In a first aspect, an embodiment of the present disclosure provides an object modeling method, comprising:
[0006] Acquire an object image of a target object at different viewing angles, wherein the object image is obtained by capturing an image of the target object placed on a target plane;
[0007] Processing the object image to obtain a mixed point cloud of the target object and the target plane;
[0008] generating a plane equation of the target plane according to the mixed point cloud;
[0009] The surface of the target object is reconstructed based on the plane equation to obtain a three-dimensional model of the target object.
[0010] In a second aspect, an embodiment of the present disclosure further provides an object modeling device, comprising:
[0011] An image acquisition module is used to acquire an object image of a target object at different viewing angles, wherein the object image is obtained by acquiring an image of the target object placed on a target plane;
[0012] An image processing module, configured to process the object image to obtain a mixed point cloud of the target object and the target plane;
[0013] an equation generating module, configured to generate a plane equation of the target plane according to the mixed point cloud;
[0014] A surface reconstruction module is used to reconstruct the surface of the target object based on the plane equation to obtain a three-dimensional model of the target object.
[0015] In a third aspect, an embodiment of the present disclosure further provides an electronic device, including:
[0016] one or more processors;
[0017] a memory for storing one or more programs,
[0018] When the one or more programs are executed by the one or more processors, the one or more processors implement the object modeling method as described in the embodiment of the present disclosure.
[0019] In a fourth aspect, an embodiment of the present disclosure further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the object modeling method as described in the embodiment of the present disclosure.
[0020] In a fifth aspect, an embodiment of the present disclosure further provides a computer program product, which, when executed by a computer, enables the computer to implement the object modeling method as described in the embodiment of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that the originals and elements are not necessarily drawn to scale.
[0022] FIG1 is a schematic flow chart of an object modeling method provided by an embodiment of the present disclosure;
[0023] FIG2 is a schematic diagram of a flow chart of another object modeling method provided by an embodiment of the present disclosure;
[0024] FIG3 is a schematic flow chart of an optional object modeling method provided in an embodiment of the present disclosure;
[0025] FIG4 is a structural block diagram of an object modeling device provided by an embodiment of the present disclosure;
[0026] FIG5 is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION
[0027] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0028] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.
[0029] As used herein, the term "including" and its variations are open-ended, i.e., "including but not limited to." The term "based on" means "based, at least in part, on." The term "one embodiment" means "at least one embodiment," the term "another embodiment" means "at least one additional embodiment," and the term "some embodiments" means "at least some embodiments." Other terms are defined in the following description.
[0030] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0031] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".
[0032] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.
[0033] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0034] For example, in response to a user's active request, a prompt message is sent to the user to clearly inform the user that the operation requested will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operations of the disclosed technical solution based on the prompt message.
[0035] As an optional but non-limiting implementation, in response to receiving a user's active request, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. Furthermore, the pop-up window may also contain a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.
[0036] It is understandable that the above notification and user authorization process are merely illustrative and do not limit the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.
[0037] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) must comply with the requirements of relevant laws, regulations and relevant provisions.
[0038] During the capture process, if data from the bottom of an object is needed, the object must be flipped over. This requires the object to be made of a relatively rigid material, meaning that the object itself cannot undergo non-rigid deformation due to flipping. Furthermore, for ease of operation, the object must be small and heavy, and not something that should be inverted, such as a refrigerator. The collection of bottom-level data is subject to numerous limitations.
[0039] However, if the bottom data is not collected, the bottom triangle mesh will be uneven during neural surface reconstruction (NSR), affecting the overall reconstruction quality. Therefore, for objects that are difficult to flip and / or whose bottom information is unimportant or does not require bottom data collection, how to perform incomplete reconstruction and reconstruct a flat bottom without bottom data is an urgent problem that needs to be solved.
[0040] FIG1 is a flow chart of an object modeling method provided by an embodiment of the present disclosure. The method can be executed by an object modeling device, wherein the device can be implemented by software and / or hardware and can be configured in an electronic device, for example, a mobile phone or tablet computer. The object modeling method provided by an embodiment of the present disclosure is suitable for scenarios where object modeling is performed, and is particularly suitable for scenarios where object modeling is performed when object bottom data is missing. As shown in FIG1 , the object modeling method provided by this embodiment may include:
[0041] S101 : Acquire object images of a target object at different viewing angles, wherein the object images are obtained by capturing images of the target object placed on a target plane.
[0042] The target object can be considered as the object to be modeled. The object image can be an image of the target object, such as a picture of the target object, which can be obtained by capturing images of the target object placed on a target plane at different viewing angles. The different viewing angles may not include the bottom view of the target object, and accordingly, the acquired object image may not include the bottom image of the target object. The target plane can be understood as the plane on which the target object is placed during image capture, such as the ground or tabletop on which the target object is placed.
[0043] In this embodiment, object images of the target object at different viewing angles may be acquired, so as to subsequently model the target object based on the acquired object images.
[0044] There is no limit to the way of acquiring the object image of the target object. For example, the object image of the target object placed on the target plane can be acquired at different viewing angles by an image acquisition device (such as a camera, etc.). Alternatively, the user can acquire the object image of the target object placed on the target plane that he wants to model at different viewing angles in advance, and can upload the acquired object image after the acquisition is completed; accordingly, the current application can acquire the object image of the target object uploaded by the user. Here, in the process of acquiring the object image, there is no need to flip the target object to acquire the bottom data of the target object, such as there is no need to acquire the bottom image of the target object.
[0045] S102: Process the object image to obtain a mixed point cloud of the target object and the target plane.
[0046] The mixed point cloud of the target object and the target plane can be understood as including both points located on the target object and points located on the target plane.
[0047] In this embodiment, the object image of the target object under at least part of the viewing angle may include an image of the target plane. For example, when capturing an image of the target object under certain viewing angles, the target plane on which the target object is placed may also be captured.
[0048] Thus, after acquiring images of the target object at different perspectives, the acquired object images can be processed to generate a hybrid point cloud of the target object and the target plane. There are no limitations on the method for processing the object images to generate the hybrid point cloud. For example, existing point cloud generation methods can be used to generate a hybrid point cloud of the target object and the target plane based on the object images of the target object.
[0049] In some embodiments, the obtained mixed point cloud of the target object and the target plane can be a Structure From Motion (SFM) sparse point cloud. In this case, after acquiring the object images of the target object at different perspectives, the acquired object images can be processed by the SFM method to obtain the SFM sparse point cloud corresponding to the acquired object images as the mixed point cloud of the target object and the target plane. At this time, optionally, the processing of the object image to obtain the mixed point cloud of the target object and the target plane includes: processing the object image using the Structure From Motion SFM method to obtain the SFM sparse point cloud of the target object and the target plane as the mixed point cloud of the target object and the target plane. Among them, the existing SFM method can be used to process the acquired object image, and the object image processing process will not be described in detail here.
[0050] S103: Generate a plane equation of the target plane according to the mixed point cloud.
[0051] In this embodiment, after obtaining a mixed point cloud of the target object and the target plane, a plane equation of the target plane can be generated according to the mixed point cloud.
[0052] Exemplarily, plane extraction can be performed on a mixed point cloud of a target object and a target plane; the target plane can be determined from the extracted planes, such as determining which plane among the extracted planes is the target plane based on the relative positional relationship between the target object and the target plane; and a plane equation of the target plane can be generated.
[0053] S104 : Reconstruct the surface of the target object based on the plane equation to obtain a three-dimensional model of the target object.
[0054] In this embodiment, after obtaining the plane equation of the target plane where the target object is placed, the surface of the target object can be reconstructed based on the plane equation to obtain a three-dimensional model of the target object.
[0055] There is no limit to the method of surface reconstruction of the target object. For example, the pose of the object image of the target object can be estimated, and the object image of the target object can be dedistorted and processed to obtain a processed image and the pose information of the processed image; based on the processed image and the pose information of the processed image, the surface of the target object is reconstructed, and during the surface reconstruction process, the bottom surface of the target object is plane-constrained based on the plane equation of the target plane to obtain a three-dimensional model of the target object.
[0056] The object modeling method provided in this embodiment obtains an object image of a target object at different perspectives. The object image is obtained by capturing an image of the target object placed on a target plane; the captured object image is processed to obtain a mixed point cloud of the target object and the target plane; a plane equation of the target plane is generated based on the mixed point cloud; and the surface of the target object is reconstructed based on the plane equation of the target plane to obtain a three-dimensional model of the target object. This embodiment utilizes the above-mentioned technical solution to obtain a plane equation of the target plane on which the object is placed based on the object image, and reconstructs the object based on the plane equation. This eliminates the need to collect bottom data of the object, and even in the absence of bottom data, a flat bottom surface can be reconstructed for the object, thereby reducing the difficulty of collecting object data and improving the quality of the constructed three-dimensional model.
[0057] FIG2 is a flow chart of another object modeling method provided by an embodiment of the present disclosure. The solution in this embodiment can be combined with one or more optional solutions in the above embodiments. Optionally, before generating the plane equation of the target plane based on the mixed point cloud, the method further includes: performing mask extraction on the target object to obtain the mask of the target object in the object image; generating the plane equation of the target plane based on the mixed point cloud includes: generating the plane equation of the target plane based on the mixed point cloud and the mask.
[0058] Optionally, reconstructing the surface of the target object based on the plane equation includes: generating a plane loss function of the target object based on the plane equation, the plane loss function being used to perform plane constraints during the process of reconstructing the surface of the target object; and reconstructing the surface of the target object based on the plane loss function.
[0059] Accordingly, as shown in FIG2 , the object modeling method provided in this embodiment may include:
[0060] S201 : Acquire object images of a target object at different viewing angles, wherein the object images are obtained by capturing images of the target object placed on a target plane.
[0061] S202: Process the object image to obtain a mixed point cloud of the target object and the target plane.
[0062] S203: Perform mask extraction on the target object to obtain a mask of the target object in the object image.
[0063] In this embodiment, after acquiring the target object's images at different viewing angles, mask extraction can be performed on the target object to obtain the target object's mask within the object image, thereby facilitating the subsequent generation of the plane equation of the target plane. The mask extraction method is not limited, and for example, existing mask extraction methods can be used to extract the target object's mask within the object image. This method will not be further described here.
[0064] It should be noted that the execution order of S202 and S203 is not limited. For example, after obtaining the object images of the target object at different perspectives, S202 and S203 can be executed simultaneously; S202 can also be executed first, and then S203 is executed after S202 is completed, or S203 can be executed first, and then S202 is executed after S203 is completed, and so on. The specific settings can be flexibly made according to needs.
[0065] S204: Generate a plane equation of the target plane according to the mixed point cloud and the mask.
[0066] In this embodiment, after processing the mixed point cloud of the target object and the target plane and extracting the mask of the target object in the object image, the plane equation of the target plane can be generated based on the processed mixed point cloud and the extracted mask. For example, the target plane on which the target object is placed can be determined based on the processed mixed point cloud and the extracted mask, thereby obtaining the plane equation of the target plane.
[0067] In some embodiments, generating the plane equation of the target plane based on the mixed point cloud and the mask includes: extracting the target point cloud in the mixed point cloud based on the mask and the point information of each point in the mixed point cloud, wherein the target point cloud is composed of points in the mixed point cloud that are not located on the target object; performing plane extraction on the target point cloud to obtain the plane equation of the target plane.
[0068] Among them, the point information of the points in the hybrid point cloud can be understood as information used to locate the points in the hybrid point cloud in the object image, such as the image identifier and pixel identifier corresponding to the points in the hybrid point cloud. The image identifier corresponding to a point in the hybrid point cloud can be the identifier of the object image corresponding to this point, which is used to indicate which object image this point corresponds to. The pixel identifier corresponding to a point in the hybrid point cloud can be the identifier of the pixel corresponding to this point. For example, the pixel identifier corresponding to a point in the hybrid point cloud can be the pixel coordinates of the pixel corresponding to this point, which is used to indicate which pixel in the object image to which this point corresponds. The target point cloud can be a point cloud composed of points in the hybrid point cloud that are not located on the target object, such as the non-main part in the hybrid point cloud. The main body can be understood as the target object.
[0069] In the above embodiment, the point cloud of the non-main part (i.e., not located on the target object) in the mixed point cloud can be extracted as the target point cloud, and the target point cloud can be plane extracted to obtain the plane equation of the target plane, so as to avoid the situation where the points located on the target object interfere with the determination of the target plane, thereby further improving the determination efficiency of the target plane and the accuracy of the determined target plane.
[0070] Specifically, the target point cloud that is not located on the target object in the mixed point cloud can be extracted based on the mask of the target object in the object image and the point information of the points in the mixed point cloud of the target object and the target plane; and plane extraction is performed on the extracted target point cloud to obtain the plane equation of the target plane.
[0071] Exemplarily, the target point cloud extraction process can be described as: for each point in the mixed point cloud, determine the corresponding object image and the corresponding pixel in the object image based on the point information of this point; based on the mask of the target object in the object image corresponding to this point, determine whether the pixel corresponding to this point in the object image is the pixel of the target object. If so, determine that this point is on the target object; if not, determine that this point is not on the target object.
[0072] After determining whether the points in the mixed point cloud are located on the target object, the target point cloud can be obtained based on this determination result. For example, the points in the mixed point cloud that are not located on the target object can be extracted, and the point cloud composed of the extracted points can be used as the target point cloud; the points in the mixed point cloud that are located on the target object can also be deleted, and the point cloud after the points located on the target object are deleted can be used as the target point cloud, and so on.
[0073] In the above embodiment, the method for extracting a plane from the target point cloud is not limited. For example, a random sample consensus (RANSAC) method can be used to extract a plane from the target point cloud, that is, RANSAC plane extraction is performed on the target point cloud to obtain a plane equation of the target plane nx+d=0. Where n is the normal vector of the target plane, that is, the first normal vector; d is the distance between the target plane and the origin, that is, the target distance.
[0074] S205 . Generate a plane loss function of the target object based on the plane equation, where the plane loss function is used to perform plane constraints during surface reconstruction of the target object.
[0075] Specifically, after obtaining the plane equation of the target plane, a plane loss function of the target object can be generated based on the plane equation. The plane loss function is used to perform plane constraints during the surface reconstruction process of the target object, such as constraining the reconstructed bottom surface of the target object during the surface reconstruction process of the target object. The plane loss function can be a loss function for constraining the reconstructed bottom surface of the target object. Exemplarily, the loss function can be used to constrain the angle between the bottom surface of the target object and the target plane and / or the smoothness of the bottom surface of the target object.
[0076] In this embodiment, exemplarily, a plane loss function can be constructed based on the normal vector of the target plane and / or the distance between the target plane and the origin. For example, the normal vector of the target plane and / or the distance between the target plane and the origin can be determined based on the plane equation of the target plane, and the plane loss function of the target object can be constructed based on this. At this time, optionally, the generating of the plane loss function of the target object based on the plane equation includes: determining the plane information of the target plane based on the plane equation, the plane information including the first normal vector of the target plane and / or the target distance between the target plane and the origin; and constructing the plane loss function of the target object based on the plane information. Among them, the plane information of the target plane can include a first normal vector and / or a target distance, the first normal vector can be considered as the normal vector of the target plane, and the target distance can be considered as the distance between the target plane and the origin.
[0077] In this embodiment, the method of constructing the plane loss function of the target object according to the plane information of the target plane can be flexibly set as needed, as long as the corresponding constraint effect can be achieved.
[0078] In some embodiments, the plane loss function of the target object may include a first plane loss function and / or a second plane loss function. In this case, optionally, constructing the plane loss function of the target object based on the plane information includes: constructing the first plane loss function of the target object based on the first normal vector and the target distance, the first plane loss function is used to constrain the angle between the second normal vector and the first normal vector, the second normal vector is the normal vector of the target space field, and the target space field includes a space field located in the bottom area of the target object; and / or constructing the second plane loss function of the target object based on the target distance, the second plane loss function is used to smoothly constrain the target space field.
[0079] The first plane loss function can be constructed based on the first normal vector of the target plane and the target distance between the target plane and the origin. The first plane loss function can be used to constrain the angle between the second normal vector and the first normal vector, for example, to constrain the second normal vector to be as close as possible to the first normal vector of the target plane.
[0080] The second plane loss function can be constructed based on the target distance between the target plane and the origin. The second loss function can be used to constrain the target space field to be smooth, such as smoothing the target space field.
[0081] The second normal vector may be a normal vector of a target space field generated during surface reconstruction of the target object. The target space field may include a space field located in a bottom region of the target object. The bottom region may be an area near the bottom of the target object, such as an area within a set distance range from the target plane. The type of the space field is not limited. Exemplarily, the space field may be a signed distance field (SDF).
[0082] Specifically, a first plane loss function of the target object can be constructed based on the first normal vector of the target plane and the target distance, so that in the subsequent process of surface reconstruction of the target object, the angle between the second normal vector and the first normal vector of the target space field near the bottom can be constrained by the first plane loss function; and / or, a second plane loss function of the target object can be constructed based on the target distance, so that in the subsequent process of surface reconstruction of the target object, the target space field near the bottom can be smoothly constrained by the second plane loss function.
[0083] In the above embodiment, the expressions of the first plane loss function and the second plane loss function are not limited, as long as they can achieve the corresponding effects.
[0084] For example, the first plane loss function L1 may be:
[0085] The second plane loss function L2 can be:
[0086] Where N is the number of sampling points on each ray during surface reconstruction, which can be set in advance. i is the i-th sampling point x on each ray during surface reconstruction i The first weight coefficient of ω″ i is the i-th sampling point x on each ray during surface reconstruction i The second weight coefficient of n is the first normal vector of the target plane. is the i-th sampling point x on each ray during surface reconstruction i The normal vector (i.e., the second normal vector) of the spatial field (such as SDF) at . is the i-th sampling point x on each ray during surface reconstruction i The second derivative of the spatial field (such as SDF) at .
[0087] In this embodiment, considering the above point x i is the sampling point obtained on the light, which is distributed everywhere in the space. Therefore, the plane loss function L1 and L2 may act on the entire object, affecting the object reconstruction result. Therefore, two weights can be added to ω′ i and ω″ i , through ω′ i and ω″ i The effects of the plane loss functions L1 and L2 are mainly limited to the area near the bottom of the target object.
[0088] The first weight coefficient ω′ i and the second weight coefficient ω″ i It can be set as needed, as long as it can achieve the above effect. For example, the first weight coefficient ω′ i and the second weight coefficient ω″ i It can be set as: i =[|d i | <T i ] T i =ae -c*k
[0089] During the surface reconstruction process, ω′ can be introduced i The spatial field (such as SDF) near the bottom is constrained, and its influence range decreases with the increase of the number of iterations. i is point x i Distance to target plane, T i is the threshold that decreases with the increase of the number of iterations, a is the preset initial value, c is the preset coefficient, and k is the number of iterations.
[0090] In the surface reconstruction process, such as in the volume rendering stage of the surface reconstruction process, the light first reaches the upper surface and side surface of the target object, and then reaches the ground of the target object. Therefore, by introducing ω″ i To reduce the impact on the light's first pass through the surface. i Point x is obtained by volume rendering i The weight of max(ω i ) is at point x i The weight of the sampling point with the largest weight obtained by volume rendering on the light line is the maximum weight obtained by volume rendering on this light line; min(ω i ) is at point x i The weight of the sampling point with the smallest weight obtained by volume rendering on the ray, that is, the minimum weight obtained by volume rendering on this ray.
[0091] S206 : Reconstruct the surface of the target object based on the plane loss function to obtain a three-dimensional model of the target object.
[0092] Specifically, after generating a plane loss function for the target object, the target object can be reconstructed. During the surface reconstruction process, the generated plane loss function is used to perform plane constraints. Thus, after the surface reconstruction of the target object is completed, a three-dimensional model of the target object can be obtained.
[0093] The object modeling method provided in this embodiment generates a plane equation of the target plane based on a mixed point cloud of the target object and the target plane and a mask of the target object in the object image, and generates a plane loss function of the target object based on the plane equation. By using the plane loss function to perform plane constraints during the surface reconstruction process of the target object, the generation efficiency of the plane equation and the accuracy of the generated plane equation can be improved, thereby further improving the modeling quality of the object modeling.
[0094] FIG3 is a flow chart of an optional object modeling method provided by an embodiment of the present disclosure. As shown in FIG3 , the process of modeling a target object can be described as follows:
[0095] A1. Obtain multi-view images of the target object whose bottom surface data is missing (ie, obtain the object images of the target object at different viewing angles).
[0096] A2. Estimate the pose of the object based on the SFM method to obtain an SFM sparse point cloud (i.e., a mixed point cloud of the target object and the target plane), as well as an object image and pose information of the object image obtained after dedistortion of the target object.
[0097] A3. Extract the sparse point cloud of the non-main part (i.e., target point cloud) using the object mask and the image identification and pixel position of each point recorded by the SFM sparse point cloud. Perform RANSAC plane extraction on the sparse point cloud of the non-main part to obtain the plane equation of the plane where the object is located (i.e., target plane).
[0098] A4. Construct a plane-constrained loss function (i.e., a plane loss function) according to the plane equation, such as constructing a first plane loss function L1 to constrain the SDF field normal vector (i.e., the second normal vector) near the bottom to be consistent with the normal vector of the target plane (i.e., the first normal vector) according to the plane equation; and / or constructing a second plane loss function L2 to smooth the SDF field near the bottom according to the plane equation.
[0099] A5. Reconstruct the surface of the target object based on the processed object image, the pose information of the object image, and the constructed plane loss function to obtain an SDF field describing the target object.
[0100] When reconstructing the surface of a target object, for example, the initial SDF field can be trained. For example, the processed object image and the pose information of the object image are input into the SDF field, volume rendering is performed, and the plane loss (Plane Loss), RGB loss (RGB Loss), Eikonal loss (Eikonal Loss) and mask loss (Mark Loss) of the volume rendering are calculated. The above losses are constrained by the plane loss function, RGB loss function, Eikonal loss function and mask loss function respectively. Thus, after multiple iterative training, the SDF field describing the target object can be obtained.
[0101] A6. Perform isosurface extraction (Marching Cubes, MC) on the obtained SDF field to obtain a three-dimensional model of the target object.
[0102] It can be seen that the object modeling method provided by this embodiment first detects the plane equation of the plane where the object is located through the SFM sparse point cloud; then, the normal of the SDF field near the bottom is constrained to be consistent with the plane normal according to the plane equation, and the SDF field near the bottom is smoothed according to the plane equation. At the same time, in order to reduce the impact on non-bottom surfaces in the above process, the two weight coefficients used can be set according to the distance from the sampling point to the bottom and the number of surfaces passed by the light. By making improvements during the model training process, this embodiment can solve the problem of missing bottom surface data affecting the overall reconstruction quality in surface reconstruction, and fundamentally solve the negative impact of missing data from end to end.
[0103] FIG4 is a block diagram of the structure of an object modeling device provided by an embodiment of the present disclosure. The device can be implemented by software and / or hardware and can be configured in an electronic device, for example, a mobile phone or a tablet computer, and can model an object by executing an object modeling method, such as modeling an object when the bottom data of the object is missing. As shown in FIG4 , the object modeling device provided by this embodiment may include: an image acquisition module 401, an image processing module 402, an equation generation module 403 and a surface reconstruction module 404, wherein,
[0104] An image acquisition module 401 is configured to acquire an image of a target object at different viewing angles, wherein the image of the target object is acquired by capturing an image of the target object placed on a target plane.
[0105] An image processing module 402 is configured to process the object image to obtain a mixed point cloud of the target object and the target plane;
[0106] An equation generating module 403 is configured to generate a plane equation of the target plane according to the mixed point cloud;
[0107] The surface reconstruction module 404 is configured to perform surface reconstruction on the target object based on the plane equation to obtain a three-dimensional model of the target object.
[0108] The object modeling device provided in this embodiment acquires the object image of the target object at different perspectives through an image acquisition module, and the object image is obtained by performing image acquisition on the target object placed on the target plane; the acquired object image is processed by an image processing module to obtain a mixed point cloud of the target object and the target plane; the plane equation of the target plane is generated according to the mixed point cloud by an equation generation module; and the surface of the target object is reconstructed based on the plane equation of the target plane by a surface reconstruction module to obtain a three-dimensional model of the target object. This embodiment utilizes the above-mentioned technical solution to obtain the plane equation of the target plane on which the object is placed based on the object image, and reconstructs the object based on the plane equation, without the need to collect the bottom data of the object. Even in the absence of bottom data, a flat bottom surface can be reconstructed for the object, which can reduce the difficulty of collecting object data and improve the quality of the constructed three-dimensional model.
[0109] Furthermore, the object reconstruction device provided in this embodiment may also include: a mask extraction module, which is used to perform mask extraction on the target object before generating the plane equation of the target plane based on the mixed point cloud to obtain the mask of the target object in the object image; the equation generation module 403 can be used to: generate the plane equation of the target plane based on the mixed point cloud and the mask.
[0110] Optionally, the equation generation module 403 includes: a point cloud extraction unit, used to extract the target point cloud in the mixed point cloud based on the mask and the point information of each point in the mixed point cloud, wherein the target point cloud is composed of points in the mixed point cloud that are not located on the target object; a plane extraction unit, used to perform plane extraction on the target point cloud to obtain the plane equation of the target plane.
[0111] Optionally, the surface reconstruction module 404 includes: a function generation unit, used to generate a plane loss function of the target object based on the plane equation, and the plane loss function is used to perform plane constraints during the process of surface reconstruction of the target object; a surface reconstruction unit, used to reconstruct the surface of the target object based on the plane loss function.
[0112] Optionally, the function generation unit includes: an information determination subunit, used to determine the plane information of the target plane based on the plane equation, the plane information including the first normal vector of the target plane and / or the target distance between the target plane and the origin; a function construction subunit, used to construct the plane loss function of the target object based on the plane information.
[0113] Optionally, the function construction subunit is specifically used to: construct a first plane loss function of the target object based on the first normal vector and the target distance, the first plane loss function is used to constrain the angle between the second normal vector and the first normal vector, the second normal vector is the normal vector of the target space field, and the target space field includes the space field located in the bottom area of the target object; and / or, construct a second plane loss function of the target object based on the target distance, the second plane loss function is used to smoothly constrain the target space field.
[0114] Optionally, the image processing module 402 is configured to process the object image using a structure from motion (SFM) method to obtain an SFM sparse point cloud of the target object and the target plane as a mixed point cloud of the target object and the target plane.
[0115] The object modeling device provided in the embodiments of the present disclosure can execute the object modeling method provided in any embodiment of the present disclosure, and has the corresponding functional modules and beneficial effects of executing the object modeling method. For technical details not fully described in this embodiment, please refer to the object modeling method provided in any embodiment of the present disclosure.
[0116] Reference is now made to FIG5 , which illustrates a schematic diagram of the structure of an electronic device (e.g., a terminal device) 500 suitable for implementing embodiments of the present disclosure. The terminal device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. The electronic device illustrated in FIG5 is merely an example and should not limit the functionality or scope of use of the embodiments of the present disclosure.
[0117] As shown in Figure 5, the electronic device 500 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage device 508 into a random access memory (RAM) 503. Various programs and data required for the operation of the electronic device 500 are also stored in the RAM 503. The processing device 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0118] Typically, the following devices may be connected to the I / O interface 505: an input device 506 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 507 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 508 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 509. The communication device 509 may allow the electronic device 500 to communicate with other devices wirelessly or by wire to exchange data. Although FIG5 shows the electronic device 500 with various devices, it should be understood that not all of the devices shown are required to be implemented or present. More or fewer devices may alternatively be implemented or present.
[0119] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 509, or installed from the storage device 508, or installed from the ROM 502. When the computer program is executed by the processing device 501, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.
[0120] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.
[0121] In some embodiments, the client and server can communicate using any currently known or later developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or later developed network.
[0122] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0123] The above-mentioned computer-readable medium carries one or more programs. When the above-mentioned one or more programs are executed by the electronic device, the electronic device is enabled to: obtain object images of the target object at different perspectives, where the object images are obtained by capturing images of the target object placed on a target plane; process the object images to obtain a mixed point cloud of the target object and the target plane; generate a plane equation of the target plane based on the mixed point cloud; and reconstruct the surface of the target object based on the plane equation to obtain a three-dimensional model of the target object.
[0124] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including, but not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0125] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0126] The units involved in the embodiments described in this disclosure may be implemented in software or hardware, wherein the name of a module does not, in some cases, limit the unit itself.
[0127] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.
[0128] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0129] According to one or more embodiments of the present disclosure, Example 1 provides an object modeling method, including:
[0130] Acquire an object image of a target object at different viewing angles, wherein the object image is obtained by capturing an image of the target object placed on a target plane;
[0131] Processing the object image to obtain a mixed point cloud of the target object and the target plane;
[0132] generating a plane equation of the target plane according to the mixed point cloud;
[0133] The surface of the target object is reconstructed based on the plane equation to obtain a three-dimensional model of the target object.
[0134] According to one or more embodiments of the present disclosure, Example 2 is the method according to Example 1, and before generating the plane equation of the target plane according to the mixed point cloud, further comprising:
[0135] Performing mask extraction on the target object to obtain a mask of the target object in the object image;
[0136] Generating the plane equation of the target plane according to the mixed point cloud includes:
[0137] A plane equation of the target plane is generated according to the mixed point cloud and the mask.
[0138] According to one or more embodiments of the present disclosure, Example 3 is the method according to Example 2, wherein generating the plane equation of the target plane according to the mixed point cloud and the mask includes:
[0139] Extracting a target point cloud from the mixed point cloud based on the mask and point information of each point in the mixed point cloud, where the target point cloud is composed of points in the mixed point cloud that are not located on the target object;
[0140] Perform plane extraction on the target point cloud to obtain a plane equation of the target plane.
[0141] According to one or more embodiments of the present disclosure, Example 4 is the method according to Example 1, wherein reconstructing the surface of the target object based on the plane equation includes:
[0142] generating a plane loss function of the target object based on the plane equation, wherein the plane loss function is used to perform plane constraints during surface reconstruction of the target object;
[0143] The surface of the target object is reconstructed based on the plane loss function.
[0144] According to one or more embodiments of the present disclosure, Example 5, according to the method of Example 4, generating the plane loss function of the target object based on the plane equation includes:
[0145] determining plane information of the target plane based on the plane equation, the plane information including a first normal vector of the target plane and / or a target distance between the target plane and an origin;
[0146] A plane loss function of the target object is constructed according to the plane information.
[0147] According to one or more embodiments of the present disclosure, Example 6 is the method according to Example 5, wherein constructing the plane loss function of the target object based on the plane information includes:
[0148] Constructing a first plane loss function of the target object according to the first normal vector and the target distance, wherein the first plane loss function is used to constrain an angle between a second normal vector and the first normal vector, where the second normal vector is a normal vector of a target space field, and the target space field includes a space field located in a bottom area of the target object; and / or
[0149] A second plane loss function of the target object is constructed according to the target distance, and the second plane loss function is used to perform smoothness constraints on the target space field.
[0150] According to one or more embodiments of the present disclosure, Example 7, according to the method of any one of Examples 1-6, processing the object image to obtain a mixed point cloud of the target object and the target plane includes:
[0151] The object image is processed using a structure from motion (SFM) method to obtain SFM sparse point clouds of the target object and the target plane as mixed point clouds of the target object and the target plane.
[0152] According to one or more embodiments of the present disclosure, Example 8 provides an object modeling device, including:
[0153] An image acquisition module is used to acquire an object image of a target object at different viewing angles, wherein the object image is obtained by acquiring an image of the target object placed on a target plane;
[0154] An image processing module, configured to process the object image to obtain a mixed point cloud of the target object and the target plane;
[0155] an equation generating module, configured to generate a plane equation of the target plane according to the mixed point cloud;
[0156] A surface reconstruction module is used to reconstruct the surface of the target object based on the plane equation to obtain a three-dimensional model of the target object.
[0157] According to one or more embodiments of the present disclosure, Example 9 provides an electronic device, including:
[0158] one or more processors;
[0159] a memory for storing one or more programs,
[0160] When the one or more programs are executed by the one or more processors, the one or more processors implement the object modeling method as described in any one of Examples 1-7.
[0161] According to one or more embodiments of the present disclosure, Example 10 provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the object modeling method as described in any one of Examples 1-7.
[0162] According to one or more embodiments of the present disclosure, Example 11 provides a computer program product. When the computer program product is executed by a computer, the computer implements the object modeling method described in any one of Examples 1-7.
[0163] The above description is merely a preferred embodiment of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also includes other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this disclosure.
[0164] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details have been included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.
[0165] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.
Claims
1. A method for object modeling, comprising: Acquire an object image of a target object at different viewing angles, wherein the object image is obtained by capturing an image of the target object placed on a target plane; Processing the object image to obtain a mixed point cloud of the target object and the target plane; generating a plane equation of the target plane according to the mixed point cloud; The surface of the target object is reconstructed based on the plane equation to obtain a three-dimensional model of the target object.
2. The method according to claim 1, wherein Before generating the plane equation of the target plane according to the mixed point cloud, the method further includes: Performing mask extraction on the target object to obtain a mask of the target object in the object image; Generating the plane equation of the target plane according to the mixed point cloud includes: A plane equation of the target plane is generated according to the mixed point cloud and the mask.
3. The method according to claim 2, wherein: Generating the plane equation of the target plane according to the mixed point cloud and the mask includes: Extracting a target point cloud from the mixed point cloud based on the mask and point information of each point in the mixed point cloud, where the target point cloud is composed of points in the mixed point cloud that are not located on the target object; Perform plane extraction on the target point cloud to obtain a plane equation of the target plane.
4. The method according to any one of claims 1 to 3, wherein The performing surface reconstruction on the target object based on the plane equation includes: generating a plane loss function of the target object based on the plane equation, wherein the plane loss function is used to perform plane constraints during surface reconstruction of the target object; The surface of the target object is reconstructed based on the plane loss function.
5. The method according to claim 4, wherein Generating a plane loss function of the target object based on the plane equation includes: determining plane information of the target plane based on the plane equation, the plane information including a first normal vector of the target plane and / or a target distance between the target plane and an origin; A plane loss function of the target object is constructed according to the plane information.
6. The method according to claim 5, wherein: The constructing a plane loss function of the target object according to the plane information includes: Constructing a first plane loss function of the target object according to the first normal vector and the target distance, wherein the first plane loss function is used to constrain an angle between a second normal vector and the first normal vector, where the second normal vector is a normal vector of a target space field, and the target space field includes a space field located in a bottom area of the target object; and / or A second plane loss function of the target object is constructed according to the target distance, and the second plane loss function is used to perform smoothness constraints on the target space field.
7. The method according to any one of claims 1 to 6, wherein The processing of the object image to obtain a mixed point cloud of the target object and the target plane includes: The object image is processed using a structure-from-motion method to obtain a sparse point cloud of the target object and the target plane as a mixed point cloud of the target object and the target plane.
8. An object modeling device comprising: an image acquisition module configured to acquire an object image of a target object at different viewing angles, wherein the object image is obtained by acquiring an image of the target object placed on a target plane; an image processing module configured to process the object image to obtain a mixed point cloud of the target object and the target plane; an equation generating module, configured to generate a plane equation of the target plane according to the mixed point cloud; The surface reconstruction module is configured to perform surface reconstruction on the target object based on the plane equation to obtain a three-dimensional model of the target object.
9. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor. The computer program is executed by the at least one processor to enable the at least one processor to perform the object modeling method according to any one of claims 1 to 7.
10. A computer-readable storage medium storing computer instructions, wherein the computer instructions are configured to enable a processor to implement the object modeling method according to any one of claims 1 to 7 when executed.
11. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements the object modeling method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Geometric modeling method based on pipe factory point cloud
CN102722907A
Laser odometer error correction method and system, storage medium and computing equipment
CN113935904A
Image-based three-dimensional modeling method and device, electronic equipment and storage medium
CN115330927A
Object modeling method and device, electronic equipment, storage medium and program product
CN117994471A
Three-dimensional object detection method and apparatus, and computer-readable storage medium
WO2023078052A1