Article three-dimensional construction method and device, electronic equipment, medium and program product

By placing a target marker between the target object and the placement surface and using the marker segmentation model to segment the image information, the problem of low three-dimensional construction efficiency caused by the integration of the target object and the placement surface is solved, and accurate and efficient three-dimensional structure construction is achieved.

CN120635318APending Publication Date: 2025-09-12BEIJING WODONG TIANJUN INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510765839.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

In the prior art, the target object and the placement surface are integrated during the three-dimensional construction process, which requires manual removal of the placement surface structure, which is time-consuming, labor-intensive, and inefficient.

Method used

By placing a target marker between the target object and the placement surface, and using a pre-trained marker segmentation model, the marker and object information in the captured image is acquired and segmented, object segmentation information is generated, and the three-dimensional structure of the target object is constructed.

Benefits of technology

It achieves spatial isolation between the target object and the placement surface, accurately and efficiently generates the three-dimensional structure of the target object, avoids manual intervention, and improves the efficiency of three-dimensional construction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120635318A_ABST
    Figure CN120635318A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an article three-dimensional construction method and device, electronic equipment, a medium and a program product. According to one specific embodiment, the method comprises the steps that shot image sets, corresponding to different shooting orientations, of a target object in a target scene are obtained, and the target scene is a scene where the target object is placed on a target marker and the target marker is placed on an object placement surface; the target marker is used for assisting in segmenting target article information in the shot image; for each shot image, executing a generation step: determining marker segmentation information in the shot image by using a marker segmentation model; generating article segmentation information corresponding to each article in the shot image according to the marker segmentation information; and executing three-dimensional structure construction for the target article according to the article segmentation information set. The implementation mode is related to three-dimensional modeling of the object, and the object placing surface and the target object can be accurately and efficiently segmented, so that three-dimensional structure construction of the target object can be conveniently realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present disclosure relate to the field of three-dimensional modeling of objects, and in particular to a method, device, electronic device, medium, and program product for constructing three-dimensional objects. Background Art

[0002] Currently, 3D construction utilizes the principles of multi-view geometry (i.e., observing an object from multiple perspectives). Images or scanned data of the object under construction are captured from these perspectives, and then combined with computer graphics and image processing to infer the corresponding 3D structure of the object. However, during the 3D construction process, the object and the surface on which it is placed are often integrated. This necessitates manual removal of the corresponding structure using model editing tools, which is time-consuming and inefficient, leading to low 3D construction efficiency.

[0003] The above information disclosed in this Background section is only for enhancement of understanding of the background of the inventive concept and therefore it may contain information that does not form the prior art that is already known in this country to a person of ordinary skill in the art. Summary of the Invention

[0004] The content of this disclosure is used to briefly introduce concepts that will be described in detail in the detailed description section below. The content of this disclosure is not intended to identify key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.

[0005] Some embodiments of the present disclosure propose methods, devices, electronic devices, media, and program products for constructing three-dimensional objects to solve the technical problems mentioned in the above background technology section.

[0006] In a first aspect, some embodiments of the present disclosure provide a method for constructing a three-dimensional object, comprising: obtaining a set of captured images of a target object corresponding to different shooting orientations in a target scene, wherein the target scene is a scene in which the target object is placed on a target marker and the target marker is placed on a placement surface, and the target marker is used to assist in segmenting the target object information in the captured image; for each captured image in the captured image set, performing the following generation steps: determining the marker segmentation information in the captured image using a pre-trained marker segmentation model; generating object segmentation information corresponding to each object in the captured image based on the marker segmentation information; and performing three-dimensional structure construction for the target object based on the obtained object segmentation information set.

[0007] Optionally, the above-mentioned marker segmentation model is a first interactive segmentation model for segmenting markers; and the above-mentioned marker segmentation model is used to determine the marker segmentation information in the above-mentioned captured image, including: generating a first prompt image corresponding to the above-mentioned captured image, wherein the above-mentioned first prompt image has a first positive point set and a first negative point set for the marker information, wherein the surface corresponding to the above-mentioned target marker has multiple marker information with a regular distribution pattern, the first positive point is a marker point in the marker information, and the first negative point is a marker point not in the marker information; the above-mentioned first prompt image and the above-mentioned captured image are input into the above-mentioned first interactive segmentation model to generate the above-mentioned marker segmentation information.

[0008] Optionally, the above-mentioned marker segmentation model is a non-interactive segmentation model for segmenting markers; and the above-mentioned use of the pre-trained marker segmentation model to determine the marker segmentation information in the above-mentioned captured image includes: inputting the above-mentioned captured image into the above-mentioned non-interactive segmentation model to generate marker segmentation information.

[0009] Optionally, the above-mentioned generating of item segmentation information corresponding to each item in the captured image based on the above-mentioned marker segmentation information includes: setting the image area corresponding to the above-mentioned marker segmentation information in the above-mentioned captured image as the background area, setting the remaining image areas as the foreground area, and obtaining foreground image information as the item segmentation information corresponding to each item.

[0010] Optionally, the above-mentioned three-dimensional structure construction for the above-mentioned target object is performed based on the obtained object segmentation information set, including: using the three-dimensional structure construction model to generate three-dimensional structure construction information based on the above-mentioned object segmentation information set; removing the plane information corresponding to at least one isolated plane from the above-mentioned three-dimensional structure construction information to obtain the three-dimensional structure information corresponding to the above-mentioned target object.

[0011] Optionally, after determining the marker segmentation information in the captured image using the pre-trained marker segmentation model, the method further includes: in response to determining that there is a pre-trained second interactive segmentation model for segmenting objects, generating a second prompt image corresponding to the captured image, wherein the second prompt image has a second positive point set and a second negative point set for the target object, and the target marker corresponds to a surface with a plurality of regularly distributed marking information, the second positive point is a marking point on the surface of the target object, and the second negative point is a marking point in the marking information; inputting the second prompt image and the captured image into the second interactive segmentation model to generate the object segmentation information.

[0012] Optionally, the above-mentioned three-dimensional structure construction for the above-mentioned target object is performed based on the obtained object segmentation information set, including: for each object segmentation information in the above-mentioned object segmentation information set, filtering out the segmentation information corresponding to the above-mentioned target object from the above-mentioned object segmentation information as the target segmentation information; inputting the shooting orientation set corresponding to the above-mentioned shooting image set, the obtained target segmentation information set and the above-mentioned shooting image set into the three-dimensional structure construction model to generate the three-dimensional structure information corresponding to the above-mentioned target object.

[0013] In a second aspect, some embodiments of the present disclosure provide a three-dimensional object construction device, comprising: an acquisition unit, configured to acquire a set of captured images of a target object corresponding to different shooting orientations in a target scene, wherein the target scene is a scene where the target object is placed on a target marker and the target marker is placed on a placement surface, and the target marker is used to assist in segmenting the target object information in the captured image; a first execution unit, configured to perform the following generation steps for each captured image in the captured image set: determine the marker segmentation information in the captured image using a pre-trained marker segmentation model; generate object segmentation information corresponding to each object in the captured image based on the marker segmentation information; and a second execution unit, configured to execute three-dimensional structure construction for the target object based on the obtained object segmentation information set.

[0014] Optionally, the above-mentioned marker segmentation model is a first interactive segmentation model for segmenting markers; and the first execution unit can be configured to: generate a first prompt image corresponding to the above-mentioned captured image, wherein the above-mentioned first prompt image has a first positive point set and a first negative point set for the marking information, wherein the above-mentioned target marker corresponds to a surface with a plurality of marking information with a regular distribution, the first positive point is a marking point in the marking information, and the first negative point is a marking point not in the marking information; input the above-mentioned first prompt image and the above-mentioned captured image into the above-mentioned first interactive segmentation model to generate the above-mentioned marker segmentation information.

[0015] Optionally, the marker segmentation model is a non-interactive segmentation model for segmenting markers; and the first execution unit may be configured to: input the captured image into the non-interactive segmentation model to generate marker segmentation information.

[0016] Optionally, the first execution unit can be configured to: set the image area corresponding to the above-mentioned marker segmentation information in the above-mentioned captured image as the background area, set the remaining image areas as the foreground area, and obtain foreground image information as the item segmentation information corresponding to each item.

[0017] Optionally, the second execution unit can be configured to: generate three-dimensional structure construction information by using a three-dimensional structure construction model based on the above-mentioned object segmentation information set; remove the plane information corresponding to at least one isolated plane from the above-mentioned three-dimensional structure construction information to obtain the three-dimensional structure information corresponding to the above-mentioned target object.

[0018] Optionally, the first execution unit can be configured to: in response to determining that there is a pre-trained second interactive segmentation model for segmenting objects, generate a second prompt image corresponding to the above-mentioned captured image, wherein the above-mentioned second prompt image has a second positive point set and a second negative point set for the above-mentioned target object, and the above-mentioned target marker corresponds to a plurality of regularly distributed marking information on the surface, the second positive point is a marking point on the surface of the target object, and the second negative point is a marking point in the marking information; input the above-mentioned second prompt image and the above-mentioned captured image into the above-mentioned second interactive segmentation model to generate the above-mentioned object segmentation information.

[0019] Optionally, the second execution unit can be configured to: for each item segmentation information in the above-mentioned item segmentation information set, filter out the segmentation information corresponding to the above-mentioned target item from the above-mentioned item segmentation information as the target segmentation information; input the shooting orientation set corresponding to the above-mentioned shooting image set, the obtained target segmentation information set and the above-mentioned shooting image set into the three-dimensional structure construction model to generate the three-dimensional structure information corresponding to the above-mentioned target item.

[0020] In a third aspect, some embodiments of the present disclosure provide an electronic device comprising: one or more processors; a storage device on which one or more programs are stored, and when the one or more programs are executed by one or more processors, the one or more processors implement the method described in any implementation manner in the first aspect.

[0021] In a fourth aspect, some embodiments of the present disclosure provide a computer-readable medium having a computer program stored thereon, wherein when the program is executed by a processor, the method described in any implementation manner in the first aspect is implemented.

[0022] In a fifth aspect, some embodiments of the present disclosure provide a computer program product, including a computer program, which implements the method described in any implementation manner in the first aspect when executed by a processor.

[0023] The above-mentioned various embodiments of the present disclosure have the following beneficial effects: through the object three-dimensional construction method of some embodiments of the present disclosure, the placement surface and the target object can be accurately and efficiently segmented to facilitate the three-dimensional structure construction of the target object. Specifically, in the process of constructing the three-dimensional structure of the target object, the construction result is often that the target object and the placement surface are integrated, resulting in the need to manually use model editing tools to cut out the structure corresponding to the placement surface, which is time-consuming and labor-intensive, resulting in low efficiency of three-dimensional construction. Based on this, the object three-dimensional construction method of some embodiments of the present disclosure, first, obtains a set of photographed images corresponding to different shooting orientations of the target object in the target scene, wherein the above-mentioned target scene is a scene in which the above-mentioned target object is placed on the target marker and the above-mentioned target marker is placed on the placement surface, and the above-mentioned target marker is used to assist in segmenting the target object information in the photographed image. Here, by obtaining a set of photographed images, the structural information corresponding to the target object in different shooting orientations is obtained. By placing a target marker between the target item and the storage surface, the target marker is used to achieve spatial isolation between the target item and the storage surface and segment the target marker, thereby avoiding the problem of contact adhesion between the target item and the storage surface during the three-dimensional construction process. Then, for each captured image in the above-mentioned captured image set, the following generation steps are performed: First, using a pre-trained marker segmentation model, the marker segmentation information in the above-mentioned captured image can be accurately determined to achieve object segmentation for the target marker in the captured image. Second, based on the above-mentioned marker segmentation information, item segmentation information corresponding to each item in the above-mentioned captured image can be accurately generated. Here, the marker segmentation information can be used to segment the target marked items to achieve spatial isolation between the storage surface and the target item. Finally, based on the obtained item segmentation information set, a three-dimensional structure construction is performed for the above-mentioned target item to accurately and efficiently generate the corresponding three-dimensional structure information of the target item. In summary, spatial isolation is achieved by adding a target marker between the target item and the storage surface. Based on this, the segmentation and recognition of target markers can be used to accurately determine the object information in the corresponding image of the target object without being affected by the placement surface, so as to facilitate the construction of three-dimensional structures. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] The above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that components and elements are not necessarily drawn to scale.

[0025] Figure 1 is a schematic diagram of an application scenario of a method for constructing a three-dimensional object according to some embodiments of the present disclosure;

[0026] Figure 2 is a flow chart of some embodiments of the method for three-dimensionally constructing an object according to the present disclosure;

[0027] Figure 3 is a schematic diagram of a target marker in some embodiments of the method for constructing a three-dimensional object according to the present disclosure;

[0028] Figure 4-Figure 5 is a schematic diagram of a target marker in a target scene according to some embodiments of the method for constructing a three-dimensional object of the present disclosure;

[0029] Figure 6 is a general schematic diagram of a method for constructing a three-dimensional object in some embodiments of the method for constructing a three-dimensional object according to the present disclosure;

[0030] Figure 7 is a flow chart of other embodiments of the method for constructing a three-dimensional object according to the present disclosure;

[0031] Figure 8 is a schematic diagram of an image corresponding to a first prompt image in some embodiments of the method for constructing a three-dimensional object according to the present disclosure;

[0032] Figure 9 Schematic diagrams of the structures of some embodiments of the device for constructing three-dimensional objects according to the present disclosure;

[0033] Figure 10 It is a structural diagram of an electronic device suitable for implementing some embodiments of the present disclosure. DETAILED DESCRIPTION

[0034] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments described herein. On the contrary, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.

[0035] It should also be noted that, for ease of description, only the parts related to the invention are shown in the drawings. In the absence of conflict, the embodiments and features in the embodiments of the present disclosure may be combined with each other.

[0036] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0037] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".

[0038] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0039] Before performing any of the collection, storage, and use of user personal information (such as captured images) involved in this disclosure, the relevant organizations or individuals must fulfill their obligations, including conducting personal information security impact assessments, fulfilling their obligations to inform the personal information subjects, and obtaining the prior authorization and consent of the personal information subjects.

[0040] The present disclosure will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.

[0041] Figure 1 It is a schematic diagram of an application scenario of the method for constructing three-dimensional objects according to some embodiments of the present disclosure.

[0042] exist Figure 1 In the application scenario, first, the electronic device 101 can obtain a set of captured images 103 of the target object 102 corresponding to different shooting orientations in the target scene. The target scene is a scene in which the target object 102 is placed on a target marker and the target marker is placed on a placement surface. The target marker is used to assist in segmenting the target object information in the captured image. Then, for each captured image in the captured image set 103, the electronic device 101 can perform the following generation steps: First, using the pre-trained marker segmentation model 104, determine the marker segmentation information in the captured image. In this application scenario, the captured image 1031 is captured image A, and the corresponding marker segmentation information 105 is marker segmentation information A. Second, based on the marker segmentation information, generate the item segmentation information corresponding to each item in the captured image. In this application scenario, the marker segmentation information 105 is marker segmentation information A, and the corresponding item segmentation information 1061 is item segmentation information A. Finally, the electronic device 101 may construct a three-dimensional structure for the target object 102 according to the obtained object segmentation information set 106 .

[0043] It should be noted that the electronic device 101 can be hardware or software. When the electronic device is hardware, it can be implemented as a distributed cluster consisting of multiple servers or terminal devices, or it can be implemented as a single server or a single terminal device. When the electronic device is embodied as software, it can be installed in the hardware devices listed above. It can be implemented as multiple software or software modules for providing distributed services, for example, or it can be implemented as a single software or software module. No specific limitation is made here.

[0044] It should be understood that Figure 1 The number of electronic devices in the embodiment is merely illustrative. Any number of electronic devices may be provided according to implementation requirements.

[0045] Continue to refer Figure 2 , shows a process 200 of some embodiments of the method for constructing a three-dimensional object according to the present disclosure. The method for constructing a three-dimensional object includes the following steps:

[0046] Step 201: Acquire a set of images of a target object in a target scene corresponding to different shooting orientations.

[0047] In some embodiments, the execution subject of the above-mentioned three-dimensional construction method of the object (for example Figure 1 The electronic device 101 shown) can obtain a set of captured images of a target object in a target scene corresponding to different shooting directions through a wired connection or a wireless connection. The target object may be an object to be constructed into a three-dimensional structure. For example, the target object may be a toy car. The target scene is a scene in which the target object is placed on a target marker and the target marker is placed on a storage surface. The target marker may be a marker with predetermined marking information on its surface. Since the target marker has predetermined marking information, it is convenient for the image segmentation model to perform semantic segmentation and marker recognition. That is, the image segmentation model can more easily learn the marking features of the predetermined marking information to facilitate segmentation and recognition tasks. For example, the target marker may be marker paper. The predetermined marking information on the surface of the target marker may be rectangles with regularly distributed colors. The target marker is used to assist in segmenting the target object information in the captured image. The target object information may be the object information of the target object in the image.

[0048] See also Figure 3 , a schematic diagram of target markers is shown.

[0049] Here, the target marker can be a marker of various colors, and the corresponding color can be set for the three-dimensional construction scene of the corresponding target object. Figure 3The target markers in the image are represented by individual square boxes. Each square box contains content that is easy to segment and identify. Furthermore, the textures corresponding to the square boxes are predetermined. Using an image segmentation model to segment target markers offers two advantages: First, target markers have regular features, making segmentation much easier than segmenting objects with uncertain textures or structures (e.g., objects placed on a surface). Second, the predetermined marker information provides a good prompt for the interactive segmentation model, effectively ensuring segmentation accuracy.

[0050] like Figure 4-Figure 5 As shown, a schematic diagram of a target marker in a target scene is shown.

[0051] Figure 4 and Figure 5 The target object in the example is a toy car. Specifically, the target toy car is placed on a ground covered with target markers to isolate the target toy car from the ground. The ground here can be a storage surface.

[0052] Step 202: For each captured image in the captured image set, perform the following generation steps:

[0053] Step 2021: Determine the marker segmentation information in the captured image using a pre-trained marker segmentation model.

[0054] In some embodiments, the execution entity may utilize a pre-trained marker segmentation model to determine marker segmentation information in the captured image. The marker segmentation model may be a neural network model that segments target markers in the captured image. The marker segmentation information may be segmentation information related to the target marker in the captured image. In practice, the marker segmentation information may be a set of marker coordinates corresponding to the target marker in the captured image. Each set of marker coordinates may represent the location of the target marker in the captured image.

[0055] In some optional implementations of some embodiments, the aforementioned marker segmentation model is a non-interactive segmentation model for segmenting markers. A non-interactive segmentation model may be a neural network model that does not require prompt information (mask instructions) to guide segmentation. In practice, the non-interactive segmentation model may be, but is not limited to, one of the following: an HRNet model, a Deeplabv3++ model, or a U-Net++ model.

[0056] Optionally, the execution entity may directly input the captured image into the non-interactive segmentation model to generate marker segmentation information.

[0057] In some optional implementations of some embodiments, after step 2021, the steps further include:

[0058] In the first step, in response to determining that there is a pre-trained second interactive segmentation model for segmenting objects, the execution entity may generate a second prompt image corresponding to the captured image. The second prompt image includes a second positive point set and a second negative point set for the target object, and the target marker corresponds to a surface with a plurality of regularly distributed marking information. The second positive point is a marking point on the surface of the target object. The second negative point is a marking point in the marking information. The second interactive segmentation model is an interactive segmentation model for segmenting the target object. The second positive point may be a point on the target object in the captured image. The second negative point is a point on the target marker in the captured image. For example, the color corresponding to the second positive point is green. The color corresponding to the second negative point is red.

[0059] In a second step, the execution entity may input the second prompt image and the captured image into the second interactive segmentation model to generate the object segmentation information.

[0060] Step 2022: Generate object segmentation information corresponding to each object in the captured image based on the marker segmentation information.

[0061] In some embodiments, the execution entity may generate item segmentation information corresponding to each item in the captured image based on the marker segmentation information. The items may be the items displayed in the captured image. For example, the items may include items corresponding to the placement surface, background items, and target items. The item segmentation information may indicate the location of each item corresponding to the item information in the captured image. The item segmentation information may include a set of item coordinates corresponding to each item.

[0062] As an example, the execution entity may first subtract the image content corresponding to the marker segmentation information from the captured image to obtain a post-subtraction captured image. The image content coordinates in the post-subtraction captured image are then used as the object segmentation information corresponding to each object.

[0063] In some optional implementations of some embodiments, the execution entity may set the image region corresponding to the marker segmentation information in the captured image as the background region, and set the remaining image regions as the foreground region, thereby obtaining foreground image information as the item segmentation information corresponding to each item. The content corresponding to the background region may be the background content in the captured image, and the content corresponding to the foreground region may be the foreground content in the captured image.

[0064] Step 203 : constructing a three-dimensional structure of the target object based on the obtained object segmentation information set.

[0065] In some embodiments, the execution entity may execute a three-dimensional structure construction for the target object based on the obtained object segmentation information set, wherein the result of the three-dimensional structure construction is the three-dimensional structure information corresponding to the target object.

[0066] As an example, first, for each item segmentation information in the above-mentioned item segmentation information set, the segmentation model corresponding to the target item is used to segment the image content corresponding to the target item from the image content corresponding to the item segmentation information. Then, based on the obtained image content set, the NeuS model is used to perform three-dimensional structure construction for the above-mentioned target item. Among them, the NeuS model is an innovative neural surface reconstruction method, which is suitable for achieving high-fidelity 3D reconstruction from 2D images, especially for scenes with complex structures and self-occlusion. The core of NeuS lies in its combination of signed distance function (SDF) and innovative volume rendering technology, which can achieve more accurate surface construction without mask supervision.

[0067] In some optional implementations of some embodiments, the execution entity may execute the three-dimensional structure construction for the target object based on the obtained object segmentation information set, including the following steps:

[0068] In the first step, based on the aforementioned object segmentation information set, a 3D structure construction model is used to generate 3D structure construction information. The 3D structure construction model can be a model that constructs the 3D structure of the target object. For example, the 3D structure construction model can be a NeuS model. The 3D structure construction information includes the 3D structure information corresponding to each object.

[0069] As an example, the aforementioned execution entity may input a shooting orientation set, an object segmentation information set, and a captured image set into the NeuS model to generate 3D structural information corresponding to each object as 3D structure construction information. The shooting orientation refers to the orientation at which the target object is captured during the image capture process. Each shooting orientation in the shooting orientation set can be predetermined or different. There is a one-to-one correspondence between the shooting orientations in the shooting orientation set and the captured images in the captured image set.

[0070] In the second step, the plane information corresponding to at least one isolated plane is removed from the three-dimensional structure construction information to obtain the three-dimensional structure information corresponding to the target object. An isolated plane may be an object present in the generated captured image due to insufficient structural information corresponding to the object. Here, the captured image set is captured specifically for the target object, and therefore contains object information for the target object at all captured orientations. However, the remaining objects in the captured image are not captured specifically for the target object, resulting in missing object information. Therefore, the missing object information only results in an isolated plane.

[0071] In some optional implementations of some embodiments, the execution entity may execute the three-dimensional structure construction for the target object based on the obtained object segmentation information set, including the following steps:

[0072] In the first step, for each item segmentation information in the item segmentation information set, the segmentation information corresponding to the target item is selected from the item segmentation information as the target segmentation information.

[0073] As an example, the execution entity inputs the image content corresponding to the object segmentation information into a target object recognition model to filter out the segmentation information corresponding to the target object as the target segmentation information. The target object recognition model may be a neural network model that recognizes the target object. For example, the target object recognition model may be a YOLO v8 model.

[0074] In the second step, the shooting orientation set corresponding to the above-mentioned shooting image set, the obtained object segmentation information set, and the above-mentioned shooting image set are input into the 3D structure construction model to generate the 3D structure information corresponding to the above-mentioned target object. Among them, the shooting orientation set corresponds to the above-mentioned different shooting orientations.

[0075] like Figure 6 As shown, a general schematic diagram of the method for three-dimensional construction of an object is shown.

[0076] First, a set of captured images is input. Each image in the set represents a photograph of the target object at different orientations. For each captured image, a corresponding first prompt image (referred to as the prompt in the figure) is generated. The prompt image and the captured image are subjected to SAM segmentation to generate object segmentation information. Using the NeuS model, a 3D object reconstruction is performed based on the resulting object segmentation information set.

[0077] The above-mentioned various embodiments of the present disclosure have the following beneficial effects: through the object three-dimensional construction method of some embodiments of the present disclosure, the placement surface and the target object can be accurately and efficiently segmented to facilitate the three-dimensional structure construction of the target object. Specifically, in the process of constructing the three-dimensional structure of the target object, the construction result is often that the target object and the placement surface are integrated, resulting in the need to manually use model editing tools to cut out the structure corresponding to the placement surface, which is time-consuming and labor-intensive, resulting in low efficiency of three-dimensional construction. Based on this, the object three-dimensional construction method of some embodiments of the present disclosure, first, obtains a set of photographed images corresponding to different shooting orientations of the target object in the target scene, wherein the above-mentioned target scene is a scene in which the above-mentioned target object is placed on the target marker and the above-mentioned target marker is placed on the placement surface, and the above-mentioned target marker is used to assist in segmenting the target object information in the photographed image. Here, by obtaining a set of photographed images, the structural information corresponding to the target object in different shooting orientations is obtained. By placing a target marker between the target item and the storage surface, the target marker is used to achieve spatial isolation between the target item and the storage surface and segment the target marker, thereby avoiding the problem of contact adhesion between the target item and the storage surface during the three-dimensional construction process. Then, for each captured image in the above-mentioned captured image set, the following generation steps are performed: First, using a pre-trained marker segmentation model, the marker segmentation information in the above-mentioned captured image can be accurately determined to achieve object segmentation for the target marker in the captured image. Second, based on the above-mentioned marker segmentation information, item segmentation information corresponding to each item in the above-mentioned captured image can be accurately generated. Here, the marker segmentation information can be used to segment the target marked items to achieve spatial isolation between the storage surface and the target item. Finally, based on the obtained item segmentation information set, a three-dimensional structure construction is performed for the above-mentioned target item to accurately and efficiently generate the corresponding three-dimensional structure information of the target item. In summary, spatial isolation is achieved by adding a target marker between the target item and the storage surface. Based on this, the segmentation and recognition of target markers can be used to accurately determine the object information in the corresponding image of the target object without being affected by the placement surface, so as to facilitate the construction of three-dimensional structures.

[0078] Further references Figure 7 , shows a process 400 of another embodiment of the method for constructing a three-dimensional object according to the present disclosure. The method for constructing a three-dimensional object includes the following steps:

[0079] Step 701: Acquire a set of images of a target object in a target scene corresponding to different shooting orientations.

[0080] Step 702: For each captured image in the captured image set, perform the following generation steps:

[0081] Step 7021: Generate a first prompt image corresponding to the captured image.

[0082] In some embodiments, the execution entity (e.g. Figure 1 The electronic device 101 shown) can generate a first prompt image corresponding to the above-mentioned captured image. The first prompt image contains a first positive point set and a first negative point set for the marking information. There are multiple marking information with a regular distribution pattern on the surface corresponding to the above-mentioned target marker, and the first positive point is a marking point in the marking information. The first negative point is a marking point that is not in the marking information. Each first positive point in the first positive point set is randomly selected and marked from at least one marking information corresponding to the target marker information in the captured image. The first positive point can be the center point in the marking information. And the color of the point corresponding to the first positive point can be pre-set. For example, the color corresponding to the first positive point is green. The first negative point may not be a point on the marking information. And the color of the point corresponding to the first negative point can be pre-set. For example, the color corresponding to the first negative point is red.

[0083] The first positive point set and the second negative point set in the first prompt image can enable the subsequent first interactive segmentation model to clearly learn the feature information corresponding to the target marker to achieve accurate segmentation of the target marker.

[0084] See also Figure 8 , shows a schematic image diagram corresponding to the first prompt image.

[0085] like Figure 8 As shown, cube 802 is the target object. Item 801 is the target marker. Point 804 is a positive example point. Points 803 and 805 are negative example points.

[0086] Step 7022: Input the first prompt image and the captured image into the first interactive segmentation model to generate the marker segmentation information.

[0087] In some embodiments, the execution entity may input the first prompt image and the captured image into the first interactive segmentation model to generate the marker segmentation information. The marker segmentation model is a first interactive segmentation model for segmenting markers. The first interactive segmentation model may be a neural network model that performs object segmentation based on the interactive information. For example, the first interactive segmentation model may be a SAM (Segment Anything Model) model.

[0088] Step 7023: Generate object segmentation information corresponding to each object in the captured image based on the marker segmentation information.

[0089] Step 703 : constructing a three-dimensional structure of the target object based on the obtained object segmentation information set.

[0090] In some embodiments, the specific implementation of steps 701, 7023 and 703 and the technical effects thereof can be referred to. Figure 2 Steps 201, 2021 and 203 in the corresponding embodiment are not described again here.

[0091] from Figure 7 It can be seen that Figure 2 Compared with the description of some corresponding embodiments, Figure 7 In the corresponding process 700 of the method for constructing a three-dimensional object in some embodiments, by setting positive points and negative points in the first prompt image, the first interactive segmentation model can more accurately segment the image information corresponding to the target marker.

[0092] Further references Figure 9 As an implementation of the methods shown in the above figures, the present disclosure provides some embodiments of a device for constructing a three-dimensional object. These device embodiments are similar to Figure 2 Corresponding to the method embodiments shown, the three-dimensional object construction device can be specifically applied to various electronic devices.

[0093] like Figure 9 As shown, a device 900 for constructing a three-dimensional object includes: an acquisition unit 901, a first execution unit 902, and a second execution unit 903. The acquisition unit 901 is configured to acquire a set of captured images of a target object in a target scene corresponding to different shooting orientations, wherein the target scene is a scene in which the target object is placed on a target marker, and the target marker is placed on a storage surface, and the target marker is used to assist in segmenting target object information in the captured images; the first execution unit 902 is configured to perform the following generation steps for each captured image in the captured image set: using a pre-trained marker segmentation model to determine marker segmentation information in the captured image; and generating object segmentation information corresponding to each object in the captured image based on the marker segmentation information; and the second execution unit 903 is configured to execute three-dimensional structure construction for the target object based on the obtained set of object segmentation information.

[0094] In some optional implementations of some embodiments, the above-mentioned marker segmentation model is a first interactive segmentation model for segmenting markers; and the above-mentioned first execution unit 902 can be further configured to: generate a first prompt image corresponding to the above-mentioned captured image, wherein the above-mentioned first prompt image has a first positive point set and a first negative point set for the marking information, wherein the above-mentioned target marker corresponds to a surface with multiple regularly distributed marking information, the first positive point is a marking point in the marking information, and the first negative point is a marking point not in the marking information; input the above-mentioned first prompt image and the above-mentioned captured image into the above-mentioned first interactive segmentation model to generate the above-mentioned marker segmentation information.

[0095] In some optional implementations of some embodiments, the above-mentioned marker segmentation model is a non-interactive segmentation model for segmenting markers; and the above-mentioned first execution unit 902 can be further configured to: input the above-mentioned captured image into the above-mentioned non-interactive segmentation model to generate marker segmentation information.

[0096] In some optional implementations of some embodiments, the above-mentioned first execution unit 902 can be further configured to: set the image area corresponding to the above-mentioned marker segmentation information in the above-mentioned captured image as the background area, set the remaining image areas as the foreground area, and obtain foreground image information as the item segmentation information corresponding to each item.

[0097] In some optional implementations of some embodiments, the above-mentioned second execution unit 903 can be further configured to: generate three-dimensional structure construction information by using a three-dimensional structure construction model based on the above-mentioned object segmentation information set; remove the plane information corresponding to at least one isolated plane from the above-mentioned three-dimensional structure construction information to obtain the three-dimensional structure information corresponding to the above-mentioned target object.

[0098] In some optional implementations of some embodiments, the first execution unit 902 may be further configured to: in response to determining that there is a pre-trained second interactive segmentation model for segmenting objects, generate a second prompt image corresponding to the captured image, wherein the second prompt image contains a second positive point set and a second negative point set for the target object, and the target marker corresponds to a surface containing a plurality of regularly distributed marking information, the second positive point is a marking point on the surface of the target object, and the second negative point is a marking point in the marking information; and input the second prompt image and the captured image into the second interactive segmentation model to generate the object segmentation information.

[0099] In some optional implementations of some embodiments, the above-mentioned second execution unit 903 can be further configured to: for each item segmentation information in the above-mentioned item segmentation information set, filter out the segmentation information corresponding to the above-mentioned target item from the above-mentioned item segmentation information as the target segmentation information; input the shooting orientation set corresponding to the above-mentioned shooting image set, the obtained target segmentation information set and the above-mentioned shooting image set into the three-dimensional structure construction model to generate the three-dimensional structure information corresponding to the above-mentioned target item.

[0100] It is understood that the units described in the object three-dimensional construction device 900 are similar to those in the reference Figure 2 Therefore, the operations, features and beneficial effects described above for the method are also applicable to the object three-dimensional construction device 900 and the units included therein, and will not be repeated here.

[0101] Reference below Figure 10 , which shows an electronic device (eg, Figure 1 Schematic diagram of the structure of the electronic device 101)1000. Figure 10 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.

[0102] like Figure 10 As shown, the electronic device 1000 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 1001, which can perform various appropriate actions and processes according to a program stored in a read-only memory 1002 or a program loaded from a storage device 1008 into a random access memory 1003. Various programs and data required for the operation of the electronic device 1000 are also stored in the random access memory 1003. The processing device 1001, the read-only memory 1002, and the random access memory 1003 are connected to each other via a bus 1004. An input / output interface 1005 is also connected to the bus 1004.

[0103] Typically, the following devices may be connected to the input / output interface 1005: an input device 1006 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 1007 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 1008 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 may allow the electronic device 1000 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 10 The electronic device 1000 is shown with various devices, but it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed instead. Figure 10Each block shown in the figure may represent one device, or may represent multiple devices as needed.

[0104] In particular, according to some embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of the present disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for executing the method shown in the flowchart. In some such embodiments, the computer program can be downloaded and installed from a network via the communication device 1009, or installed from the storage device 1008, or installed from the read-only memory 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of some embodiments of the present disclosure are performed.

[0105] It should be noted that in some embodiments of the present disclosure, the computer-readable medium mentioned above may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In some embodiments of the present disclosure, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device, or device. In some embodiments of the present disclosure, the computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.

[0106] In some embodiments, the client and server can communicate using any currently known or future developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.

[0107] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device. The computer-readable medium carries one or more programs, and when the one or more programs are executed by the electronic device, the electronic device: obtains a set of captured images of a target object in a target scene corresponding to different shooting orientations, wherein the target scene is a scene in which the target object is placed on a target marker and the target marker is placed on a placement surface, and the target marker is used to assist in segmenting the target object information in the captured image; for each captured image in the captured image set, performs the following generation steps: determines the marker segmentation information in the captured image using a pre-trained marker segmentation model; generates object segmentation information corresponding to each object in the captured image based on the marker segmentation information; and performs three-dimensional structure construction for the target object based on the obtained object segmentation information set.

[0108] Computer program code for performing the operations of some embodiments of the present disclosure may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0109] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0110] The units described in some embodiments of the present disclosure may be implemented in software or in hardware. The described units may also be provided in a processor. For example, they may be described as follows: a processor includes an acquisition unit, a first execution unit, and a second execution unit. The names of these units do not, in some cases, constitute limitations on the units themselves. For example, the acquisition unit may also be described as a "unit for acquiring a set of photographed images of a target object corresponding to different shooting orientations in a target scene."

[0111] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0112] Some embodiments of the present disclosure further provide a computer program product, comprising a computer program, which implements any of the above-mentioned methods for constructing a three-dimensional object when executed by a processor.

[0113] The above description is only an illustration of some preferred embodiments of the present disclosure and the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but should also cover other technical solutions formed by any combination of the above-mentioned technical features or their equivalent features without departing from the above-mentioned inventive concept. For example, the above-mentioned features are replaced with (but not limited to) technical features with similar functions disclosed in the embodiments of the present disclosure.

Claims

1. A method for constructing a three-dimensional object, comprising: Acquire a set of images of a target object in a target scene corresponding to different shooting orientations, wherein the target scene is a scene in which the target object is placed on a target marker and the target marker is placed on a placement surface, and the target marker is used to assist in segmenting target object information from the captured images; For each captured image in the captured image set, the following generation steps are performed: Determining marker segmentation information in the captured image using a pre-trained marker segmentation model; generating object segmentation information corresponding to each object in the captured image according to the marker segmentation information; Based on the obtained object segmentation information set, a three-dimensional structure of the target object is constructed.

2. The method according to claim 1, wherein The marker segmentation model is a first interactive segmentation model for segmenting markers; as well as The determining the marker segmentation information in the captured image by using a pre-trained marker segmentation model includes: Generate a first prompt image corresponding to the captured image, wherein the first prompt image has a first positive example point set and a first negative example point set for the marking information, wherein the target marker object has a plurality of marking information with a regular distribution pattern corresponding to the surface, the first positive example point is a marking point in the marking information, and the first negative example point is a marking point not in the marking information; The first prompt image and the captured image are input into the first interactive segmentation model to generate the marker segmentation information.

3. The method according to claim 1, wherein The marker segmentation model is a non-interactive segmentation model for segmenting markers; as well as The determining the marker segmentation information in the captured image by using a pre-trained marker segmentation model includes: The captured image is input into the non-interactive segmentation model to generate marker segmentation information.

4. The method according to claim 1, wherein Generating object segmentation information corresponding to each object in the captured image according to the marker segmentation information includes: The image area corresponding to the marker segmentation information in the captured image is set as the background area, and the remaining image areas are set as the foreground area to obtain foreground image information as the item segmentation information corresponding to each item.

5. The method according to claim 1, wherein The step of constructing a three-dimensional structure of the target object based on the obtained object segmentation information set includes: generating three-dimensional structure construction information by using a three-dimensional structure construction model according to the object segmentation information set; Plane information corresponding to at least one isolated plane is removed from the three-dimensional structure construction information to obtain three-dimensional structure information corresponding to the target object.

6. The method according to claim 1, wherein After determining the marker segmentation information in the captured image using the pre-trained marker segmentation model, the method further includes: In response to determining that a pre-trained second interactive segmentation model for segmenting objects exists, generating a second prompt image corresponding to the captured image, wherein the second prompt image includes a second positive point set and a second negative point set for the target object, and the target marker object includes a plurality of regularly distributed marking information corresponding to the surface, the second positive points are marking points on the surface of the target object, and the second negative points are marking points in the marking information; The second prompt image and the captured image are input into the second interactive segmentation model to generate the object segmentation information.

7. The method according to claim 1, wherein The step of constructing a three-dimensional structure of the target object based on the obtained object segmentation information set includes: For each item segmentation information in the item segmentation information set, filtering out the segmentation information corresponding to the target item from the item segmentation information as target segmentation information; The shooting orientation set corresponding to the shot image set, the obtained target segmentation information set, and the shot image set are input into a three-dimensional structure construction model to generate three-dimensional structure information corresponding to the target object.

8. A device for constructing a three-dimensional object, comprising: an acquisition unit configured to acquire a set of images of a target object in a target scene corresponding to different shooting orientations, wherein the target scene is a scene in which the target object is placed on a target marker and the target marker is placed on a placement surface, and the target marker is used to assist in segmenting target object information from the captured images; The first execution unit is configured to perform the following generation steps for each captured image in the captured image set: determine marker segmentation information in the captured image using a pre-trained marker segmentation model; and generate object segmentation information corresponding to each object in the captured image based on the marker segmentation information; The second execution unit is configured to execute three-dimensional structure construction for the target object according to the obtained object segmentation information set.

9. An electronic device comprising: one or more processors; a storage device having one or more programs stored thereon, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 7.

10. A computer-readable medium having a computer program stored thereon, wherein: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

11. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 7.