Three-dimensional image reconstruction method and device based on large model, training method and device, equipment and medium

Through the three-dimensional image reconstruction method based on the big model, implicit prompt information is used for inference segmentation and geometric reconstruction, combined with the global inference of the big language model, the problem of low processing and computing efficiency of implicit language instruction in the existing technology is solved, and efficient and accurate three-dimensional image reconstruction is achieved.

CN120198592APending Publication Date: 2025-06-24BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510337028.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The existing three-dimensional image reconstruction technology has shortcomings in processing implicit language instructions and improving computing efficiency and accuracy. Especially in applications in the fields of metaverse, virtual reality and augmented reality, it is difficult to meet the needs of efficient and accurate three-dimensional image reconstruction.

Method used

A three-dimensional image reconstruction method based on large models is adopted, and the reconstructed images are inferred segmented through implicit prompt information, combined with geometric reconstruction, fused the segmented image and the three-dimensional grid model, and finally input the results into the large language model to obtain the three-dimensional reconstruction results of the target object.

Benefits of technology

This method can effectively reduce the computational amount, improve the efficiency and accuracy of three-dimensional image reconstruction, is suitable for processing implicit language instructions, and achieve higher applicability in multiple scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198592A_ABST
    Figure CN120198592A_ABST
Patent Text Reader

Abstract

The invention provides a three-dimensional image reconstruction method based on a large model, and relates to the technical field of artificial intelligence, in particular to the technical fields of computer vision, deep learning, large models, meta universe, virtual reality, augmented reality and the like. According to the specific implementation scheme, reasoning segmentation is carried out on a to-be-reconstructed image based on implicit prompt information, at least one candidate object in the to-be-reconstructed image is determined, a mask is added to the at least one candidate object, and a segmented image is obtained; geometric reconstruction is carried out on the to-be-reconstructed image to obtain a three-dimensional grid model, and the three-dimensional grid model represents geometric information of the at least one candidate object and geometric information of a reconstruction scene of the to-be-reconstructed image; fusing the segmented image and the three-dimensional grid model to obtain a three-dimensional reconstruction result of the at least one candidate object; and inputting the implicit prompt information and the three-dimensional reconstruction result of the at least one candidate object into the first large language model, and outputting the three-dimensional reconstruction result of the target object indicated by the implicit prompt information. The invention further provides a training method and device based on the large model, electronic equipment and a storage medium.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technologies, and particularly to technical fields such as computer vision, deep learning, large models, metaverse, virtual reality, and augmented reality, and can be applied to scenarios such as artificial intelligence-based content generation. More specifically, the present disclosure provides a three-dimensional image reconstruction method, a training method, a device, an electronic device, and a storage medium based on a large model. Background Art

[0002] Three-dimensional image reconstruction technology is a technology that can extract three-dimensional information from two-dimensional images or data sets and construct a three-dimensional model or scene based on this. With the development of artificial intelligence technology, three-dimensional image reconstruction technology has become increasingly mature and has broad prospects in technical fields such as computer vision, deep learning, large models, metaverse, and augmented reality. How to achieve high-quality and high-efficiency three-dimensional image reconstruction has become a key issue in this field. Summary of the Invention

[0003] The present disclosure provides a three-dimensional image reconstruction method, a training method, a device, a device, and a storage medium based on a large model.

[0004] According to one aspect of the present disclosure, a three-dimensional image reconstruction method based on a large model is provided, including: performing inference segmentation on an image to be reconstructed based on implicit prompt information, determining at least one candidate object in the image to be reconstructed and adding a mask to the at least one candidate object to obtain a segmented image; performing geometric reconstruction on the image to be reconstructed to obtain a three-dimensional mesh model, where the three-dimensional mesh model represents the geometric information of at least one candidate object and the geometric information of the reconstruction scene of the image to be reconstructed; fusing the segmented image and the three-dimensional mesh model to obtain a three-dimensional reconstruction result of at least one candidate object; and inputting the implicit prompt information and the three-dimensional reconstruction result of at least one candidate object into a first large language model to output a three-dimensional reconstruction result of a target object indicated by the implicit prompt information.

[0005] According to another aspect of the present disclosure, a model training method is provided, including: obtaining a first training data set, where the first training data set includes a data pair composed of a first mask and first training prompt information, and the first training prompt information implicitly includes the category represented by the first mask; and training a model based on the first training data set to obtain a trained inference segmentation model, where the trained inference segmentation model is used to perform inference segmentation on an image to be reconstructed based on implicit prompt information to obtain a segmented image.

[0006] According to another aspect of the present disclosure, there is provided a three-dimensional image reconstruction device based on a large model, including: an inference segmentation module for performing inference segmentation on an image to be reconstructed based on implicit prompt information, determining at least one candidate object in the image to be reconstructed and adding masks to the at least one candidate object to obtain a segmented image; a geometric reconstruction module for performing geometric reconstruction on the image to be reconstructed to obtain a three-dimensional mesh model, the three-dimensional mesh model characterizing the geometric information of at least one candidate object and the geometric information of the scene where the image to be reconstructed is located; a fusion module for fusing the segmented image and the three-dimensional mesh model to obtain three-dimensional reconstruction results of at least one candidate object; and an input / output module for inputting the implicit prompt information and the three-dimensional reconstruction results of the at least one candidate object into a first large language model and outputting three-dimensional reconstruction results of a target object indicated by the implicit prompt information.

[0007] According to another aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method provided by the present disclosure.

[0008] According to another aspect of the present disclosure, there is provided a model training device, including: an acquisition module for acquiring a first training data set, the first training data set including a data pair composed of a first mask and first training prompt information, the first training prompt information implicitly including the category characterized by the first mask; and a training module for training a model based on the first training data set to obtain a trained inference segmentation model, wherein the trained inference segmentation model is used to perform inference segmentation on an image to be reconstructed based on implicit prompt information in the three-dimensional image reconstruction method provided by the present disclosure to obtain a segmented image.

[0009] According to another aspect of the present disclosure, there is provided an electronic device, including at least one display device; and a three-dimensional image reconstruction device based on a large model, communicatively connected to the at least one display device; the three-dimensional image reconstruction device based on a large model is used to execute the three-dimensional image reconstruction method provided by the present disclosure to obtain three-dimensional reconstruction results and output the three-dimensional reconstruction results to the at least one display device for display.

[0010] According to another aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method provided by the present disclosure.

[0011] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the method provided according to the present disclosure.

[0012] According to another aspect of the present disclosure, there is provided a computer program product including a computer program which, when executed by a processor, implements the method provided according to the present disclosure.

[0013] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. Description of the Drawings

[0014] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0015] Figure 1 is a schematic diagram of an exemplary system architecture to which the three-dimensional image reconstruction method and apparatus based on a large model according to an embodiment of the present disclosure can be applied;

[0016] Figure 2 is a flowchart of a three-dimensional image reconstruction method based on a large model according to an embodiment of the present disclosure;

[0017] Figure 3 schematically shows the process of an inference segmentation method for an image to be reconstructed according to an embodiment of the present disclosure;

[0018] Figure 4 is a flowchart of a three-dimensional image reconstruction method based on a large model according to another embodiment of the present disclosure;

[0019] Figure 5 is a flowchart of a three-dimensional image reconstruction method based on a large model according to still another embodiment of the present disclosure;

[0020] Figure 6 is a flowchart of model training according to an embodiment of the present disclosure;

[0021] Figure 7 is a block diagram of a three-dimensional image reconstruction apparatus based on a large model according to an embodiment of the present disclosure;

[0022] Figure 8 is a block diagram of a model training apparatus based on a large model according to an embodiment of the present disclosure;

[0023] Figure 9 is a block diagram of an electronic device according to an embodiment of the present disclosure; and

[0024] Figure 10A schematic block diagram of an exemplary electronic device 1000 that can be used to implement embodiments of the present disclosure is shown. Detailed implementation manners

[0025] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to assist understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, descriptions of well-known functions and structures are omitted in the following description for clarity and conciseness.

[0026] In one example, three-dimensional image reconstruction mainly relies on explicit instructions (displaying prompt information) or queryable maps to implement instruction-guided online three-dimensional image reconstruction. This method first applies a Simultaneous Localization and Mapping (SLAM) system, and then matches the language instructions with a predefined map to obtain the three-dimensional reconstruction result corresponding to the instructions. Since this method obtains the three-dimensional image reconstruction result of the corresponding object based on a queryable map, it often needs to first obtain the three-dimensional image reconstruction results of all objects, and then match the objects with the language instructions, including a large amount of information unrelated to the language instructions, wasting some computational power on unrelated objects, limiting its performance on the target object, with a large amount of computation, low three-dimensional image reconstruction accuracy and reconstruction efficiency. In addition, this method can only process explicit language instructions. For example, the explicit language instruction is "Help me find the cup", but it is difficult to respond to implicit language instructions such as "I'm a little thirsty. Help me find something to drink", and the obtained three-dimensional image reconstruction result has poor accuracy, limiting its application in actual scenarios.

[0027] In view of this, embodiments of the present disclosure provide a three-dimensional image reconstruction method based on a large model, including: performing inference segmentation on the image to be reconstructed based on implicit prompt information, determining at least one candidate object in the image to be reconstructed and adding masks to the at least one candidate object to obtain a segmented image. Performing geometric reconstruction on the image to be reconstructed to obtain a three-dimensional mesh model, where the three-dimensional mesh model represents the geometric information of the at least one candidate object and the geometric information of the reconstruction scene of the image to be reconstructed. Fusing the segmented image and the three-dimensional mesh model to obtain the three-dimensional reconstruction result of the at least one candidate object. Inputting the implicit prompt information and the three-dimensional reconstruction result of the at least one candidate object into a first large language model, and outputting the three-dimensional reconstruction result of the target object indicated by the implicit prompt information.

[0028] Figure 1 It is a schematic diagram of an exemplary system architecture to which the three-dimensional image reconstruction method and device based on a large model can be applied according to an embodiment of the present disclosure. It should be noted that Figure 1The figure shown is only an example of the system architecture to which the embodiments of the present disclosure can be applied, to help those skilled in the art understand the technical content of the present disclosure, but it does not mean that the embodiments of the present disclosure cannot be used in other devices, systems, environments or scenarios.

[0029] As Figure 1 shown, the system architecture 100 according to this embodiment may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.

[0030] Users can use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. The terminal devices 101, 102, 103 may be various electronic devices with a display screen and supporting web browsing, including but not limited to smartphones, tablets, laptop computers, and desktop computers, etc.

[0031] The server 105 may be a server providing various services, such as a background management server (only an example) that provides support for three-dimensional image reconstruction of the to-be-reconstructed images sent by users using the terminal devices 101, 102, 103. The background management server may perform three-dimensional image reconstruction on the received to-be-reconstructed images and feedback the three-dimensional image reconstruction results to the terminal devices 101, 102, 103 for display on the display screens of the terminal devices 101, 102, 103.

[0032] It should be noted that the three-dimensional image reconstruction method based on a large model provided by the embodiments of the present disclosure can generally be executed by the server 105. Correspondingly, the three-dimensional image reconstruction device based on a large model provided by the embodiments of the present disclosure can generally be set in the server 105. The three-dimensional image reconstruction method based on a large model provided by the embodiments of the present disclosure can also be executed by a server or a server cluster different from the server 105 and capable of communicating with the terminal devices 101, 102, 103 and / or the server 105. Correspondingly, the three-dimensional image reconstruction device based on a large model provided by the embodiments of the present disclosure can also be set in a server or a server cluster different from the server 105 and capable of communicating with the terminal devices 101, 102, 103 and / or the server 105.

[0033] It should be understood that Figure 1 the number and types of the terminal devices, network, and server in

[0034] Figure 2It is a flowchart of a 3D image reconstruction method based on a large model according to an embodiment of the present disclosure.

[0035] As Figure 2 shown, the method 200 may include operation S210 to operation S240.

[0036] In operation S210, perform inference segmentation on the image to be reconstructed based on implicit prompt information, determine at least one candidate object in the image to be reconstructed, and add a mask to the at least one candidate object to obtain a segmented image.

[0037] According to an embodiment of the present disclosure, the image to be reconstructed may be a single image or multiple images with time series in a video.

[0038] According to an embodiment of the present disclosure, implicit prompt information can be understood as it is difficult to directly obtain the true meaning of implicit prompt information from the text of implicit prompt information, and it is necessary to perform parsing and reasoning on the implicit prompt information to obtain the true meaning of the implicit prompt information. For example, the language instruction "Help me find a cup" issued by the user is an explicit prompt information, and its true meaning can be obtained from the literal meaning. The language instruction "I'm a bit thirsty. Help me find something to drink" issued by the user is an implicit prompt information, and it is not clear from the literal meaning what specific thing is wanted to drink, that is, it is difficult to directly obtain its true meaning from the literal meaning.

[0039] According to an embodiment of the present disclosure, performing inference segmentation on the image to be reconstructed can be understood as fusing multi-modal data (implicit prompt information and the image to be reconstructed), understanding the implicit intention of the user by parsing complex implicit prompt information, and segmenting the graph based on the implicit intention of the user to achieve the segmentation of the objects included in the image to be reconstructed.

[0040] According to an embodiment of the present disclosure, the image to be reconstructed can be processed to obtain an RGB image corresponding to the image to be reconstructed for providing the color of the object, and then perform inference segmentation on the RGB image based on the implicit prompt information. Combining the visual information such as color, texture, and shape in the RGB image and the semantic information in the implicit prompt information, pixel-level classification of different regions in the RGB image can be obtained. A mask (mask) can be added to each pixel point, which can also be called a specific class label. The mask can be a binary image, that is, the generated segmented image can be a binary image. For example, the region in the foreground of the image to be reconstructed that matches the implicit prompt information is marked as 1, and the background is marked as 0.

[0041] According to an embodiment of the present disclosure, a candidate object can be understood as an object in the image to be reconstructed whose matching degree with the true meaning of the implicit prompt information is higher than the matching degree threshold. This object may be used as the target object and needs to be further matched subsequently to determine the target object from the candidate objects.

[0042] In operation S220, geometric reconstruction is performed on the image to be reconstructed to obtain a three-dimensional mesh model.

[0043] According to an embodiment of the present disclosure, geometric reconstruction can be understood as reconstructing a three-dimensional mesh model of an object by using mathematical and computer graphics methods with information such as feature points, contours, and textures in the image to be reconstructed.

[0044] According to an embodiment of the present disclosure, the three-dimensional mesh model can represent the geometric information of at least one candidate object and the geometric information of the reconstruction scene of the image to be reconstructed. The three-dimensional mesh model can precisely define the shape and contour of an object or scene in three-dimensional space through the combination of its vertices, edges, and faces. The three-dimensional mesh model provides an infrastructure for the rendering engine, enabling the rendering algorithm to calculate effects such as the color, lighting, and shadows of each vertex, thereby generating realistic three-dimensional images.

[0045] According to an embodiment of the present disclosure, a depth map of the image to be reconstructed can be generated based on the image to be reconstructed, and geometric reconstruction is performed based on the depth map. The depth map can be an image or image channel containing information related to the distance from the viewpoint to the surface of the scene object. It is similar to a grayscale image, but each pixel value represents the actual distance of the sensor from the object. The depth map mainly provides the spatial position information of the object and is an important basis for tasks such as three-dimensional reconstruction, object recognition, and scene understanding.

[0046] It should be understood that both the RGB image and the depth map are generated based on the image to be reconstructed, that is, the RGB image and the depth map capture the same scene and reflect different aspects of characteristics. The RGB image provides color information, while the depth map provides spatial position information, and there is a one-to-one correspondence between them at the pixel level.

[0047] In operation S230, the segmented image and the three-dimensional mesh model are fused to obtain the three-dimensional reconstruction result of at least one candidate object.

[0048] According to an embodiment of the present disclosure, the segmented image is obtained by performing two-dimensional segmentation on the image to be reconstructed, and the three-dimensional mesh model is obtained based on three-dimensional geometric reconstruction. They are obtained based on different types of images and different methods, and are in different coordinate systems. During the fusion process, the segmented image and the three-dimensional mesh model can be converted into the same coordinate system, so as to combine the two-dimensional inference segmentation result with the three-dimensional geometric reconstruction result. The segmented image and the three-dimensional mesh model can be converted into the same coordinate system by using projection or inverse projection methods.

[0049] In operation S240, the implicit hint information and the three-dimensional reconstruction result of at least one candidate object are input into the first large language model, and the three-dimensional reconstruction result of the target object indicated by the implicit hint information is output.

[0050] According to an embodiment of the present disclosure, during the process of inferring and segmenting the image to be reconstructed, the field of view of the inference segmentation model may be limited, and the objects within the field of view of the current frame are seen, resulting in some candidate objects that conform to the implicit hint information. In order to determine the target object that best conforms to the implicit hint information from the candidate objects, the three-dimensional reconstruction results of at least one candidate object can be input into the first large language model for inference, thereby obtaining the final target object and the three-dimensional image reconstruction result of the target object. In order to enable the first large language model to better perform global inference on the three-dimensional reconstruction results of the candidate objects, the implicit hint information can be further input into the first large language model, and the global inference of the three-dimensional reconstruction results of the candidate objects can be further guided by reusing the implicit hint information.

[0051] Through the three-dimensional image reconstruction method of the embodiments of the present disclosure, inferring and segmenting the image to be reconstructed based on the implicit hint information, and combining the geometric reconstruction results to obtain the three-dimensional reconstruction results of the candidate objects in advance, without obtaining the reconstruction results of all objects in the image to be reconstructed, a large amount of information irrelevant to the hint information is proposed, reducing the computational amount and complexity of the three-dimensional image reconstruction, and improving the efficiency and accuracy of the three-dimensional image reconstruction. Moreover, adjusting the three-dimensional reconstruction results of the candidate objects based on the implicit hint information can better perform global inference on the three-dimensional reconstruction results of the candidate objects, further improving the accuracy of the three-dimensional image reconstruction. In addition, the three-dimensional image reconstruction method can utilize more complex implicit hint information for three-dimensional image reconstruction, improving the applicability of the three-dimensional image reconstruction, and enabling the three-dimensional image reconstruction method to be applicable to more scenarios.

[0052] The following will combine Figure 3 to schematically describe the inference segmentation method of the image to be reconstructed in the embodiments of the present disclosure. Figure 3 Schematically shows the flow of the inference segmentation method of the image to be reconstructed according to an embodiment of the present disclosure.

[0053] In some embodiments, inferring and segmenting the image to be reconstructed based on the implicit hint information may include:

[0054] Input the first encoded image obtained by visually encoding the image to be reconstructed and the implicit hint information into the large language model, and output the masks of at least one candidate object indicated by the implicit hint information.

[0055] Decode the second encoded image obtained by image encoding the image to be reconstructed based on the mask to generate a segmented image.

[0056] Such as Figure 3As shown, the RGB image 311a corresponding to the image to be reconstructed can be input into the vision encoder 314 for visual encoding. After projecting the visual encoding result, the first encoded image 311c can be obtained. The second encoded image 311c is input into the second large language model 315, and the implicit prompt information 316 is input into the second large language model 315. The second large language model 315 performs reasoning based on the input first encoded image 311c and implicit prompt information 316, obtains the masks 317 of at least one candidate object indicated by the output implicit prompt information, and inputs the masks 317 of at least one candidate object into the mask decoder 313.

[0057] The RGB image 311a corresponding to the image to be reconstructed can be input into the image encoder 312 for image encoding, and the second encoded image 311b can be obtained. The second encoded image 311b is input into the mask decoder 313.

[0058] The mask decoder 313 fuses the input first encoded image 311c and the masks 317 of at least one candidate object, and outputs the segmentation image 311d.

[0059] According to an embodiment of the present disclosure, the image encoder 312 can preprocess the RGB image 311a in advance, such as resizing, denoising, etc. Then, the RGB image can be encoded to achieve data compression.

[0060] According to an embodiment of the present disclosure, the vision encoder 314 can extract features from the RGB image 311a to obtain a series of features that can represent the image content. These features may include color, texture, shape, etc.

[0061] According to an embodiment of the present disclosure, the second large language model can integrate the multimodal input composed of the second encoded image 316 and the implicit prompt information 316 to form a unified representation. Based on this unified representation, the implicit prompt information is recognized, and the second encoded image is inferred according to the recognition result to obtain the masks 317 of at least one candidate object indicated by the implicit prompt information.

[0062] Through the three-dimensional image reconstruction method of the embodiments of the present disclosure, by combining image encoding and visual encoding, and based on the implicit prompt information, a large amount of information irrelevant to the implicit prompt information can be excluded, thereby improving the efficiency and accuracy of three-dimensional image reconstruction.

[0063] The following will combine Figure 4 A schematic description will be given of the large model-based three-dimensional image reconstruction method of another embodiment of the embodiments of the present disclosure. Figure 4It is a flowchart of a large model-based 3D image reconstruction method according to another embodiment of the present disclosure.

[0064] As Figure 4 shown, the method 400 may include operation S410 to operation S420.

[0065] In operation S410, according to the first quantity of the objects included in the image to be reconstructed and the scene corresponding to the image to be reconstructed, determine the second quantity of masks.

[0066] In operation S420, input the second encoded image and the implicit prompt information into the second large language model so that the second large language model outputs the second quantity of masks.

[0067] According to the implementation of the present disclosure, in 3D image reconstruction, the quantity of masks required for different reconstruction tasks may be different. In a simple 3D image reconstruction scenario, perhaps only one mask is needed to identify the target. In a complex 3D image reconstruction scenario, multiple masks may be needed to separately identify different targets. For example, in interactive applications such as medical image analysis and autonomous driving, the model's ability to output multiple masks can help users more intuitively understand the target distribution and relationships in the image. When dealing with certain abnormal situations, the input image may not contain the target to be segmented, and in this case, no mask may be needed to identify the target. Based on this, the quantity of output masks can be flexibly adjusted in advance based on different images to be reconstructed and scene requirements.

[0068] Through the 3D image reconstruction method of the embodiments of the present disclosure, corresponding quantities of masks are output based on different images to be reconstructed and scene requirements for 3D reconstruction, which can meet the requirements of different 3D image reconstruction tasks and improve the flexibility and adaptability of the 3D reconstruction method. Moreover, for different 3D image reconstruction task requirements, the quantity of masks is targeted. While reducing computing resources, it can improve the segmentation accuracy and efficiency for the same 3D image reconstruction task requirements, thereby improving the accuracy and efficiency of 3D image reconstruction.

[0069] Based on the above embodiments, in some embodiments, the training dataset of the second large language model includes a first training dataset, and the first training dataset includes data pairs composed of first masks and first training prompt information, and the first training prompt information implicitly includes the categories represented by the first masks.

[0070] According to the embodiments of the present disclosure, the second large language model can be trained by constructing specific data pairs, enabling the second large language model to have the ability to recognize implicit prompt information, that is, to recognize the implicit category information contained in the implicit prompt information. In this way, the prompt information does not need to contain explicit category information, but rather poses a requirement, and the specific target to be segmented is inferred by the second large language model.

[0071] For example, the process of constructing specific data can be as follows: for an image whose target category is known, the known target category can be input into the large model, and questions can be asked of the large model regarding the target represented by this target category. The questions can be asked in terms of functions and uses, and the name of the known target category cannot be explicitly included in the questions. By repeatedly asking questions in this way, implicit prompt information - mask data pairs are constructed and used as the training data for the second large language model to train the second large language model.

[0072] For example, since the quality of the mask output by the large model through questioning may not be high, a segmentation model can be used to refine the mask to obtain a high-quality mask.

[0073] Through the three-dimensional image reconstruction method of the embodiments of the present disclosure, by constructing data in a specific format to train the model, the model can be enabled to have the ability to recognize implicit prompt information, enhance the model's reasoning ability for implicit prompt information, and thus be able to better perform three-dimensional image reconstruction based on implicit prompt information.

[0074] Based on the above embodiments, in some embodiments, the training data set of the second large language model includes a second training data set. The second training data set includes data pairs composed of a second mask and second training prompt information. The second training prompt information implicitly includes the category represented by the second mask, and there is a one-to-one or one-to-many relationship between the second training prompt information and the second mask.

[0075] In the traditional solution, for an image, when performing inference segmentation, inputting one prompt information corresponds to outputting one mask, and the number of output masks is fixed. However, in three-dimensional image reconstruction, the number of masks required for different reconstruction tasks may vary.

[0076] According to the embodiments of the present disclosure, specific data can be constructed as the training data of the model to train the model, enabling the model to output one mask, multiple masks, or zero masks when inputting one implicit prompt information. The construction of specific data is similar to the above and will not be elaborated here.

[0077] Through the three-dimensional image reconstruction method of the embodiments of the present disclosure, by constructing data in a specific format to train the model, the model can be enabled to flexibly output the corresponding number of masks when inputting one implicit prompt information, improving the flexibility and adaptability of the three-dimensional reconstruction method.

[0078] The following will combine Figure 5 to schematically describe the large model-based three-dimensional image reconstruction method of another embodiment of the embodiments of the present disclosure. Figure 5A flowchart of a large model-based 3D image reconstruction method according to another embodiment of the present disclosure.

[0079] In some embodiments, the 3D image reconstruction method may further include:

[0080] Segment the text included in the implicit prompt information and convert the text into at least one semantic token.

[0081] Input at least one semantic token into a second large language model so that the second large language model outputs a mask of at least one candidate object indicated by the implicit prompt information based on the second encoded image and the at least one semantic token.

[0082] According to an embodiment of the present disclosure, the implicit prompt information may be text obtained by converting a language instruction issued by a user, or text directly input by the user. The original text has a certain length. When processing the text, the large language model may have certain limitations on the length of the text.

[0083] Based on this, in some embodiments, as Figure 5 shown, the RGB image 511a corresponding to the image to be reconstructed can be input into the visual encoder 514 for visual encoding, and after projecting the visual encoding result, the first encoded image 511c can be obtained. The first encoded image 511c is input into the second large language model 515.

[0084] The RGB image 511a corresponding to the image to be reconstructed can be input into the image encoder 512 for image encoding, and the second encoded image 511b can be obtained. The second encoded image 511b is input into the mask decoder 513.

[0085] Before inputting the implicit prompt information 516 into the second large language model 515, the implicit prompt information 516 can be first input into the tokenizer 518 for semantic segmentation to obtain a plurality of semantic tokens 519, and then the semantic tokens 519 are input into the second large language model 515. The second large language model 515 performs reasoning based on the input first encoded image 511c and the semantic tokens 519 to obtain a mask 517 of at least one candidate object indicated by the output implicit prompt information, and the mask 519 of at least one candidate object is input into the mask decoder 513. The mask decoder 513 fuses the input second encoded image 511b and the mask 517 of at least one candidate object, and outputs the segmented image 511d.

[0086] In other embodiments, decoding the first encoded image based on the mask to generate a segmented image includes:

[0087] Convert the first encoded image into an image embedding, and the image embedding represents the features included in the first encoded image.

[0088] Fuse the image embedding and the mask based on a specific operation to obtain a fused feature, where the fused feature includes the features in the first encoded image affected by the mask.

[0089] Generate a segmentation image based on the fused feature.

[0090] As Figure 5 shown, before inputting the second encoded image 511b into the mask decoder 513, the second encoded image 511b can be pre-converted into an image embedding 511e (Image Embeddings). Image embedding is the process of converting image data into a dense vector representation. These vectors are usually represented in a low-dimensional space but still retain the key features and information of the image. Image embedding can be achieved through deep learning models such as convolutional neural networks (CNNs), which can learn the features in the image and map these features into a vector space. The model can extract local and global features in the image and combine them into a vector representation.

[0091] Through the three-dimensional image reconstruction method of the embodiments of the present disclosure, splitting the long text into multiple semantic token sequences and then fusing the input with the visual coding results can ensure that each sequence is within the processing capacity of the model, thus avoiding processing failures or efficiency degradation caused by overly long text. At the same time, the semantically tokenized text is more easily processed in parallel by the model, further improving the processing efficiency. In addition, splitting the long text into multiple semantic token sequences helps the large language model to more accurately understand the vocabulary, phrases, and sentence structures in the text, enabling it to generate more precise and coherent responses, thereby improving the accuracy of three-dimensional image reconstruction. Image embedding helps to achieve cross-modal information retrieval and understanding. After converting the image coding results into image embeddings, it can better achieve the fusion between image coding and the mask, realizing accurate image segmentation.

[0092] Based on the above embodiments, in some embodiments, geometric reconstruction of the image to be reconstructed to obtain a three-dimensional mesh model may include:

[0093] Construct a voxel block grid according to the size and resolution of the scene where the image to be reconstructed is located. Each voxel in the voxel block grid is used to store the distance information from the position where the voxel is located to the object surface and the corresponding weight.

[0094] Project the depth image corresponding to the image to be reconstructed onto the voxel block grid, determine the voxel positions corresponding to each projection point, and calculate the distance information from the position where the voxel corresponding to each projection point is located to the object surface according to the voxel positions corresponding to each projection point.

[0095] Update the distance information between the position of the voxel corresponding to the projection point and the object surface and the corresponding weights stored in the voxel in the voxel block grid to obtain an updated voxel block grid.

[0096] Extract the three-dimensional grid model from the updated voxel block grid.

[0097] According to an embodiment of the present disclosure, a Truncated Signed Distance Function (TSDF) can be used in combination with a Fragment Bounding Volume (FBV) for geometric reconstruction. In the context of three-dimensional image reconstruction, the FBV can be used to define and limit the spatial extent of each local fragment (or segment). These fragments are extracted from images or video frames captured from different perspectives, and together they form the final three-dimensional model. By calculating the cubic bounding volume (i.e., FBV) of each key-frame frustum, only the regions within these volumes need to be considered during the reconstruction process.

[0098] According to an embodiment of the present disclosure, the distance information between the position of the voxel and the object surface can be a TSDF value. The truncation range of the TSDF value can be determined by setting a truncation distance λ, that is, voxels within λ from the object surface can be considered for reconstruction. The size and number of voxels determine the accuracy and computational cost of the reconstruction. The smaller the voxel, the higher the modeling accuracy, but the greater the computational cost.

[0099] When performing three-dimensional reconstruction on a video, the image to be reconstructed can be a sequence of consecutive video frames with time series. To provide sufficient motion parallax and maintain multi-view co-visibility for language-guided reconstruction, if the relative translation of a certain frame is greater than tmax and the relative rotation angle is greater than Rmax, then this frame is selected as a key frame. A window with N key frames is defined as a local fragment, and the global reconstruction result is obtained by fusing all local fragments. The maximum visible depth of each view is set to dmax, and all view frustums in the fragment are restricted to a cubic shape and voxelized FBV. For the FBV incoming at time t and the depth maps from N views, depth map fusion is performed using standard TSDF integration with a truncation distance of λ. To perform global TSDF fusion, only the global TSDF values within the current FBV are updated. For the previous global TSDF fusion results, the parts within the FBV participate in the fusion process. We gradually recover the scene geometry, only retaining TSDF values less than λ. The Marching Cubes algorithm can be executed to reconstruct the three-dimensional grid model. Based on the geometric reconstruction result, the target object that conforms to the implicit hint information can be further subdivided.

[0100] Through the 3D image reconstruction method of the embodiments of the present disclosure, geometric reconstruction of the image to be reconstructed can be realized in parallel based on the voxel block grid combined with the truncated signed distance function, which can optimize the calculation efficiency and reduce redundancy, with high geometric reconstruction efficiency and accuracy, thus providing the efficiency and accuracy of 3D image reconstruction.

[0101] Based on the above embodiments, in some embodiments, fusing the segmented image and the 3D mesh model to obtain the 3D reconstruction result of at least one candidate object may include:

[0102] Adding a mask projection of at least one candidate object in the segmented image to the 3D mesh model, and calculating the overlap degree between any two masks in the 3D space.

[0103] In response to the overlap degree being greater than the overlap degree threshold, fusing any two masks into one mask to obtain the 3D reconstruction result of at least one candidate object.

[0104] According to the embodiments of the present disclosure, the 2D segmentation mask can be projected onto the 3D mesh model according to the internal and external parameters of the camera. The internal parameters of the camera describe the internal optical characteristics of the camera, such as the focal length and the optical center, and the external parameters describe the position and orientation of the camera in the world coordinate system. The intersection over union (IoU) between any two masks in the 3D space can be calculated as the overlap degree.

[0105] For example, setting the overlap degree threshold to 0.5, when the overlap degree between two masks is greater than 0.5, it can be considered that these two masks correspond to the same object in the 3D space, and these two masks can be fused as the mask of the object. The fusion method can be, for example, weighted fusion, voting, or more complex geometric fusion. During the fusion process, attention needs to be paid to processing the boundary regions to avoid cracks or inconsistencies in the overlapping parts.

[0106] Through the 3D image reconstruction method of the embodiments of the present disclosure, fusing the masks of candidate objects based on the overlap map between any two masks can integrate information from multiple perspectives, thus more accurately reconstructing the 3D shape of the object and improving the accuracy of 3D image reconstruction. Moreover, it can reduce redundant data, significantly reduce the computational complexity, and improve the efficiency of 3D image reconstruction.

[0107] Based on the above embodiments, in some embodiments, inputting the implicit prompt information and the 3D reconstruction result of at least one candidate object into the first large language model may include:

[0108] Identifying and globally reasoning about the keywords, phrases, and semantic relationships included in the implicit prompt information to obtain the semantic information associated with at least one candidate object, and the semantic information can be used to describe the attributes of at least one candidate object.

[0109] Match the attributes of at least one candidate object described by semantic information with the geometric features of at least one candidate object in the 3D reconstruction result, and adjust the 3D reconstruction result according to the matching result to obtain the 3D reconstruction result of the target object.

[0110] According to an embodiment of the present disclosure, since the field of view of the model is limited during the 2D inference segmentation process and the objects within the field of view of the current frame are seen, some candidate objects that conform to the implicit hint information can be obtained. Which specific candidate object best conforms to the implicit hint information still needs to be comprehensively judged after obtaining the global information. A large language model can be used for global inference, that is, the implicit hint information and the 3D reconstruction result are input into the large language model for global inference to obtain the final target instance and its 3D image reconstruction result.

[0111] According to an embodiment of the present disclosure, the training data set of the first large language model may include a third training data set, and the third training data set includes data pairs composed of a third mask and third training hint information, and the third training hint information implicitly includes the category represented by the first mask.

[0112] According to an embodiment of the present disclosure, the first large language model can be trained by constructing specific data pairs so that the first large language model has the ability to recognize implicit hint information, that is, it can recognize the implicit category information contained in the implicit hint information. In this way, the hint information does not need to include explicit category information, but puts forward a requirement, and the specific target to be segmented is inferred by the first large language model.

[0113] Through the 3D reconstruction method of the embodiment of the present disclosure, based on the large language model, the 3D reconstruction result of the candidate object is adjusted according to the implicit hint information, which can better realize the global inference of the 3D reconstruction result of the candidate object and further improve the accuracy of 3D image reconstruction.

[0114] Based on the above embodiments, in some embodiments, before inputting the 3D reconstruction result into the first large language model, it may further include:

[0115] Convert the 3D reconstruction result into structured information.

[0116] Input the structured information and the implicit hint information into the first large language model, and output the 3D reconstruction result of the target object indicated by the implicit hint information.

[0117] According to an embodiment of the present disclosure, structured information can present key data, attributes, and relationships in the 3D reconstruction result in a standardized format, and describes each element and its mutual relationship in the 3D reconstruction result through explicit semantic relationships. This helps the large model to more accurately understand the role and position of these elements in the overall structure, enabling the large model to more easily understand and process this information.

[0118] Through the 3D image reconstruction method according to the embodiment of the present disclosure, by organizing the 3D reconstruction result into structured information and inputting it into the large model, the data processing efficiency can be significantly improved, the model understanding ability can be enhanced, thereby improving the efficiency and accuracy of 3D image reconstruction.

[0119] Based on the above embodiment, in some embodiments, the 3D image reconstruction method may further include:

[0120] Obtain the next image to be reconstructed.

[0121] Perform incremental reconstruction on the 3D reconstruction result of the target object based on the next image to be reconstructed, and obtain the 3D reconstruction result of the target object in the next image to be reconstructed.

[0122] According to an embodiment of the present disclosure, incremental reconstruction can gradually add the next image, point cloud, or other data sources on the basis of the existing 3D image reconstruction result to update and improve the 3D model. This method is particularly suitable for scenarios that require processing a large amount of data or real-time updating of the 3D model. For example, in scenarios that require real-time updating of the 3D model, such as in the fields of augmented reality (AR) and virtual reality (VR), incremental reconstruction can be adopted. Another example is that in large-scale 3D reconstruction, incremental reconstruction can be used to gradually integrate data and reduce the computational complexity. Still another example is that in dynamically changing scenarios, such as in the fields of autonomous driving and robot navigation, incremental reconstruction can update the 3D model in real time to adapt to the scene changes.

[0123] According to the 3D image reconstruction method of the embodiment of the present disclosure, the incremental reconstruction method can better realize real-time 3D scenes, large-scale 3D reconstruction scenes, and dynamic 3D reconstruction scenes.

[0124] The following will combine Figure 6 A model training method of an embodiment of the present disclosure will be described schematically. Figure 6 It is a flowchart of model training according to an embodiment of the present disclosure.

[0125] As Figure 6 shown, the model training method 600 may include operation S610 to operation S620.

[0126] In operation S610, obtain a first training data set, and the first training data set includes a data pair composed of a first mask and first training prompt information.

[0127] In operation S620, the model is trained based on the first training dataset to obtain a trained inference segmentation model.

[0128] According to an embodiment of the present disclosure, the first training prompt information implicitly includes the category of the first mask representation.

[0129] According to an embodiment of the present disclosure, the inference segmentation model can be trained by constructing specific data pairs, so that the inference segmentation model has the ability to recognize implicit prompt information, that is, it can recognize the implicit category information included in the implicit prompt information. In this way, the prompt information does not need to include explicit category information, but puts forward a requirement, and the specific target to be segmented is inferred by the inference segmentation model.

[0130] For example, the process of constructing specific data can be as follows: for an image, the target categories included therein are known. The known target categories can be input into a large model, and questions can be asked about the targets represented by these target categories. The questions can be asked from the aspects of functions and uses, and the name of the known target category cannot be explicitly included in the questions. In this way, implicit prompt information-mask data pairs are constructed through repeated questioning, and used as the training data pairs of the inference segmentation model to train the inference segmentation model. Since the quality of the mask output by the large model through questioning may not be high, for example, a segmentation model can be used to refine the mask to obtain a high-quality mask.

[0131] Based on the above embodiments, in some embodiments, the model training method may further include:

[0132] Obtain a second training dataset, where the second training dataset includes data pairs composed of a second mask and second training prompt information, and the second training prompt information implicitly includes the category of the second mask representation.

[0133] Train the model based on the first training dataset and the second training dataset to obtain a trained inference segmentation model.

[0134] According to an embodiment of the present disclosure, there is a one-to-one or one-to-many relationship between the second training prompt information and the second mask. Specific data can be constructed as the training data of the model to train the model, so that when an implicit prompt information is input, the model can output one mask, multiple masks, or zero masks. The construction of specific data is similar to the above, and will not be elaborated here.

[0135] It should be noted that the trained inference segmentation model can be used to perform inference segmentation on an image based on implicit prompt information to obtain a segmented image. The segmented image can be used in scenarios such as image recognition and 3D image reconstruction. That is, the specific application scenarios of the trained inference segmentation model can be determined according to actual application requirements, and the present disclosure does not make any restrictions.

[0136] When the trained inference segmentation model is used to perform inference segmentation on an image based on implicit prompt information and the segmented image is used for 3D image reconstruction, the trained inference segmentation model can be used as the second large language model in the aforementioned 3D image reconstruction method to perform inference segmentation on the image to be reconstructed. For specific details, please refer to the method embodiment section and will not be elaborated here.

[0137] The following will be combined with Figure 7 to schematically describe a 3D image reconstruction device based on a large model according to an embodiment of the present disclosure. Figure 7 It is a block diagram of a 3D image reconstruction device based on a large model according to an embodiment of the present disclosure.

[0138] As Figure 7 shown, the 3D image reconstruction device 700 based on a large model may include an inference segmentation module 710, a geometric reconstruction module 720, a fusion module 730, and an input / output module 740.

[0139] The inference segmentation module 710 is configured to perform inference segmentation on the image to be reconstructed based on implicit prompt information, determine at least one candidate object in the image to be reconstructed, and add a mask to the at least one candidate object to obtain a segmented image.

[0140] The geometric reconstruction module 720 is configured to perform geometric reconstruction on the image to be reconstructed to obtain a 3D mesh model, and the 3D mesh model represents the geometric information of at least one candidate object and the geometric information of the scene where the image to be reconstructed is located.

[0141] The fusion module 730 is configured to fuse the segmented image and the 3D mesh model to obtain a 3D reconstruction result of at least one candidate object.

[0142] The input / output module 740 is configured to input the implicit prompt information and the 3D reconstruction result of at least one candidate object into the first large language model, and output the 3D reconstruction result of the target object indicated by the implicit prompt information.

[0143] According to an embodiment of the present disclosure, the inference segmentation module 710 performs inference segmentation on the image to be reconstructed based on implicit prompt information, determines at least one candidate object in the image to be reconstructed, and adds a mask to the at least one candidate object to obtain a segmented image, including:

[0144] Perform image encoding on the image to be reconstructed to obtain a first encoded image.

[0145] Perform visual encoding on the image to be reconstructed to obtain a second encoded image.

[0146] Input the second encoded image and the implicit prompt information into the second large language model, and output masks of at least one candidate object indicated by the implicit prompt information.

[0147] Decode the first encoded image based on the masks to generate a segmentation image.

[0148] According to an embodiment of the present disclosure, the three-dimensional image reconstruction device 700 further includes:

[0149] A determination module, configured to determine a second quantity of the masks according to a first quantity of objects included in the image to be reconstructed and a scene corresponding to the image to be reconstructed.

[0150] The input / output module 740 is further configured to input the second encoded image and the implicit prompt information into the second large language model, so that the second large language model outputs masks of the second quantity.

[0151] According to an embodiment of the present disclosure, the training data set of the second large language model includes a first training data set, and the first training data set includes data pairs composed of first masks and first training prompt information, and the first training prompt information implicitly includes categories represented by the first masks.

[0152] According to an embodiment of the present disclosure, the training data set of the second large language model includes a second training data set, and the second training data set includes data pairs composed of second masks and second training prompt information, and the second training prompt information implicitly includes categories represented by the second masks, and there is a one-to-one relationship or a one-to-many relationship between the second training prompt information and the second masks.

[0153] According to an embodiment of the present disclosure, the three-dimensional image reconstruction device 700 further includes:

[0154] A segmentation module, configured to segment the text included in the implicit prompt information and convert the text into at least one semantic token.

[0155] The input / output module 740 is further configured to input at least one semantic token into the second large language model, so that the second large language model outputs masks of at least one candidate object indicated by the implicit prompt information according to the second encoded image and at least one semantic token.

[0156] According to an embodiment of the present disclosure, the geometric reconstruction module 720 performs geometric reconstruction on the image to be reconstructed to obtain a three-dimensional mesh model, which may include:

[0157] Construct a voxel block grid according to the size and resolution of the scene where the image to be reconstructed is located. Each voxel in the voxel block grid is used to store the distance information from the position of the voxel to the object surface and the corresponding weight.

[0158] Project the depth image corresponding to the image to be reconstructed into the voxel block grid, determine the voxel positions corresponding to each projection point, and calculate the distance information from the position of the voxel corresponding to each projection point to the object surface.

[0159] Update the distance information from the position of the voxel stored in the voxel block grid to the object surface and the corresponding weight according to the distance information from the position of the voxel corresponding to the projection point to the object surface, and obtain the updated voxel block grid.

[0160] Extract a three-dimensional grid model from the updated voxel block grid.

[0161] According to an embodiment of the present disclosure, the fusion module 730 is used to fuse the segmentation image and the three-dimensional grid model to obtain the three-dimensional reconstruction result of at least one candidate object, which may include:

[0162] Add a mask projection of at least one candidate object in the segmentation image into the three-dimensional grid model, and calculate the overlap degree between any two masks in the three-dimensional space.

[0163] In response to the overlap degree being greater than the overlap degree threshold, fuse any two masks into one mask to obtain the three-dimensional reconstruction result of at least one candidate object.

[0164] According to an embodiment of the present disclosure, the input / output module 740 is used to input the implicit prompt information and the three-dimensional reconstruction result of at least one candidate object into the first large language model, which may include:

[0165] Identify and globally reason about the keywords, phrases, and semantic relationships included in the implicit prompt information to obtain semantic information associated with at least one candidate object, and the semantic information is used to describe the attributes of at least one candidate object.

[0166] Match the attributes of at least one candidate object described by the semantic information with the geometric features of at least one candidate object in the three-dimensional reconstruction result, and adjust the three-dimensional reconstruction result according to the matching result to obtain the three-dimensional reconstruction result of the target object.

[0167] According to an embodiment of the present disclosure, the three-dimensional image reconstruction device 700 may further include:

[0168] A conversion module for converting the three-dimensional reconstruction result into structured information.

[0169] The input / output module 740 is further configured to input the structured information and the implicit prompt information into the first large language model, and output the three-dimensional reconstruction result of the target object indicated by the implicit prompt information.

[0170] According to an embodiment of the present disclosure, the three-dimensional image reconstruction apparatus 700 may further include:

[0171] An incremental reconstruction module, configured to perform incremental reconstruction on the three-dimensional reconstruction result of the target object based on the next image to be reconstructed, to obtain the three-dimensional reconstruction result of the target object in the next image to be reconstructed.

[0172] It should be noted that the details of other embodiments of the three-dimensional image reconstruction apparatus and the technical effects brought are the same as or similar to the details of the embodiments of the three-dimensional image reconstruction method, and will not be elaborated here.

[0173] The following will be combined with Figure 8 to schematically describe a model training apparatus according to an embodiment of the present disclosure. Figure 8 It is a block diagram of a model training apparatus based on a large model according to an embodiment of the present disclosure.

[0174] As Figure 8 shown, the model training apparatus 800 may include an acquisition module 810 and a training module 820.

[0175] The acquisition module 810 is configured to acquire a first training data set, where the first training data set includes a data pair composed of a first mask and first training prompt information, and the first training prompt information implicitly includes the category represented by the first mask.

[0176] The training module 820 is configured to train a model based on the first training data set to obtain a trained inference segmentation model, where the trained inference segmentation model is used to perform inference segmentation on the image to be reconstructed based on the implicit prompt information in the three-dimensional image reconstruction method to obtain a segmented image.

[0177] It should be noted that the details of other embodiments of the model training apparatus and the technical effects brought are the same as or similar to the details of the embodiments of the model training method, and will not be elaborated here.

[0178] The following will be combined with Figure 9 to schematically describe an electronic device according to an embodiment of the present disclosure. Figure 9 It is a block diagram of an electronic device according to an embodiment of the present disclosure.

[0179] As Figure 9 shown, the electronic device 900 may include

[0180] At least one display device 910, and a large model-based three-dimensional image reconstruction device 920 communicatively connected to the at least one display device 710.

[0181] The large model-based three-dimensional image reconstruction device 920 is configured to execute the three-dimensional image reconstruction method described in the foregoing embodiments, obtain a three-dimensional reconstruction result, and output the three-dimensional reconstruction result to the at least one display device for display.

[0182] It should be noted that the details of other embodiments of the electronic device and the technical effects brought are the same as or similar to the details of the embodiments of the three-dimensional reconstruction method, and will not be elaborated here.

[0183] In the technical solution of the present disclosure, the processing of collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0184] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0185] Figure 10 A schematic block diagram of an exemplary electronic device 1000 that can be used to implement the embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0186] As Figure 10 shown, the device 1000 includes a computing unit 1001, which can execute various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 1002 or the computer program loaded from the storage unit 1008 into the random access memory (RAM) 1003. In the RAM 1003, various programs and data required for the operation of the device 1000 can also be stored. The computing unit 1001, the ROM 1002, and the RAM 1003 are connected to each other through a bus 1004. The input / output (I / O) interface 1005 is also connected to the bus 1004.

[0187] Multiple components in device 1000 are connected to I / O interface 1005, including: an input unit 1006, such as a keyboard, a mouse, etc.; an output unit 1007, such as various types of displays, speakers, etc.; a storage unit 1008, such as a disk, an optical disc, etc.; and a communication unit 1009, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1009 allows device 1000 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0188] The computing unit 1001 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1001 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1001 executes the various methods and processes described above, such as a text processing method and / or a deployment method of a deep learning framework. For example, in some embodiments, the text processing method and / or the deployment method of the deep learning framework can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 1008. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 1000 via the ROM 1002 and / or the communication unit 1009. When the computer program is loaded into the RAM 1003 and executed by the computing unit 1001, one or more steps of the text processing method and / or the deployment method of the deep learning framework described above can be executed. Alternatively, in other embodiments, the computing unit 1001 can be configured to execute the text processing method and / or the deployment method of the deep learning framework in any other suitable manner (e.g., by means of firmware).

[0189] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard parts (ASSPs), system on chip (SOC) systems, complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.

[0190] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code can execute entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine or entirely on the remote machine or server.

[0191] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory (EPROM) or flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0192] To provide for interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a cathode ray tube (CRT) monitor or a liquid crystal display (LCD)) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0193] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0194] A computer system may include a client and a server. The client and the server are generally far apart from each other and typically interact via a communication network. The relationship between the client and the server is generated by computer programs that run on respective computers and have a client-server relationship with each other.

[0195] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitations are imposed herein.

[0196] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.

Claims

1. A three-dimensional image reconstruction method based on a large model, comprising: Performing reasoning segmentation on the image to be reconstructed based on implicit prompt information, determining at least one candidate object in the image to be reconstructed and adding a mask to the at least one candidate object to obtain a segmented image; Performing geometric reconstruction on the image to be reconstructed to obtain a three-dimensional mesh model, wherein the three-dimensional mesh model represents geometric information of the at least one candidate object and geometric information of a reconstruction scene of the image to be reconstructed; fusing the segmented image and the three-dimensional mesh model to obtain a three-dimensional reconstruction result of the at least one candidate object; as well as The implicit prompt information and the three-dimensional reconstruction result of the at least one candidate object are input into a first large language model, and the three-dimensional reconstruction result of the target object indicated by the implicit prompt information is output.

2. The method according to claim 1, wherein: The method of performing reasoning segmentation on the image to be reconstructed based on implicit prompt information, determining at least one candidate object in the image to be reconstructed and adding a mask to the at least one candidate object to obtain a segmented image includes: Inputting a first encoded image obtained by visually encoding the image to be reconstructed and the implicit prompt information into a second large language model, and outputting a mask of at least one candidate object indicated by the implicit prompt information; and The second encoded image obtained by performing image encoding on the image to be reconstructed is decoded based on the mask to generate the segmented image.

3. The method according to claim 2, further comprising: Determining a second number of the masks according to a first number of objects included in the image to be reconstructed and a scene corresponding to the image to be reconstructed; Inputting the first encoded image and the implicit hint information into the second large language model so that the second large language model outputs the second number of masks.

4. The method according to claim 2 or 3, wherein: The training data set of the second largest language model includes a first training data set, wherein the first training data set includes a data pair consisting of a first mask and first training prompt information, and the first training prompt information implicitly includes a category represented by the first mask.

5. The method according to claim 3 or 4, wherein: The training data set of the second largest language model includes a second training data set, the second training data set includes a data pair consisting of a second mask and second training prompt information, the second training prompt information implicitly includes the category represented by the second mask, and the second training prompt information and the second mask have a one-to-one relationship or a one-to-many relationship.

6. The method according to claim 2, further comprising: Segmenting the text contained in the implicit prompt information, and converting the text into at least one semantic tag; as well as The at least one semantic tag is input into a second large language model, so that the second large language model outputs a mask of at least one candidate object indicated by the implicit prompt information according to the second encoded image and the at least one semantic tag.

7. The method according to claim 2, wherein: The step of decoding the second encoded image obtained by encoding the image to be reconstructed based on the mask to generate the segmented image comprises: Converting the second encoded image into an image embedding, wherein the image embedding represents features contained in the second encoded image; fusing the image embedding with the mask to obtain a fused feature, the fused feature comprising features of the second encoded image affected by the mask; and The segmented image is generated based on the fused features.

8. The method according to claim 1, wherein: The step of geometrically reconstructing the image to be reconstructed to obtain a three-dimensional mesh model includes: Constructing a voxel block grid according to the physical shape and resolution of the scene in which the image to be reconstructed is located, wherein each voxel in the voxel block grid is used to store distance information between the position of the voxel and the surface of the object and a corresponding weight; Projecting the depth image corresponding to the image to be reconstructed onto the voxel block grid, and determining the distance information between the position of the voxel corresponding to each projection point and the surface of the object according to the voxel position corresponding to each projection point; updating the distance information and corresponding weights stored in the voxels in the voxel block grid according to the distance information to obtain an updated voxel block grid; and The three-dimensional mesh model is extracted from the updated voxel block grid.

9. The method according to claim 8, wherein: The fusing the segmented image and the three-dimensional mesh model to obtain a three-dimensional reconstruction result of the at least one candidate object includes: Adding a mask to the at least one candidate object in the segmented image and projecting it into the three-dimensional grid model, and calculating the overlap between any two masks in the three-dimensional space; and When the overlap degree is greater than an overlap degree threshold, the arbitrary two masks are fused into one mask to obtain a three-dimensional reconstruction result of the at least one candidate object.

10. The method according to claim 1, wherein: Inputting the implicit prompt information and the three-dimensional reconstruction result of the at least one candidate object into a first large language model, comprising: Identifying and globally reasoning the keywords, phrases, and semantic relationships contained in the implicit prompt information to obtain semantic information associated with the at least one candidate object, wherein the semantic information is used to describe the attributes of the at least one candidate object; and The attributes of at least one candidate object described by the semantic information are matched with the geometric features of at least one candidate object in the three-dimensional reconstruction result, and the three-dimensional reconstruction result is adjusted according to the matching result to obtain the three-dimensional reconstruction result of the target object.

11. The method according to claim 1 or 10, further comprising: Converting the three-dimensional reconstruction result into structural information; as well as The structured information and the implicit prompt information are input into the first large language model, and a three-dimensional reconstruction result of the target object indicated by the implicit prompt information is output.

12. The method according to claim 1, further comprising: Acquire the next image to be reconstructed; as well as The three-dimensional reconstruction result of the target object is incrementally reconstructed based on the next image to be reconstructed to obtain the three-dimensional reconstruction result of the target object in the next image to be reconstructed.

13. A model training method, comprising: Acquire a first training data set, where the first training data set includes a data pair consisting of a first mask and first training prompt information, where the first training prompt information implicitly includes a category represented by the first mask; as well as The model is trained based on the first training data set to obtain a trained inference segmentation model, wherein the trained inference segmentation model is used to perform inference segmentation on the image to be reconstructed based on implicit prompt information to obtain a segmented image.

14. The method according to claim 13, further comprising: Obtaining a second training data set, the second training data set comprising a data pair consisting of a second mask and second training prompt information, the second training prompt information implicitly including a category represented by the second mask, and a one-to-one relationship or a one-to-many relationship between the second training prompt information and the second mask; as well as The model is trained based on the first training data set and the second training data set to obtain a trained inference segmentation model.

15. A three-dimensional image reconstruction device based on a large model, comprising: An inference segmentation module, configured to perform inference segmentation on the image to be reconstructed based on implicit prompt information, determine at least one candidate object in the image to be reconstructed and add a mask to the at least one candidate object to obtain a segmented image; A geometric reconstruction module, configured to geometrically reconstruct the image to be reconstructed to obtain a three-dimensional mesh model, wherein the three-dimensional mesh model represents geometric information of the at least one candidate object and geometric information of a scene in which the image to be reconstructed is located; a fusion module, configured to fuse the segmented image and the three-dimensional mesh model to obtain a three-dimensional reconstruction result of the at least one candidate object; as well as The input-output module is used to input the implicit prompt information and the three-dimensional reconstruction result of the at least one candidate object into the first large language model, and output the three-dimensional reconstruction result of the target object indicated by the implicit prompt information.

16. A model training device, comprising: An acquisition module, configured to acquire a first training data set, wherein the first training data set includes a data pair consisting of a first mask and first training prompt information, wherein the first training prompt information implicitly includes a category represented by the first mask; as well as A training module is used to train the model based on the first training data set to obtain a trained inference segmentation model, wherein the trained inference segmentation model is used to perform inference segmentation on the image to be reconstructed based on implicit prompt information to obtain a segmented image.

17. An electronic device comprising at least one display device; and The large model-based three-dimensional image reconstruction device as claimed in claim 15, being communicatively connected to the at least one display device; The large model-based three-dimensional image reconstruction device is used to execute the method described in any one of claims 1 to 12 to obtain a three-dimensional reconstruction result, and output the three-dimensional reconstruction result to the at least one display device for display.

18. An electronic device, comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 14.

19. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 14.

20. A computer program product comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1 to 14.

Citation Information

Cited By

  • Gaussian splash dynamic three-dimensional reconstruction method and system based on spiking neurons

    CN121010707A

  • A method and system for dynamic three-dimensional reconstruction of Gaussian splashes based on spiking neurons

    CN121010707B

  • Three-dimensional point cloud data segmentation model training and application method, equipment and medium

    CN121482066A