Image processing method, electronic equipment and computer readable storage medium
Through multiple depth estimation and camera parameters, the problem of high cost of obtaining three-dimensional scene data sets and insufficient diversity is solved, low-cost, large-scale three-dimensional scene data generation is achieved, and the development of spatial intelligence technology is promoted.
Patent Information
- Application Number
- CN202510460071.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-07-18
AI Technical Summary
In the prior art, the acquisition of three-dimensional scene data sets is expensive, the scene diversity is insufficient and the scale is small, making it difficult to support the training of complex spatial intelligent models.
By acquiring two-dimensional images, multiple depth estimation strategies are used to perform scene depth estimation, three-dimensional scene data is generated in combination with camera parameters, and large-scale training is used to achieve low-cost and large-scale three-dimensional scene data generation.
The generated three-dimensional scene data is highly authentic and has clear details, which significantly reduces the cost of data acquisition and improves the development of spatial intelligence technology.
Smart Images

Figure CN120339515A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical fields of image processing, artificial intelligence, and large model technology. Specifically, it relates to an image processing method, an electronic device, and a computer-readable storage medium. Background Art
[0002] Spatial intelligence is an important development direction of artificial intelligence. However, the development of spatial intelligence is limited by the lack of large-scale, high-quality, and well-annotated three-dimensional (3D) scene datasets. Compared with the vast amount of easily obtainable two-dimensional (2D) image data, the acquisition cost of 3D scene data is high. Usually, special sensors such as lidar and depth cameras are required, and the scale of 3D scene datasets is generally small and the scene diversity is insufficient, making it difficult to support the training of complex spatial intelligence models.
[0003] For the above problems, no effective solutions have been proposed yet. Summary of the Invention
[0004] Embodiments of this application provide an image processing method, an electronic device, and a computer-readable storage medium to at least solve the technical problems of high acquisition cost, insufficient scene diversity, and small scale of three-dimensional scene datasets in related technologies.
[0005] According to one aspect of the embodiments of this application, an image processing method is provided, including: obtaining a two-dimensional image; performing scene depth estimation on the two-dimensional image by using a multiple depth estimation strategy to obtain a target depth map, where the multiple depth estimation strategy is used to obtain scene depth information of a given image through different types of depth estimation methods; generating three-dimensional scene data based on the camera parameters of the two-dimensional image and the target depth map, where the three-dimensional scene data is training data applied to three-dimensional perception and understanding scenarios.
[0006] According to another aspect of the embodiments of this application, an image processing method is further provided, including: obtaining a two-dimensional image; performing scene depth estimation on the two-dimensional image by using a multiple depth estimation strategy to obtain a target depth map, where the multiple depth estimation strategy is used to obtain scene depth information of a given image through different types of depth estimation methods; generating three-dimensional scene data based on the camera parameters of the two-dimensional image and the target depth map; training an initial three-dimensional multimodal language model based on the three-dimensional scene data to obtain a target three-dimensional multimodal language model, where the target three-dimensional multimodal language model is used to process at least one of the following tasks: three-dimensional vision tasks, spatial reasoning tasks, cross-modal tasks.
[0007] According to another aspect of the embodiments of the present application, there is also provided an image processing method, including: obtaining an image processing request through a first application programming interface, where the request data carried in the image processing request includes: a two-dimensional image; returning an image processing response through a second application programming interface, where the response data carried in the image processing response includes: three-dimensional scene data, and the three-dimensional scene data is generated according to the image processing method of any one of the above.
[0008] According to another aspect of the embodiments of the present application, there is also provided an image processing method, including: obtaining a currently input image processing dialogue request, where the request data carried in the image processing dialogue request includes: a two-dimensional image; in response to the image processing dialogue request, returning an image processing dialogue reply, where the information carried in the image processing dialogue reply includes: three-dimensional scene data, and the three-dimensional scene data is generated according to the image processing method of any one of the above;
[0009] Display the three-dimensional scene data within the graphical user interface.
[0010] According to another aspect of the embodiments of the present application, there is also provided an image processing method, including: in response to an input instruction acting on an operation interface, displaying a two-dimensional image on the operation interface; in response to a processing instruction acting on the operation interface, displaying three-dimensional scene data on the operation interface; where the three-dimensional scene data is generated according to the image processing method of any one of the above.
[0011] According to another aspect of the embodiments of the present application, there is also provided an image processing system, including: a client for sending a two-dimensional image; a server connected to the client, for performing scene depth estimation on the two-dimensional image by using a multiple depth estimation strategy to obtain a target depth map, and generating three-dimensional scene data based on the camera parameters of the two-dimensional image and the target depth map, where the multiple depth estimation strategy is used to obtain the scene depth information of a given image through different types of depth estimation methods, and the three-dimensional scene data is training data applied to the three-dimensional perception and understanding scene; the client is further used to output the three-dimensional scene data.
[0012] According to another aspect of the embodiments of the present application, there is also provided an electronic device, including: a memory storing an executable program; a processor connected to the memory through a bus for running the program, where when the program runs, it executes the image processing method of any one of the above.
[0013] According to another aspect of the embodiments of the present application, there is also provided a computer-readable storage medium, the computer-readable storage medium including a stored executable program, where when the executable program runs, it controls the device where the computer-readable storage medium is located to execute the image processing method of any one of the above.
[0014] According to another aspect of the embodiments of the present application, there is also provided a computer program product, including a computer program which, when executed by a processor, implements any one of the above-mentioned image processing methods.
[0015] In the embodiments of the present application, by obtaining a two-dimensional image and adopting a multiple depth estimation strategy to perform scene depth estimation on the two-dimensional image to obtain a target depth map, and the multiple depth estimation strategy is used to obtain scene depth information of a given image through different types of depth estimation methods, and finally, three-dimensional scene data is generated based on the camera parameters of the two-dimensional image and the target depth map. Among them, the three-dimensional scene data is training data applied to the three-dimensional perception and understanding scene. Thus, the purpose of efficiently and low-costly generating three-dimensional scene data based on two-dimensional images is achieved, thereby realizing the low-cost acquisition of large-scale and high-quality three-dimensional scene data. Moreover, the generated three-dimensional scene data not only has the scale of the real world but also has clear details, significantly promoting the technical effect of the development of spatial intelligence technology, and further solving the technical problems of high acquisition cost, insufficient scene diversity and small scale in the three-dimensional scene data set in the related art.
[0016] It is easy to notice that the above general description and the following detailed description are only for exemplifying and explaining the present application, and do not constitute a limitation to the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:
[0018] Figure 1 is a schematic diagram of an application scenario of an image processing method according to an embodiment of the present application;
[0019] Figure 2 is a flowchart of an image processing method according to an embodiment of the present application;
[0020] Figure 3 is a schematic flowchart of generating 3D scene data based on 2D images according to an embodiment of the present application;
[0021] Figure 4 is a flowchart of an image processing method according to an embodiment of the present application;
[0022] Figure 5 is a flowchart of an image processing method according to an embodiment of the present application;
[0023] Figure 6 is a flowchart of an image processing method according to an embodiment of the present application;
[0024] Figure 7It is a flowchart of an image processing method according to an embodiment of the present application;
[0025] Figure 8 It is a schematic structural diagram of an image processing system according to an embodiment of the present application;
[0026] Figure 9 It is a schematic structural diagram of an image processing device according to an embodiment of the present application;
[0027] Figure 10 It is a schematic structural diagram of another image processing device according to an embodiment of the present application;
[0028] Figure 11 It is a schematic structural diagram of another image processing device according to an embodiment of the present application;
[0029] Figure 12 It is a schematic structural diagram of another image processing device according to an embodiment of the present application;
[0030] Figure 13 It is a schematic structural diagram of another image processing device according to an embodiment of the present application;
[0031] Figure 14 It is a block diagram of the structure of a computing device according to an embodiment of the present application;
[0032] Figure 15 It is a block diagram of the structure of an electronic device according to an embodiment of the present application. Detailed implementation manners
[0033] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0034] It should be noted that the terms "first", "second", etc. in the description, claims and the above-mentioned drawings of this application are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0035] The technical solution provided by this application is mainly implemented by large model technology. Here, the large model refers to a deep learning model with a large number of model parameters, usually including hundreds of millions, tens of billions, hundreds of billions, trillions or even more than one quadrillion model parameters. The large model can also be called a Foundation Model. Through large-scale pre-training on unlabeled corpora, a pre-trained model with more than hundreds of millions of parameters is produced. This model can adapt to a wide range of downstream tasks and has good generalization ability, such as Large Language Model (LLM), multi-modal pre-training model, etc.
[0036] It should be noted that in actual applications, the large model can be fine-tuned with a small number of samples on the pre-trained model, so that the large model can be applied to different tasks. For example, the large model can be widely applied in fields such as Natural Language Processing (NLP), computer vision, and speech processing. Specifically, it can be applied to tasks in the field of computer vision such as Visual Question Answering (VQA), Image Caption (IC), and image generation. It can also be widely applied to tasks in the field of natural language processing such as text-based sentiment classification, text summary generation, and machine translation. Therefore, the main application scenarios of the large model include but are not limited to digital assistants, intelligent robots, search, online education, office software, e-commerce, intelligent design, etc. In the embodiments of this application, taking the image processing by a large model that can implement the image processing method proposed in the embodiments of this application in the 3D scene data generation scenario as an example for explanation.
[0037] First, some nouns or terms that appear in the process of describing the embodiments of this application are applicable to the following explanations:
[0038] Spatial Intelligence: Refers to the perception, understanding, reasoning, and interaction capabilities of artificial intelligence in a three-dimensional environment, including the modeling and analysis of the spatial layout, scale of objects in a three-dimensional scene, and the interaction process with the environment.
[0039] 3D Computer Vision: Refers to the perception tasks carried out by computer vision in a three-dimensional environment, such as 3D object recognition, 3D semantic segmentation, 3D scene reconstruction, etc.
[0040] Multimodal Large Language Model (MLLM): Refers to a large language model that can simultaneously process various types of modal information (such as text, images, speech, video, 3D data, etc.), capable of cross-modal understanding and reasoning.
[0041] 3D Multimodal Large Language Model (3D-MLLM): A multimodal large language model specifically for three-dimensional scenes, with the ability to analyze 3D data and fuse and understand it with other modal information such as text and images.
[0042] Data Generation: Refers to the process of converting or enhancing existing data or raw sensor data into a new data form that meets specific application requirements through a certain algorithmic process. In the embodiments of this application, data generation mainly refers to converting 2D images into 3D point clouds, depth maps, and 3D annotations with real scale and extrinsic parameter information.
[0043] Red-Green-Blue-Depth (RGB-D) Data: A data format that combines Red-Green-Blue (RGB) images and depth information. The RGB image provides color information of the scene, while the D (depth) information represents the exact distance from each pixel point to the camera. This data format is very suitable for three-dimensional space understanding and perception tasks because the RGB image can provide rich visual details, and the depth information can help construct the three-dimensional structure of the scene to achieve functions such as object detection, tracking, and obstacle avoidance.
[0044] Spatial intelligence is an important development direction of artificial intelligence. However, its current development is limited by the lack of large-scale, high-quality, and well-annotated 3D scene datasets. Compared with the vast amount of easily available 2D image data, the acquisition cost of 3D scene data is high. Usually, special sensors such as lidar and depth cameras are required, and cumbersome post-annotation is needed, resulting in the generally small scale and insufficient scene diversity of existing 3D scene datasets, making it difficult to support the training of complex spatial intelligence models. At the same time, the rise of multimodal large models indicates that effectively using cross-modal big data (such as images, text, speech, etc.) can significantly improve the capabilities of artificial intelligence (AI). However, in the 3D field, due to data scarcity, the development of models such as 3D-MLLM lags far behind the 2D field.
[0045] Therefore, there is an urgent need for a method that can "elevate" relatively accurate 3D scene data that is metrically consistent with the real scene from a large amount of existing 2D image data at low cost and automatically.
[0046] Currently, there are the following various ways to generate and obtain 3D data.
[0047] Emulator approach: By using a robot simulation platform and a 3D simulation environment, etc., synthetic 3D data can be quickly generated under controllable conditions. However, the 3D data generated by the emulator usually has a simplified geometric structure and material texture, with a cartoonish style, and there is an obvious gap from the real world. That is, there is a "simulation-to-reality" gap. Models trained solely with simulation data often have insufficient generalization ability when facing the complex and changeable real environment.
[0048] AI-generated 3D assets approach: Some AI generation technologies can generate 3D models or representations such as Neural Radiance Fields (NeRF) from text or images through deep learning. However, this method is mostly limited to the generation of single objects and it is difficult to construct complex complete 3D scene data. Even when generating simple scenes, problems such as disproportion and unrealistic appearance often occur. For example, the generated texture and lighting are difficult to accurately restore natural conditions, and the layout of objects in the scene may not conform to real-world logic. In addition, many AI-generated 3D results also have a cartoon rendering style and are difficult to be directly used in applications with high requirements for realism.
[0049] Ways for sensors to collect 3D data: Using lidar and depth cameras to obtain point clouds or RGB-D data of the real environment is a traditional way to directly obtain high-fidelity 3D data. For example, the classic indoor scan dataset only contains about 1,500 indoor scenes. Although the data obtained by this method is real and accurate, the hardware equipment is expensive, the data collection is time-consuming and requires professional operation, and the subsequent manual annotation cost is also very high. Therefore, the scale of the 3D dataset collected by sensors is limited, and the scene types are biased towards specific fields (such as indoor rooms), making it difficult to cover a wide range of outdoor or complex environments.
[0050] It can be seen that the 3D scene datasets obtained in the related technologies have the following defects.
[0051] Defect 1: The authenticity of simulation generation is insufficient. The models and textures in the simulated environment are often too simplified, resulting in a visual domain difference between the generated data and the real world. If the model is trained entirely on the simulation data, it is prone to recognition and understanding errors in real-world applications, that is, the "simulation to reality" migration is difficult.
[0052] Defect 2: The integrity of the scenes generated by AI is insufficient. Existing AI methods for generating 3D can usually only generate single objects or simple scenes, and it is difficult to represent complex scenes with multiple object interactions. In addition, the generated results lack scale constraints, and problems such as unreasonable object size ratios and abnormal position layouts may occur, limiting their practicality.
[0053] Defect 3: The cost of sensor data is high and the coverage is limited. Although the data collected by relying on lidar or depth cameras has high quality, the acquisition cost and annotation cost are huge, making it difficult to expand the dataset on a large scale. The existing 3D data is mostly limited to a few indoor scenes or specific urban roads, lacking cross-scene diversity and making it difficult to meet the training needs of general spatial intelligence.
[0054] Defect 4: Data consistency and annotation problems. Although means such as data augmentation and style transfer can enrich the data to a certain extent, they often introduce problems such as inconsistent styles or distortion, and cannot guarantee the accuracy of the generated data in terms of geometric structure and scale. Moreover, these methods usually cannot automatically generate accurate 3D annotations and still require manual intervention.
[0055] In view of the above defects, no effective solution has been proposed before this application, that is, there is currently no technology that can simultaneously achieve high authenticity, low cost, large scale, and automatic annotation of 3D scene data generation.
[0056] According to an embodiment of the present application, an image processing method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0057] Considering that the number of model parameters of the large model is huge and the computing resources of the mobile terminal are limited, the above method provided by the embodiment of the present application can be applied to Figure 1 the application scenarios shown, but not limited thereto. In the Figure 1 application scenarios shown, the large model is deployed in the server 10. The server 10 can connect to one or more client devices 20 through a local area network connection, a wide area network connection, an Internet connection, or other types of data networks. Here, the client devices 20 can include but are not limited to: smart phones, tablets, laptops, palmtop computers, personal computers, smart home devices, in-vehicle devices, etc. The client device 20 can interact with the user through a graphical user interface to implement the call of the large model, and further implement the method provided by the embodiment of the present application.
[0058] In the embodiment of the present application, the system composed of the client device and the server can execute the following steps: The client device executes steps such as acquiring a two-dimensional image and sending the two-dimensional image to the server. The server executes performing scene depth estimation on the two-dimensional image by using a multiple depth estimation strategy to obtain a target depth map, where the multiple depth estimation strategy is used to obtain the scene depth information of a given image through different types of depth estimation methods, and generating three-dimensional scene data based on the camera parameters of the two-dimensional image and the target depth map. The three-dimensional scene data is training data applied to the three-dimensional perception and understanding scene, and returning the three-dimensional scene data to the client device, etc. It should be noted that in the case where the operating resources of the client device can meet the deployment and operation conditions of the large model, the embodiment of the present application can be performed in the client device.
[0059] It should be noted that with the rapid development of high-performance computing units, in other application scenarios, the above method provided by the embodiment of the present application can also be applied to a model all-in-one machine. In an optional embodiment, multiple models are built in the model all-in-one machine. The user can select and adjust one model according to needs to obtain his own model. Thus, the high-performance computing unit built in the model all-in-one machine can directly call the adjusted model to execute the above method provided by the embodiment of the present application. In another optional embodiment, a trained model is built in the large model all-in-one machine. Thus, the high-performance computing unit built in the model all-in-one machine can directly call the model to execute the above method provided by the embodiment of the present application.
[0060] Further, when the user needs to train their own model, they can also upload their own dataset through the client. This dataset is sent from the client to the server, enabling the server to adjust the pre-trained model with this dataset to obtain the user's own model, which can then be deployed to the production environment. To facilitate the user's model adjustment requirements, the server can provide complete adjustment tools, development frameworks, and processes, and support multiple adjustment strategies, enabling the adjusted model to better adapt to different field applications and achieve high customization.
[0061] Under the above operating environment, the present application provides an Figure 2 image processing method as shown below. Figure 2 It is a flowchart of an image processing method according to an embodiment of the present application. As Figure 2 shown, the method may include the following steps:
[0062] Step S21, obtain a two-dimensional image;
[0063] Step S22, perform scene depth estimation on the two-dimensional image using a multiple depth estimation strategy to obtain a target depth map, where the multiple depth estimation strategy is used to obtain the scene depth information of a given image through different types of depth estimation methods;
[0064] Step S23, generate three-dimensional scene data based on the camera parameters of the two-dimensional image and the target depth map, where the three-dimensional scene data is training data applied to three-dimensional perception and understanding scenarios.
[0065] In an embodiment of the present application, a two-dimensional image (2D Image) can be understood as an image with information in two dimensions of length and width, usually including RGB color channels. Exemplarily, the two-dimensional image can be an ordinary photographic photo or any other source of planar image, such as an image in a public image dataset, an image uploaded by a user, an image crawled from the Internet, a satellite or aerial image, etc., which is not limited here. The visual content in the two-dimensional image can include daily scenes, industrial and architectural scenes, extreme or rare scenes, professional scenes, etc., which is not limited here.
[0066] It can be understood that the obtained two-dimensional image can be a two-dimensional image with basic annotations, that is, additional information is attached to the two-dimensional image to indicate or describe specific objects, regions, or features in the two-dimensional image. Exemplarily, bounding boxes, segmentation masks, keypoints, semantic labels, attribute annotations, relationship annotations, etc. can be attached to the two-dimensional image, which is not limited here.
[0067] By obtaining two-dimensional images from different data sets or the Internet, a large amount of basic data can be provided for subsequent processing, thereby realizing the large-scale generation of three-dimensional scene data. Moreover, the acquisition of two-dimensional images in this application does not depend on specific acquisition devices or environments, reducing the cost of data acquisition and improving the flexibility and efficiency of data generation.
[0068] The multiple depth estimation strategy is used to obtain the scene depth information of a given image through different types of depth estimation methods, that is, the multiple depth estimation strategy can include multiple scene depth estimation methods. Taking two methods as an example, the multiple depth estimation strategy can include, for example, relative depth estimation and absolute depth estimation, which are used to obtain the relative distance information and the actual metric scale distance information of the objects in the image respectively, and are not limited here. In addition, the multiple depth estimation strategy can also include the Stereo Matching method, which estimates the depth by comparing image pairs (stereo image pairs) taken from different perspectives at the same moment, and can also include Structured Light Depth Estimation, which uses structured light technology, that is, projects light rays with known patterns (such as grids or stripes) onto the scene, and then calculates the depth by analyzing the deformation of these patterns in the scene, and is not limited here.
[0069] The target depth map can be understood as an image containing the depth information of each pixel point, reflecting the distance between each object in the scene and the observer. The obtained target depth map not only has the real-world scale, but also has very clear image details.
[0070] Using the multiple depth estimation strategy to perform scene depth estimation on the two-dimensional image to obtain the target depth map, the multiple depth estimation strategy adopted in this application can effectively make up for the deficiencies of a single method in terms of details and global scale, and improve the accuracy and stability of three-dimensional reconstruction. Moreover, using the multiple depth estimation strategy to perform scene depth estimation on the two-dimensional image, the reconstructed target depth map contains both fine local geometric features and ensures the overall metric scale, thereby enhancing the authenticity and reliability of the data.
[0071] The camera parameters of the two-dimensional image include internal parameters (such as focal length, image sensor size, etc.) and external parameters (such as the position and attitude of the camera, etc.), which are used to describe the characteristics of the camera and its geometric relationship relative to the scene.
[0072] The three-dimensional scene data is training data applied to the three-dimensional perception and understanding scene, and can be understood as a kind of data containing information such as the position, size, color, and texture of objects in the three-dimensional space, and can be widely applied to artificial intelligence products and services that require three-dimensional perception and understanding, such as being suitable for model training for three-dimensional perception, understanding, and interaction.
[0073] Generate 3D scene data based on the camera parameters of a 2D image and the target depth map. The use of camera parameters ensures the correct projection and scale consistency of the 3D scene data, facilitating subsequent model understanding and processing of 3D spatial relationships. Through the combined action of camera parameters and depth information, this application achieves precise positioning of each pixel point in the 3D coordinate system in the scene, thereby generating 3D scene data. The generated 3D scene data retains the visual features of the 2D image while assigning 3D spatial coordinates to each pixel point, greatly enhancing the practicality of the data, providing richer spatial information for the model, and helping to improve the performance and generalization ability of the model in 3D vision tasks.
[0074] In the embodiments of this application, a 2D image is obtained, and a multiple depth estimation strategy is used to perform scene depth estimation on the 2D image to obtain a target depth map. The multiple depth estimation strategy is used to obtain the scene depth information of a given image through different types of depth estimation methods. Finally, 3D scene data is generated based on the camera parameters of the 2D image and the target depth map, where the 3D scene data is training data applied to 3D perception and understanding scenarios. It can be seen that by obtaining existing large-scale 2D image datasets, this application avoids using expensive 3D acquisition devices for data acquisition, significantly reducing the hardware investment for obtaining 3D training data and the labor and material costs required for the acquisition process. Moreover, this application can process existing 2D images on a large scale and quickly generate a large amount of 3D scene data. In addition, this application uses multiple depth estimation methods to perform scene depth estimation on 2D images, extracts depth information from 2D images, and the 3D scene data generated according to camera parameters and depth information retains the texture and details of the original 2D image, is closer to the real world than simulated or synthetic data, effectively narrowing the visual domain gap between simulated data and actual applications, and at the same time achieving precise positioning of each pixel point in the 3D coordinate system. Therefore, it can be widely applied to artificial intelligence products and services that require 3D perception and understanding.
[0075] That is, this application overcomes the limitations of high cost, limited dataset scale, and insufficient realism in existing 3D scene data acquisition, provides a large amount of high-quality and low-cost 3D scene data for training spatial intelligence models, and significantly promotes the development of spatial intelligence technology.
[0076] It can be understood that the image processing method provided in the embodiments of this application can be applied to large models or other deep learning models, which is not limited here.
[0077] The above image processing method provided by the embodiments of the present application can be but is not limited to being applied to application scenarios involving the generation of three-dimensional scene data in fields such as e-commerce services, education services, legal services, medical services, conference services, social network services, financial product services, logistics services, and navigation services. For example, image processing related to e-commerce services, image processing related to education services, image processing related to legal services, etc., which are not limited here.
[0078] By adopting the embodiments of the present application, by obtaining a two-dimensional image and using a multiple depth estimation strategy to perform scene depth estimation on the two-dimensional image to obtain a target depth map, and the multiple depth estimation strategy is used to obtain the scene depth information of a given image through different types of depth estimation methods, and finally, based on the camera parameters of the two-dimensional image and the target depth map, three-dimensional scene data is generated, where the three-dimensional scene data is training data applied to three-dimensional perception and understanding scenarios. Thus, the purpose of efficiently and low-costly generating three-dimensional scene data based on two-dimensional images is achieved, thereby realizing the low-cost acquisition of large-scale and high-quality three-dimensional scene data, and the generated three-dimensional scene data has both real-world scale and clear details, significantly promoting the development of spatial intelligence technology. Furthermore, the technical problems of high acquisition cost, insufficient scene diversity, and small scale of three-dimensional scene data sets in the related art are solved.
[0079] In an alternative embodiment, in step S22, using a multiple depth estimation strategy to perform scene depth estimation on the two-dimensional image to obtain a target depth map includes the following method steps:
[0080] Step S221, using a relative depth estimation model to perform relative depth estimation on the two-dimensional image to obtain a relative depth map;
[0081] Step S222, using a metric depth estimation model to perform metric depth estimation on the two-dimensional image to obtain a metric depth map;
[0082] Step S223, fusing the relative depth map and the metric depth map to obtain a target depth map.
[0083] In the embodiments of the present application, when performing scene depth estimation on a two-dimensional image using a multiple depth estimation strategy to obtain a target depth map, a relative depth estimation model can be first used to perform relative depth estimation on the two-dimensional image to obtain a relative depth map. The relative depth estimation model is used to predict a relative depth map from the two-dimensional image. The predicted relative depth map does not contain an absolute scale but has rich details. It can be understood that the relative depth estimation model is used to predict the relative depth difference between pixels in the two-dimensional image. The relative depth estimation model usually does not need to know the absolute distance of an object, but learns the relative distance relationship between objects. Therefore, the output depth map can reflect details but lacks a measurement scale. Exemplarily, the relative depth estimation model can be a monocular relative depth estimation model, a large model, or other deep learning models, which are not limited herein.
[0084] A relative depth map can be understood as the depth map output by the relative depth estimation model, where the value of each pixel represents its depth relationship relative to other pixels, but there is no absolute depth value or unit. It can be seen that the relative depth map often has rich details and can capture the subtle structures and texture changes in the image.
[0085] The present application uses a relative depth estimation model to perform relative depth estimation on a two-dimensional image, predicts a relative depth map from the two-dimensional image, and thus obtains a relative depth map. By capturing the subtle structures and textures in the two-dimensional image, it helps to reconstruct the shape and surface features of the object in the subsequent process. At the same time, since the relative depth estimation model does not require metric depth information for training, it can utilize a wider range of data sets, including those two-dimensional images without accurate distance annotations, thereby reducing the training threshold of the relative depth estimation model.
[0086] Exemplarily, when using a relative depth estimation model to perform relative depth estimation on a two-dimensional image, a three-dimensional point cloud can be first estimated for the two-dimensional image to provide a richer geometric representation, and then the relative depth map can be derived from the three-dimensional point cloud. By imposing a multi-scale local geometric loss on the local differences in the three-dimensional point cloud under an independent affine transformation, relatively high local geometric accuracy can be achieved, and a relatively accurate relative depth map with a robust three-dimensional geometric structure can be obtained, which is not limited herein.
[0087] Meanwhile, a metric depth estimation model can be used to perform metric depth estimation on a two-dimensional image to obtain a metric depth map. Among them, the metric depth estimation model is used to predict an approximate scaled depth map. The metric depth estimation model is trained in advance on data with true scale data (such as images with laser point clouds) and can provide a global distance scale for the scene. However, the absolute depth model is often less detailed than the relative depth model. That is to say, the metric depth estimation model is used to predict the absolute depth of each pixel in the image and is usually trained on a training data set with metric scale information, so it can output a depth map with actual distance units. Exemplarily, the metric depth estimation model can be a large model or other deep learning models, which is not limited here.
[0088] The metric depth map can be understood as the depth map output by the metric depth estimation model, which contains the actual distance from each pixel to the observer. The unit can be meters or centimeters, which is not limited here. It can be seen that the metric depth map is more accurate on a global scale but may not be as rich in details as the relative depth map.
[0089] This application uses a metric depth estimation model to perform metric depth estimation (i.e., absolute depth estimation) on a two-dimensional image, predicts the absolute depth of each pixel in the two-dimensional image, and thus obtains a metric depth map, which can accurately determine the true size and distance of objects and the environment, and helps object size measurement and spatial layout analysis. At the same time, the accuracy of the metric depth map helps to effectively fuse information of different modalities (such as text descriptions, image appearances) with three-dimensional geometric information, thus promoting the training of multimodal large language models.
[0090] Exemplarily, when using the metric depth estimation model to perform metric depth estimation on a two-dimensional image, a relative depth map can be generated based on the two-dimensional image, and then the metric depth can be fine-tuned on a data set containing true depth. Or, the focal length can be incorporated into the input, and an end-to-end training method can be used to jointly predict the metric depth and surface normal, so as to combine with the camera internal parameters to obtain a reasonable scale and generate a more robust metric depth map, which is not limited here.
[0091] After obtaining the relative depth map and the metric depth map, the relative depth map and the metric depth map are fused to obtain the target depth map. The target depth map not only retains rich detail features but also has accurate metric information, ensuring the realism and scale consistency of 3D reconstruction. It can be understood that this application adopts two different depth estimation methods: relative depth estimation and metric depth estimation. Relative depth estimation focuses on image details and structures, while metric depth estimation emphasizes the accuracy of scale and distance. By integrating the advantages of the two depth maps, the generated target depth map not only retains rich detail features but also has accurate metric information, ensuring the realism and scale consistency of 3D reconstruction.
[0092] Exemplarily, various strategies can be adopted for fusion, such as methods based on weighted average, scale factor adjustment, depth Figure 1 consistency optimization, etc., which are not limited herein.
[0093] It can be seen that this application combines the fine local structure of the relative depth map with the global scale information of the metric depth map to obtain a target depth map that not only retains details but also has accurate metrics, providing high-quality depth estimation for subsequent 3D scene data generation. Moreover, the accurate scale information of the target depth map ensures the realism of 3D reconstruction, making the subsequent generated 3D scene data closer to the real world and improving the performance of the model in practical applications. In addition, as high-quality input, the target depth map can be used to train and validate models that require 3D perception capabilities, such as 3D multimodal large language models, to promote the improvement of model performance.
[0094] In an alternative embodiment, in step S223, fusing the relative depth map and the metric depth map to obtain the target depth map includes the following method steps:
[0095] Step S2231, determining a scaling factor based on the relative depth map and the metric depth map, where the relative depth map and the metric depth map have the same dimension;
[0096] Step S2232, performing scale calibration on the relative depth map based on the scaling factor to obtain the target depth map.
[0097] In the embodiment of this application, when fusing the relative depth map and the metric depth map to obtain the target depth map, the scaling factor can be determined first based on the relative depth map and the metric depth map. The relative depth map and the metric depth map have the same dimension, which can be understood as that for the same two-dimensional image, this application will generate relative depth and metric depth maps of the same size.
[0098] The scaling factor can be understood as a key parameter used to adjust the results of image depth estimation or feature matching, ensuring that these results can match the real world or other reference data in scale. In the embodiments of the present application, the scaling factor is used to adjust the scale of the relative depth map so that the relative depth map is consistent with the absolute scale of the metric depth map, thereby enabling the subsequent generated three-dimensional scene data to have an accurate metric scale and being applicable to application scenarios that require a real scale.
[0099] After determining the scaling factor, scale calibration is performed on the relative depth map based on the scaling factor to obtain the target depth map. Among them, in the embodiments of the present application, scale calibration can be understood as using the determined scaling factor to adjust the depth values of the relative depth map so that they match the scale of the metric depth map, thereby generating a target depth map with an accurate scale.
[0100] It can be seen that in the present application, by determining the scaling factor and then performing scale calibration on the relative depth map based on the scaling factor, that is, adjusting the numerical range of the relative depth map based on the scaling factor to align it with the scale of the metric depth map. Thereby, the consistency in scale between the relative depth map and the metric depth map is ensured, that is, the adjusted relative depth map and the metric depth map fuse depth information at the same position, so that the generated target depth map not only retains details but also has an accurate metric scale.
[0101] It can be understood that the scale calibration through the scaling factor in the present application is a fast and effective fusion method, avoiding complex optimization processes and improving the overall efficiency of three-dimensional data generation.
[0102] In an alternative embodiment, in step S2231, determining the scaling factor based on the relative depth map and the metric depth map includes the following method steps:
[0103] Step S22311, identifying an effective point set from the two-dimensional image;
[0104] Step S22312, calculating the average effective relative depth of the effective point set based on the relative depth map, and calculating the average effective metric depth of the effective point set based on the metric depth map;
[0105] Step S22313, calculating the scaling factor using the average effective relative depth and the average effective metric depth.
[0106] In the embodiments of the present application, when determining the scaling factor based on the relative depth map and the metric depth map, an effective point set can be first identified from the two-dimensional image. Among them, the effective point set (Valid Point Set) can be understood as a set of pixel points that have depth information and reasonable and credible depth values in both the relative depth map and the metric depth map. Effective points are usually located on the surface of the object in the two-dimensional image, and the depth values have been successfully predicted by the depth estimation model.
[0107] The present application first analyzes the input two-dimensional image to identify the pixel points in the two-dimensional image that can provide effective depth information in both relative depth estimation and metric depth estimation. These pixel points form the effective point set, which is the basis for calculating the scaling factor in the subsequent steps. It can be seen that by identifying the effective point set, the noise points and outliers in the depth map can be excluded, improving the accuracy of the scaling factor calculation.
[0108] Then, calculate the average effective relative depth of the effective point set based on the relative depth map, and calculate the average effective metric depth of the effective point set based on the metric depth map. Among them, the average effective relative depth can be understood as the average value of the relative depth calculated on the effective point set, which is used to reflect the average depth relationship of the effective points in the relative depth map. The average effective metric depth can be understood as the average value of the metric depth calculated on the effective point set, which is used to reflect the average actual depth value of the effective points in the metric depth map, usually in meters or centimeters, and is not limited here.
[0109] It can be seen that by statistically analyzing the depth values of the effective point set in the relative depth map and the metric depth map respectively, the respective average values are calculated. These average values provide an overview of the depth information of the effective point set on the two depth maps and are important data for subsequent calculation of the scaling factor.
[0110] Finally, use the average effective relative depth and the average effective metric depth to calculate the scaling factor. Exemplarily, the scaling factor s can be calculated according to formula (1), where V represents the effective point set, |V| represents the number of effective points in the effective point set, d m,i represents the metric depth of the effective point i, d r,i represents the relative depth of the effective point i, and the average effective relative depth is represented as The average effective metric depth is represented as
[0111]
[0112] It can be seen that the scaling factor is calculated through the average effective relative depth and the average effective metric depth, and then the relative depth map is adjusted using the scaling factor to make it scale with the metric depth Figure 1As a result, the quality of depth map fusion is improved, ensuring the geometric details and scale accuracy of the fused target depth map.
[0113] In an optional embodiment, the camera parameters include: the camera internal parameters of the two-dimensional image and the camera external parameters of the two-dimensional image. The image processing method further includes the following method steps:
[0114] Step S24: Use the camera internal parameter estimation model to estimate the internal parameters of the two-dimensional image to obtain the camera internal parameters, and use the field of view method to estimate the external parameters of the two-dimensional image to obtain the camera external parameters.
[0115] When projecting a 2D image into 3D space, accurate camera parameters are crucial. The camera parameters include: the camera internal parameters of the two-dimensional image and the camera external parameters of the two-dimensional image. The camera internal parameters and the camera external parameters together determine how the two-dimensional image aligns with the three-dimensional structure of the real world. Among them, the camera internal parameters of the two-dimensional image are used to describe the optical characteristics during camera imaging. The camera internal parameters may include information such as focal length, pixel center point coordinates, and the size of the image sensor, which are not limited here. The camera external parameters of the two-dimensional image are used to describe the position and orientation of the camera relative to the world coordinate system. The camera external parameters may include the position coordinates and orientation angles of the camera, which are not limited here.
[0116] In the embodiment of the present application, when determining the camera parameters of the two-dimensional image, the camera internal parameter estimation model can be used to estimate the internal parameters of the two-dimensional image to obtain the camera internal parameters. Among them, the camera internal parameter estimation model can be a large model or other deep learning models for estimating the internal parameters of the camera, that is, the camera internal parameters. Exemplarily, the camera internal parameter estimation model can use the Wild-Camera method to predict the camera internal parameters, and utilize the sensitivity of Wild-Camera to scale and the ability to detect cropping to accurately recover the 2D principal point and focal length, which are not limited here.
[0117] It can be seen that the present application uses the camera internal parameter estimation model to estimate the internal parameters of the two-dimensional image, automatically inferring the internal characteristic parameters of the camera, such as focal length and the center point of the two-dimensional image, from the input two-dimensional image, which helps to convert the two-dimensional image coordinates into spatial coordinates and perform the conversion from depth map to point cloud in the subsequent process. At the same time, the use of the camera internal parameter estimation model enables the automatic acquisition of necessary internal parameter information even in the absence of camera parameter metadata, which is particularly important for processing images with unknown parameters and improves the versatility of the solution.
[0118] Meanwhile, the external parameters of the camera can be estimated for the two-dimensional image in the field of view mode to obtain the external camera parameters. Among them, the field of view mode is used to infer the orientation and position of the camera by analyzing features such as the distribution of objects in the image and the position of the horizon, and then estimate the external camera parameters. Exemplarily, the field of view mode can be the PerspectiveFields mode, and PerspectiveFields can provide the upward vector and latitude value of each pixel, from which a rotation matrix can be constructed to align the generated point cloud with the orientation of the standard 3D data set (z-axis upward), so as to ensure that the reconstructed scene matches the perspective of the real world, which is not limited here.
[0119] It can be seen that the present application uses the field of view mode to estimate the external parameters of the two-dimensional image to obtain the external camera parameters, that is, the position and orientation of the camera in the three-dimensional space. Accurate external camera parameters can ensure that the point clouds between different views can be correctly aligned, so as to construct a more complete and detailed three-dimensional scene.
[0120] Accurate internal camera parameters and external camera parameters are the basis for three-dimensional reconstruction and coordinate transformation. By estimating accurate internal camera parameters and external camera parameters, the present application can more accurately convert the pixel points in the two-dimensional image into coordinate points in the three-dimensional space, making the generated point cloud and depth map closer to the real scene. At the same time, the accurate estimation of the camera parameters helps the generated three-dimensional data to maintain a consistent scale and perspective with the real world, which is crucial for enhancing the realism of the three-dimensional data and improving the performance of the model trained based on these data in real-world applications.
[0121] In an optional embodiment, in step S23, generating three-dimensional scene data based on the camera parameters of the two-dimensional image and the target depth map includes the following method steps:
[0122] Step S231, generating an initial three-dimensional point cloud representation corresponding to the two-dimensional image based on the internal camera parameters, the external camera parameters, and the target depth map;
[0123] Step S232, post-processing the initial three-dimensional point cloud representation to obtain a target three-dimensional point cloud representation, where the post-processing includes at least some of the following processes: noise filtering processing, noise smoothing processing, and missing area filling processing.
[0124] In the embodiment of the present application, when generating three-dimensional scene data based on the camera parameters of the two-dimensional image and the target depth map, an initial three-dimensional point cloud representation corresponding to the two-dimensional image can be first generated based on the internal camera parameters, the external camera parameters, and the target depth map. It can be understood that by using the aforementioned estimated internal camera parameters and external camera parameters, combined with the depth information of each pixel in the target depth map, an initial three-dimensional point cloud representation is converted and generated.
[0125] It is understandable that the initial three-dimensional point cloud representation may contain some outliers or noise (such as holes or floating points caused by depth estimation errors). Therefore, post-processing of the initial three-dimensional point cloud representation is also required. Among them, the post-processing includes at least some of the following processes: noise filtering process, noise smoothing process, and missing area filling process.
[0126] The noise filtering process is used to identify and remove those unreasonable and abnormal three-dimensional point clouds, such as misjudged points caused by depth estimation errors or point clouds generated from unstructured areas in two-dimensional images, so as to improve the purity and reliability of the point cloud.
[0127] The noise smoothing process is used to smooth the noise points in the point cloud through the geometric relationship of neighboring points, ensure the continuity and smoothness of the point cloud surface, and enhance the visual effect of the three-dimensional scene.
[0128] The missing area filling process is used to fill the holes or missing areas in the point cloud caused by incomplete depth estimation or occlusion. Through algorithms such as interpolation and diffusion, based on the existing point cloud information, the three-dimensional structure of the missing area is reasonably inferred to enhance the integrity of the point cloud.
[0129] That is to say, in this application, by filtering out obviously incorrect points (such as points beyond a reasonable distance range), smoothing neighborhood noise, and optionally filling small gaps through depth continuity in the initial three-dimensional point cloud representation, the target three-dimensional point cloud representation, that is, three-dimensional scene data, is finally obtained. Each point in the target three-dimensional point cloud representation carries the color information of the corresponding pixel, presenting an appearance texture that conforms to the original two-dimensional image.
[0130] It can be seen that based on the initial three-dimensional point cloud generated from the camera parameters and the target depth map, and through a series of post-processings such as noise filtering, smoothing, and missing area filling on the initial three-dimensional point cloud, the finally generated target three-dimensional point cloud representation has high purity, high coherence, and integrity, greatly improving the quality of the generated three-dimensional data and its performance in practical applications.
[0131] In an alternative embodiment, in step S231, generating the initial three-dimensional point cloud representation corresponding to the two-dimensional image based on the camera internal parameters, camera external parameters, and target depth map includes the following method steps:
[0132] Step S2311, according to the camera internal parameters and the target depth map, project the pixel coordinates of the two-dimensional image into the camera coordinate system to generate a three-dimensional point set in the camera coordinate system;
[0133] Step S2312, based on the camera external parameters, convert the three-dimensional point set in the camera coordinate system to the initial three-dimensional point cloud representation in the world coordinate system.
[0134] In the embodiments of the present application, when generating an initial three-dimensional point cloud representation corresponding to a two-dimensional image based on the camera intrinsic parameters, camera extrinsic parameters, and the target depth map, the pixel coordinates of the two-dimensional image can be projected onto the camera coordinate system according to the camera intrinsic parameters and the target depth map to generate a three-dimensional point set in the camera coordinate system. The camera coordinate system can be understood as a three-dimensional rectangular coordinate system centered on the camera, where the X-axis, Y-axis, and Z-axis represent the horizontal, vertical, and depth directions, respectively.
[0135] Exemplarily, the present application first analyzes the depth information of each pixel point in the two-dimensional image using the camera intrinsic parameters and the target depth map to obtain the depth information of each pixel point. Then, using the focal length and other camera parameters in the camera intrinsic parameters, these pixel points are projected from the two-dimensional image plane into the camera coordinate system to generate a point set with three-dimensional coordinates. Thus, by associating the depth information with the pixel coordinates, the effective utilization of the depth information from the two-dimensional image to the three-dimensional space is achieved.
[0136] After that, based on the camera extrinsic parameters, the three-dimensional point set in the camera coordinate system is converted into an initial three-dimensional point cloud representation in the world coordinate system. The world coordinate system can be understood as a fixed three-dimensional rectangular coordinate system used to describe the absolute positions of all objects and points in a three-dimensional scene.
[0137] It can be seen that the generated three-dimensional point set in the camera coordinate system is transformed through the camera extrinsic parameters into the world coordinate system to obtain the initial three-dimensional point cloud representation of the scene. Thus, it is ensured that each point in the point cloud has an accurate position relative to the world coordinate system, providing global positioning information for subsequent scene fusion, object recognition, and spatial understanding. At the same time, based on the camera extrinsic parameters, the perspective consistency between the point cloud data and the original image can be maintained, which is crucial for multi-view fusion and the correct display of 3D data.
[0138] In an alternative embodiment, the image processing method further includes the following method steps:
[0139] Step S233, determining the maximum depth value and the minimum depth value of the target three-dimensional point cloud representation;
[0140] Step S234, constructing a three-dimensional bounding box according to the maximum depth value and the minimum depth value;
[0141] Step S235, converting the three-dimensional bounding box into three-dimensional annotation information based on the two-dimensional annotation information of the two-dimensional image.
[0142] In the embodiments of the present application, while generating 3D scene data, the corresponding 3D annotations can be automatically generated using the annotations of the original 2D images, that is, 3D automatic annotation migration is realized. Further, first determine the maximum depth value and the minimum depth value represented by the target three-dimensional point cloud. Among them, the maximum depth value can be understood as the depth value of the point farthest from the camera in the three-dimensional point cloud. The minimum depth value can be understood as the depth value of the point closest to the camera in the three-dimensional point cloud.
[0143] By determining the maximum depth value and the minimum depth value represented by the target three-dimensional point cloud, it helps to comprehensively understand the depth range of the scene and provides key coordinate information for the subsequent construction of the three-dimensional bounding box.
[0144] Then, based on the maximum depth value and the minimum depth value, construct a three-dimensional bounding box (3D Bounding Box). Among them, the three-dimensional bounding box can be understood as a coordinate box used to enclose and locate objects in three-dimensional space, usually defined by 8 vertices, and is applicable to object detection and positioning in 3D scenes.
[0145] Based on the determined maximum depth value and minimum depth value, construct one or more three-dimensional bounding boxes to enclose the objects or regions in the point cloud. By determining the boundaries of the objects in three-dimensional space, the position, size, and shape of the objects can be more accurately described, and thus the relative position and spatial relationship of the objects in three-dimensional space can be more intuitively understood and modeled.
[0146] Finally, convert the three-dimensional bounding box into three-dimensional annotation information based on the two-dimensional annotation information of the two-dimensional image. Among them, the two-dimensional annotation information of the two-dimensional image can be understood as the object position and category annotation on the two-dimensional image, such as the Bounding Box and semantic segmentation. The three-dimensional annotation information (3D annotations) can be understood as the markings of the position, size, pose, and category of the objects in three-dimensional space, usually including the 3D bounding box and semantic labels of the objects.
[0147] By using the two-dimensional annotation information on the two-dimensional image, the previously constructed three-dimensional bounding box is further converted into three-dimensional annotation information including the object category and pose. Exemplarily, the object bounding box in the two-dimensional image can be projected onto the three-dimensional point cloud, and the actual range of the object in three-dimensional space can be determined by combining the depth information, and the semantic label in the two-dimensional image can be assigned to the corresponding three-dimensional bounding box. Thus, by fusing the annotation information of the two-dimensional image with the three-dimensional point cloud data, the transition of semantic information from two-dimensional to three-dimensional is realized, enriching the semantic content of the point cloud. At the same time, there is no need to manually re-label the objects in three-dimensional space, and the corresponding three-dimensional annotations can be automatically generated based on the two-dimensional annotation information, greatly improving the efficiency and automation degree of data annotation.
[0148] Figure 3It is a schematic flowchart of generating 3D scene data based on 2D images according to an embodiment of the present application. It can be seen that the embodiment of the present application proposes an automated method to "elevate" a single 2D image into 3D scene data. The core idea is to combine advanced depth estimation algorithms, camera parameter calibration, and scale calibration techniques to reconstruct a complete, as real as possible in scale and appearance, three-dimensional representation from ordinary RGB images, including point clouds, camera poses (attitudes), depth maps, and pseudo RGB-D data, etc. Figure 3 The specific process includes the following parts.
[0149] The input and preprocessing part (see the part of preparing data in Figure 3 ), that is, the input is a two-dimensional image with basic annotations ([in Figure 3 the daily scene photos from the two-dimensional dataset are taken as an example). If the camera parameters (such as focal length, sensor size) of the two-dimensional image are unknown, the present application first uses camera calibration techniques to infer or estimate the internal and external parameters of the image, so as to ensure the correct projection relationship of the subsequent generated 3D point cloud.
[0150] The depth estimation and fusion part (see the part of depth prediction and scaling in Figure 3 ), the present application adopts a multiple depth estimation strategy to obtain high-quality scene depth information. On the one hand, a relative depth estimation model is used to predict a relative depth map (depth distribution without absolute scale but rich in details) from the two-dimensional image. On the other hand, a metric depth estimation model (also called an absolute depth estimation model) is used to predict an approximate scaled depth map, that is, a metric depth map. This absolute depth prediction is trained in advance on data with real scale data (such as images with laser point clouds), which can provide a global distance scale for the scene. However, the absolute depth model is often not as fine in details as the relative depth model. Therefore, the present application fuses the depth maps predicted by the two models (that is, the relative depth map and the metric depth map), not only retaining the local geometric details in the relative depth map, but also adjusting its overall numerical range to match the scale of the metric depth map. The fusion process can be achieved through methods such as scale factor correction, that is, performing a linear transformation on the relative depth map to make it consistent with the metric depth map in the overlapping area. After fusion and scale calibration, a target depth map with both real-world scale and clear details can be obtained.
[0151] The point cloud reconstruction and screening part (see Figure 3(in the camera parameter section and the 3D representation generation and post - processing section), using the calibrated camera parameters and the corrected depth map, project the depth information into a three - dimensional coordinate system to generate a dense three - dimensional point cloud. Specifically, according to the camera intrinsic parameters, convert the image pixel coordinates and the corresponding depth values into (X, Y, Z) points in the camera coordinate system to obtain the initial three - dimensional point cloud representation. The initially generated point cloud may contain some outliers or noise (such as holes or floating points caused by depth estimation errors). Therefore, this application needs to post - process the initial three - dimensional point cloud representation, such as filtering out significantly incorrect points (such as points beyond a reasonable distance range), smoothing neighborhood noise, and optionally filling small gaps through depth continuity. Thus, a three - dimensional point cloud representation of the scene is obtained, that is, the target three - dimensional point cloud representation, where each point carries the color information of the corresponding pixel, presenting an appearance texture consistent with the original image.
[0152] The camera pose determination part (see Figure 3 in the camera parameter section). While generating the point cloud, this application also records or calculates the pose information of the camera. Usually, the input image can be regarded as an observation perspective in the reconstructed scene. The camera pose includes the position and orientation (attitude) of the camera, usually represented in the world coordinate system. If a standard coordinate system is used, the camera can be placed at the origin, and the orientation is inferred from the content of the image (for example, the horizon position is used to determine the pitch angle). The camera pose is of reference significance for subsequent multi - view fusion or scene expansion.
[0153] The pseudo RGB - D generation part (see Figure 3 in the depth prediction and scaling section). Combining the original input RGB image and the predicted target depth map, this application can also generate the corresponding pseudo RGB - D data. The pseudo RGB - D data can be regarded as a four - channel image, where the RGB channels come from the original two - dimensional image, and the D channel is the estimated depth value. With the RGB - D representation, during the subsequent training of the spatial intelligence model, it can be directly used as the output of a real RGB - D sensor, providing dual information of color and geometry for the spatial intelligence model. And the pseudo RGB - D data is visually close to the results collected by a real depth camera.
[0154] The 3D automatic annotation transfer part (see Figure 3In the 3D representation generation and post-processing section), a major advantage of the present application is that while generating 3D scene data, the corresponding 3D annotations are automatically generated using the annotations of the original 2D image. For example, if the input image carries the bounding box, segmentation mask or semantic label of an object, the present application projects these 2D annotations onto the three-dimensional point cloud through depth information. Specifically, the 3D points converted from the pixels belonging to an object will inherit the label of the object, thereby forming a "point cloud instance" or semantic point cloud with category labels in the point cloud. As a result, the present application can obtain a completely annotated 3D scene data, which includes not only geometry and texture, but also semantic information of each major object or area. Such automatic 3D annotation eliminates the need for manual point-by-point labeling in three-dimensional space, greatly improving data annotation efficiency.
[0155] It can be seen that after the processing of the above parts, a single 2D image is converted into rich 3D scene data. The final output data includes: reconstructed point cloud (including color texture and semantic labels), internal and external parameters / pose of the corresponding camera, depth map, and RGB-D map formed by the combination of the original RGB image and depth. This output set can be regarded as a simulation of a "3D scan", whose scale is aligned with the real world, the visual effect is close to real shooting, and it comes with complete 3D annotations, which can be directly used to train AI models related to spatial intelligence.
[0156] Therefore, the rich annotations of existing mature 2D data (such as COCO, Objects365-v2, etc.) are combined with technologies such as multi-view depth estimation, camera extrinsic parameter estimation, and scale calibration to achieve large-scale, low-cost generation of three-dimensional point clouds, depth maps, and various three-dimensional annotations, thereby facilitating the training and verification of spatial intelligence and AI models with spatial reasoning capabilities. It can be seen that this application combines the rich annotations of existing mature 2D data with technologies such as multi-view depth estimation, camera extrinsic parameter estimation, and scale calibration to achieve large-scale, low-cost generation of three-dimensional point clouds, depth maps, and various three-dimensional annotations, thereby facilitating the training and verification of spatial intelligence and AI models with spatial reasoning capabilities. That is, this application can generate multimodal three-dimensional data representations from a single image to provide support for various downstream three-dimensional vision tasks.
[0157] It can be seen that the present application can realize sensorless monocular 3D reconstruction, that is, the present application has achieved a breakthrough in generating 3D scene data with only a single RGB image, getting rid of the dependence on hardware such as laser radar and depth camera. Algorithm-driven data generation significantly reduces the cost of obtaining three-dimensional training data.
[0158] This application can also achieve scale calibration and fine depth fusion. That is, this application proposes a method of fusing the relative depth and absolute depth estimation results and performing scale calibration, generating a 3D representation with both fine geometric details and real-world measurement scales. Thereby, it ensures that the synthetic data is credible in terms of physical size and distance and can be used for tasks that require accurate distance measurement and size judgment.
[0159] This application can also achieve automated 3D annotation transfer. That is, this application uses the rich annotations of the original 2D data and automatically projects them onto the generated 3D scene to achieve out-of-the-box three-dimensional annotation data. Each 3D scene generated in this way comes with object-level and pixel-level semantic information, eliminating the need for manual re-annotation and greatly improving the data annotation efficiency and consistency.
[0160] This application also has general scene adaptability. That is, this application is not limited to a specific environment and can perform 3D reconstruction on 2D images of various scenes such as indoor, outdoor, artificial, and natural. Through its application on a diverse dataset, the generality and robustness of this application are demonstrated, filling the gap in the lack of 3D data in diverse scenarios.
[0161] This application also has the ability to expand large-scale data. That is, this application can massively convert a large corpus of 2D images into corresponding 3D datasets, achieving exponential 3D data expansion capabilities. This scalability plays a crucial role in training models that require massive data such as future 3D-MLLMs, laying a data foundation for the spatial cognition of general artificial intelligence.
[0162] In summary, the 2D-to-3D data generation method of this application comprehensively overcomes the defects of traditional technologies in terms of photorealism, scale accuracy, scene complexity, data scale, and automatic annotation, providing an efficient and reliable solution for the training of spatial intelligent AI.
[0163] It is understandable that the 3D scene data generated by this application can be widely applied to artificial intelligence products and services that require three-dimensional perception and understanding. Exemplarily, it can be applied to intelligent robots, such as autonomous mobile robots, service robots, etc. By using the 3D scene data generated by this application for training in a simulated or real environment, better obstacle avoidance, path planning, and human-computer interaction capabilities can be obtained. It can also be applied to autonomous driving. Autonomous driving vehicles require a large amount of reliable 3D data to perceive the surrounding environment. This application can quickly generate 3D assets in various scenarios (such as urban streets, indoor and outdoor parking lots, etc.) to assist in the training of perception and decision-making networks. It can also be applied to augmented reality (AR) / virtual reality (VR) scenarios. In augmented reality and virtual reality systems, it is necessary to model the real world and interact with virtual objects. This application can extract three-dimensional scenes from a large number of 2D images to enhance the authenticity and diversity of AR / VR content. It can also be applied to 3D multimodal large language models. As the demand for spatial understanding in MLLMs becomes stronger, the massive, highly realistic, and metric-scale 3D data generated by this application can be used as a training set to improve the performance of the model in three-dimensional vision, spatial reasoning, and cross-modal tasks (such as 3D question answering, 3D description, three-dimensional retrieval, etc.), without limitation here.
[0164] It is easy to understand that the beneficial effects of the image processing method provided by this application include the following points.
[0165] Beneficial effect (1), in terms of authenticity, this application uses 2D images of the real world as the starting point, retaining the texture and details of the original photos, thus avoiding the problems of over-simplification and cartoonization in the data generated by simulators. The 3D scene data generated by this application is highly consistent with the real environment in terms of visual appearance, greatly narrowing the gap between "virtual data and real applications".
[0166] Beneficial effect (2), in terms of scene complexity and rationality, this application reconstructs 3D based on the actually photographed scenes. Therefore, the layout and scale ratio of objects in the scene naturally conform to the physical logic of the real world. This makes up for the unreasonable phenomena such as disproportion and object suspension that may occur in the purely imagined 3D generated by AI. Through depth estimation fusion and scale calibration, this application ensures that the size and distance of each reconstructed object are as close to the real value as possible, achieving the restoration of the real-world scale, which is difficult to achieve by pure generation methods.
[0167] Advantageous effect (3): In terms of data cost and scale, compared with sensor acquisition, the present application is completely based on software methods, without the need for expensive equipment and on-site acquisition processes. As long as there is a large amount of 2D image data, the present application can automatically be converted into corresponding 3D data. This means that the present application can "upgrade" the existing huge 2D database to a 3D database at low cost, that is, the present application can convert a million-level image dataset into corresponding 3D point cloud data, thereby obtaining a three-dimensional scene set far exceeding the previous scale at one time. This scale effect makes it possible to train larger and more complex spatial intelligent models.
[0168] Advantageous effect (4): In terms of annotation efficiency and consistency, the present application has a built-in mechanism for migrating 2D annotations to 3D, realizing automated 3D annotation. In contrast, the 3D data collected traditionally often requires manual point-by-point or object-by-object marking, which is time-consuming and laborious. By the method of the present application, the labels originally used for 2D tasks (such as detection boxes, segmentation masks) are automatically mapped to the 3D point cloud, and the three-dimensional position, size, and category labels of each object will be generated accordingly. This kind of annotation is consistent and accurate because it directly comes from the original image annotation and the calculated depth, without subjective errors. This not only reduces the labor cost of data preparation but also ensures the semantic alignment between 2D and 3D annotations and improves the data quality.
[0169] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or fully authorized by all parties. And the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or reject.
[0170] In addition, it should also be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, certain steps can be carried out in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present application.
[0171] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in various embodiments of the present application.
[0172] According to an embodiment of the present application, there is also provided Figure 4 an image processing method as shown in Figure 4 a flowchart of an image processing method according to an embodiment of the present application. As shown in Figure 4 the figure, the method includes:
[0173] Step S41, obtaining a two-dimensional image;
[0174] Step S42, performing scene depth estimation on the two-dimensional image by using a multiple depth estimation strategy to obtain a target depth map, where the multiple depth estimation strategy is used to obtain scene depth information of a given image through different types of depth estimation methods;
[0175] Step S43, generating three-dimensional scene data based on the camera parameters of the two-dimensional image and the target depth map;
[0176] Step S44, training an initial three-dimensional multimodal language model according to the three-dimensional scene data to obtain a target three-dimensional multimodal language model, where the target three-dimensional multimodal language model is used to process at least one of the following tasks: three-dimensional vision tasks, spatial reasoning tasks, and cross-modal tasks.
[0177] In the embodiments of the present application, the three-dimensional scene data generated based on the foregoing embodiments can be used to train a three-dimensional multimodal language model. Specifically, the description of the steps for generating three-dimensional scene data can refer to the description of the foregoing embodiments and will not be elaborated here.
[0178] Training an initial three-dimensional multimodal language model according to the three-dimensional scene data to obtain a target three-dimensional multimodal language model, where the initial three-dimensional multimodal language model can be understood as a model with the basic ability to process modal information such as three-dimensional data and text. Training the initial three-dimensional multimodal language model with the three-dimensional scene data can train a target three-dimensional multimodal language model.
[0179] The target three-dimensional multimodal language model is used to process at least one of the following tasks: three-dimensional vision tasks, spatial reasoning tasks, and cross-modal tasks. Among them, three-dimensional vision tasks can be understood as computer vision tasks carried out in three-dimensional space, involving the understanding and analysis of three-dimensional scenes. Spatial reasoning tasks focus on understanding and processing the layout, relationships, and dynamics of objects in three-dimensional space. Cross-modal tasks involve processing and integrating information from different modalities (such as images, text, sounds, touch, etc.).
[0180] It can be seen that this application can start from two-dimensional images, generate three-dimensional scene data, and finally use it for the training of the three-dimensional multimodal language model, thus fully demonstrating the effective utilization of three-dimensional data. Moreover, the generated three-dimensional data is used to train the model, significantly improving the model's intelligent perception and interaction capabilities in a three-dimensional environment, and promoting the development of the field of spatial intelligence. In addition, the trained target three-dimensional multimodal language model can be applied to more extensive and complex scenarios, such as augmented reality, virtual reality, robot vision, etc., accelerating the commercialization process of related technologies.
[0181] The above image processing method provided by the embodiments of this application can be but is not limited to being applied to application scenarios involving the generation of three-dimensional scene data in fields such as e-commerce services, education services, legal services, medical services, conference services, social network services, financial product services, logistics services, and navigation services. For example: image processing related to e-commerce services, image processing related to education services, image processing related to legal services, etc., which are not limited here.
[0182] By adopting the embodiments of this application, by obtaining two-dimensional images, a multiple depth estimation strategy is used to estimate the scene depth of the two-dimensional images to obtain a target depth map. Among them, the multiple depth estimation strategy is used to obtain the scene depth information of a given image through different types of depth estimation methods, and then three-dimensional scene data is generated based on the camera parameters of the two-dimensional images and the target depth map. Finally, an initial three-dimensional multimodal language model is trained based on the three-dimensional scene data to obtain a target three-dimensional multimodal language model. Among them, the target three-dimensional multimodal language model is used to process at least one of the following tasks: three-dimensional vision tasks, spatial reasoning tasks, and cross-modal tasks. Thus, the purpose of efficiently and low-costly generating three-dimensional scene data based on two-dimensional images is achieved, thereby realizing the low-cost acquisition of large-scale and high-quality three-dimensional scene data, and the generated three-dimensional scene data has both real-world scale and clear details, significantly promoting the development of spatial intelligence technology, and further solving the technical problems of high acquisition cost, insufficient scene diversity, and small scale in the three-dimensional scene dataset in related technologies.
[0183] It should be noted that the preferred implementation manners of this embodiment can be referred to the relevant descriptions in the embodiment, which will not be elaborated here.
[0184] According to an embodiment of the present application, there is also provided an Figure 5 image processing method as shown Figure 5 in the flowchart of an image processing method according to an embodiment of the present application. As shown Figure 5 in the figure, the method includes:
[0185] Step S51, obtaining an image processing request through a first application programming interface. Among them, the request data carried in the image processing request includes: a two-dimensional image;
[0186] Step S52, returning an image processing response through a second application programming interface. Among them, the response data carried in the image processing response includes: three-dimensional scene data, and the three-dimensional scene data is generated according to the image processing method of any one of the above.
[0187] The above first application programming interface and the second application programming interface can be either the same application programming interface or different application programming interfaces. In an optional embodiment, the interface parameters in the above first application programming interface and the second application programming interface may include, but are not limited to: interface global identifier, interface signature key, interface timestamp, interface request identifier, system call credential identifier, etc. The above first application programming interface can use GET or POST as the interface request method to obtain the file processing request. The above second application programming interface can use the JSON format to feedback the file processing response.
[0188] In the embodiment of the present application, the image processing request can be understood as a request sent by an external system to the image processing service, aiming to generate three-dimensional scene data according to the two-dimensional image. The image processing response can be understood as a response message to the image processing request, that is, a response message containing the generation result (three-dimensional scene data).
[0189] Specifically, the description of the steps for generating three-dimensional scene data can be referred to the description of the foregoing embodiments, and will not be elaborated here.
[0190] It can be seen that the service-oriented architecture implemented by the present application through the application programming interface (API) enables the processing of two-dimensional images to three-dimensional data to be performed far from the original data source, supporting distributed and remote data processing requirements. And the service-oriented mode allows efficient utilization of computing resources, processing image data on demand, avoiding unnecessary resource waste and cost expenditure. In addition, the API interface enables different systems and platforms to share the three-dimensional data conversion service, promoting cross-domain data collaboration and applications, such as sharing a unified three-dimensional data processing process in fields such as AR / VR, autonomous driving, and robotics.
[0191] The above image processing method provided by the embodiments of the present application can be, but is not limited to, applied to application scenarios involving the generation of three-dimensional scene data in fields such as e-commerce services, education services, legal services, medical services, conference services, social network services, financial product services, logistics services, and navigation services. For example, image processing related to e-commerce services, image processing related to education services, image processing related to legal services, etc., which are not limited here.
[0192] By adopting the embodiments of the present application, an image processing request is obtained through a first application programming interface. Among them, the request data carried in the image processing request includes: a two-dimensional image. Then, an image processing response is returned through a second application programming interface. Among them, the response data carried in the image processing response includes: three-dimensional scene data, and the three-dimensional scene data is generated according to the image processing method of any one of the above. Thus, the purpose of efficiently and low-costly generating three-dimensional scene data based on two-dimensional images is achieved, thereby realizing the low-cost acquisition of large-scale and high-quality three-dimensional scene data. Moreover, the generated three-dimensional scene data has both real-world scale and clear details, significantly promoting the development of spatial intelligence technology, and further solving the technical problems of high acquisition cost, insufficient scene diversity, and small scale in the three-dimensional scene data set in the related art.
[0193] It should be noted that the preferred implementation manners of this embodiment can be referred to the relevant descriptions in the embodiment, which will not be elaborated here.
[0194] According to the embodiments of the present application, there is also provided an image processing method as Figure 6 shown. Figure 6 is a flowchart of an image processing method according to the embodiments of the present application, as Figure 6 shown, and the method includes:
[0195] Step S61, obtaining the currently input image processing dialogue request. Among them, the request data carried in the image processing dialogue request includes: a two-dimensional image;
[0196] Step S62, in response to the image processing dialogue request, returning an image processing dialogue reply. Among them, the information carried in the image processing dialogue reply includes: three-dimensional scene data, and the three-dimensional scene data is generated according to the image processing method of any one of the above;
[0197] Step S63, displaying the three-dimensional scene data in the graphical user interface.
[0198] In the embodiments of the present application, the image processing dialogue request can be understood as a request initiated by a user in an interactive application for the purpose of generating three-dimensional scene data based on a two-dimensional image.
[0199] The image processing dialogue reply can be understood as the reply content to the image processing dialogue request, that is, the generation result (3D scene data) returned by the system to the user after generating 3D scene data based on a 2D image.
[0200] The graphical user interface can be understood as a visual interface that allows users to interact with a program through intuitive operations such as mouse clicks and drags.
[0201] Specifically, the steps for generating 3D scene data can be referred to the description of the foregoing embodiments, which will not be elaborated here.
[0202] It can be seen that users can not only initiate an image processing dialogue request through simple graphical user interface operations, but also obtain and intuitively view the processing results in real time. This greatly facilitates users' acquisition and understanding of 3D scene data, and at the same time improves the interactivity and usability of the system in the field of image processing.
[0203] The above image processing method provided by the embodiments of the present application can be but is not limited to being applied to application scenarios involving 3D scene data generation in fields such as e-commerce services, education services, legal services, medical services, conference services, social network services, financial product services, logistics services, and navigation services. For example: image processing related to e-commerce services, image processing related to education services, image processing related to legal services, etc., which will not be limited here.
[0204] By adopting the embodiments of the present application, by obtaining the currently input image processing dialogue request, wherein the request data carried in the image processing dialogue request includes: a 2D image, and in response to the image processing dialogue request, returning an image processing dialogue reply, wherein the information carried in the image processing dialogue reply includes: 3D scene data, and the 3D scene data is generated according to the image processing method of any one of the above, and the 3D scene data is displayed within the graphical user interface, thereby achieving the purpose of efficiently and low-costly generating 3D scene data based on 2D images, thus realizing low-cost acquisition of large-scale and high-quality 3D scene data, and the generated 3D scene data has both real-world scale and clear details, significantly promoting the development of spatial intelligence technology, and further solving the technical problems of high acquisition cost, insufficient scene diversity and small scale in the 3D scene dataset in the related art.
[0205] It should be noted that the preferred implementation manners of this embodiment can be referred to the relevant descriptions in the embodiments, which will not be elaborated here.
[0206] According to the embodiments of the present application, there is also provided an Figure 7 image processing method as shown. Figure 7 is a flowchart of an image processing method according to the embodiments of the present application, as Figure 7 shown, the method includes:
[0207] Step S71: In response to an input instruction applied to the operation interface, display a two-dimensional image on the operation interface.
[0208] Step S72: In response to a processing instruction applied to the operation interface, display three-dimensional scene data on the operation interface; wherein, the three-dimensional scene data is generated according to the image processing method of any one of the above.
[0209] In the embodiments of the present application, the operation interface can be understood as an interface with which the user can interact, allowing the user to input instructions and receive system feedback. The input instruction can be understood as an instruction issued by the user to request the system to display a two-dimensional image to be processed on the operation interface. The processing instruction can be understood as an instruction issued by the user on the operation interface, requiring the system to process the displayed two-dimensional image to generate corresponding three-dimensional scene data.
[0210] Specifically, for the description of the steps of generating three-dimensional scene data, reference can be made to the description of the foregoing embodiments, which will not be elaborated here.
[0211] It can be seen that by providing the user with an intuitive operation interface, it also supports real-time conversion and display from two-dimensional images to three-dimensional scene data, which not only improves the user experience, but also provides a more efficient and intuitive tool for data processing, teaching demonstrations, and application development.
[0212] The image processing method provided by the embodiments of the present application can be, but is not limited to, applied to application scenarios involving the generation of three-dimensional scene data in fields such as e-commerce services, education services, legal services, medical services, conference services, social network services, financial product services, logistics services, and navigation services. For example: image processing related to e-commerce services, image processing related to education services, image processing related to legal services, etc., which will not be limited here.
[0213] By adopting the embodiments of the present application, by responding to an input instruction applied to the operation interface, a two-dimensional image is displayed on the operation interface, and by responding to a processing instruction applied to the operation interface, three-dimensional scene data is displayed on the operation interface; wherein, the three-dimensional scene data is generated according to the image processing method of any one of the above, thereby achieving the purpose of efficiently and low-costly generating three-dimensional scene data based on two-dimensional images, thus realizing low-cost acquisition of large-scale and high-quality three-dimensional scene data, and the generated three-dimensional scene data has both real-world scale and clear details, significantly promoting the development of spatial intelligence technology, and further solving the technical problems of high acquisition cost, insufficient scene diversity, and small scale in the three-dimensional scene dataset in the related art.
[0214] It should be noted that the preferred implementation manners of this embodiment can be referred to the relevant descriptions in the embodiments, which will not be elaborated here.
[0215] According to an embodiment of the present application, there is also provided an Figure 8 image processing system as shown. Figure 8 FIG. is a schematic structural diagram of an image processing system according to an embodiment of the present application. As Figure 8 shown, the method includes:
[0216] A client for sending a two-dimensional image;
[0217] A server, connected to the client, for performing scene depth estimation on the two-dimensional image by using a multiple depth estimation strategy to obtain a target depth map, and generating three-dimensional scene data based on the camera parameters of the two-dimensional image and the target depth map, where the multiple depth estimation strategy is used to obtain scene depth information of a given image through different types of depth estimation methods, and the three-dimensional scene data is training data applied to three-dimensional perception and understanding scenarios;
[0218] The client is further configured to output the three-dimensional scene data.
[0219] The image processing system in the embodiment of the present application is used to execute the image processing method proposed in the embodiment of the present application. For details, reference may be made to the description of the foregoing embodiments, which will not be elaborated herein.
[0220] The above image processing system provided by the embodiment of the present application can be applied, but is not limited to, application scenarios involving the generation of three-dimensional scene data in fields such as e-commerce services, education services, legal services, medical services, conference services, social network services, financial product services, logistics services, and navigation services. For example: image processing related to e-commerce services, image processing related to education services, image processing related to legal services, etc., which are not limited herein.
[0221] By adopting the embodiment of the present application, through the image processing system, the purpose of efficiently and low-cost generating three-dimensional scene data based on two-dimensional images is achieved, thereby realizing the low-cost acquisition of a large-scale and high-quality three-dimensional scene data, and the generated three-dimensional scene data has both real-world scale and clear details, significantly promoting the development of spatial intelligence technology, and further solving the technical problems of high acquisition cost, insufficient scene diversity, and small scale of three-dimensional scene datasets in the related art.
[0222] It should be noted that the preferred implementation manners of this embodiment can be referred to the relevant descriptions in the embodiment, which will not be elaborated herein.
[0223] According to an embodiment of the present application, there is also provided an apparatus embodiment for implementing the above image processing method. Figure 9 FIG. is a schematic structural diagram of an image processing apparatus according to an embodiment of the present application. As Figure 9 shown, the apparatus includes:
[0224] The first acquisition module 901 is configured to acquire a two-dimensional image;
[0225] The first estimation module 902 is configured to perform scene depth estimation on the two-dimensional image by using a multiple depth estimation strategy to obtain a target depth map, where the multiple depth estimation strategy is used to obtain scene depth information of a given image through different types of depth estimation methods;
[0226] The first generation module 903 is configured to generate three-dimensional scene data based on the camera parameters of the two-dimensional image and the target depth map, where the three-dimensional scene data is training data applied to a three-dimensional perception and understanding scene.
[0227] Optionally, the first estimation module 902 is further configured to: perform relative depth estimation on the two-dimensional image by using a relative depth estimation model to obtain a relative depth map; perform metric depth estimation on the two-dimensional image by using a metric depth estimation model to obtain a metric depth map; fuse the relative depth map and the metric depth map to obtain a target depth map.
[0228] Optionally, the first estimation module 902 is further configured to: determine a scaling factor based on the relative depth map and the metric depth map, where the relative depth map and the metric depth map have the same dimension; perform scale calibration on the relative depth map based on the scaling factor to obtain a target depth map.
[0229] Optionally, the first estimation module 902 is further configured to: identify a valid point set from the two-dimensional image; calculate an effective relative depth average value of the valid point set according to the relative depth map, and calculate an effective metric depth average value of the valid point set according to the metric depth map; calculate a scaling factor by using the effective relative depth average value and the effective metric depth average value.
[0230] Optionally, the camera parameters include: the camera internal parameters of the two-dimensional image and the camera external parameters of the two-dimensional image, and the apparatus further includes: a second estimation module, configured to perform internal parameter estimation on the two-dimensional image by using a camera internal parameter estimation model to obtain camera internal parameters, and perform external parameter estimation on the two-dimensional image by using a field of view method to obtain camera external parameters.
[0231] Optionally, the first generation module 903 is further configured to: generate an initial three-dimensional point cloud representation corresponding to the two-dimensional image based on the camera internal parameters, the camera external parameters, and the target depth map; perform post-processing on the initial three-dimensional point cloud representation to obtain a target three-dimensional point cloud representation, where the post-processing includes at least some of the following processes: noise filtering processing, noise smoothing processing, missing area filling processing.
[0232] Optionally, the above-mentioned first generation module 903 is further configured to: project the pixel coordinates of the two-dimensional image into the camera coordinate system based on the target depth map involved in the camera, and generate a three-dimensional point set in the camera coordinate system; and convert the three-dimensional point set in the camera coordinate system into an initial three-dimensional point cloud representation in the world coordinate system based on the external camera parameters.
[0233] Optionally, the device further includes: a conversion module, configured to determine the maximum depth value and the minimum depth value of the target three-dimensional point cloud representation; construct a three-dimensional bounding box based on the maximum depth value and the minimum depth value; and convert the three-dimensional bounding box into three-dimensional annotation information based on the two-dimensional annotation information of the two-dimensional image.
[0234] By adopting the embodiment of the present application, a two-dimensional image is obtained, and a scene depth estimation is performed on the two-dimensional image by using a multiple depth estimation strategy to obtain a target depth map. The multiple depth estimation strategy is used to obtain the scene depth information of a given image through different types of depth estimation methods. Finally, three-dimensional scene data is generated based on the camera parameters of the two-dimensional image and the target depth map. The three-dimensional scene data is training data applied to the three-dimensional perception and understanding scene. Thus, the purpose of efficiently and low-costly generating three-dimensional scene data based on two-dimensional images is achieved, thereby realizing the low-cost acquisition of large-scale and high-quality three-dimensional scene data. The generated three-dimensional scene data has both real-world scale and clear details, significantly promoting the development of spatial intelligence technology. Furthermore, the technical problems of high acquisition cost, insufficient scene diversity, and small scale of the three-dimensional scene data set in the related art are solved.
[0235] It should be noted here that the above-mentioned first acquisition module 901, first estimation module 902, and first generation module 903 correspond to steps S21 to S23 in the embodiment. The examples and application scenarios implemented by the three modules and the corresponding steps are the same, but are not limited to the content disclosed in the above embodiment. It should be noted that the above-mentioned module or unit may be a hardware component or a software component stored in the memory and processed by one or more processors. The above-mentioned module may also run in the server 10 provided in the embodiment.
[0236] According to the embodiment of the present application, another device embodiment for implementing the above-mentioned image processing method is further provided. Figure 10 It is a schematic structural diagram of another image processing device according to the embodiment of the present application. As Figure 10 shown, the device includes:
[0237] A second acquisition module 1001, configured to acquire a two-dimensional image;
[0238] A third estimation module 1002 is configured to perform scene depth estimation on a two-dimensional image by adopting a multiple depth estimation strategy to obtain a target depth map, where the multiple depth estimation strategy is used to obtain scene depth information of a given image through different types of depth estimation methods;
[0239] A second generation module 1003 is configured to generate three-dimensional scene data based on the camera parameters of the two-dimensional image and the target depth map;
[0240] A training module 1004 is configured to train an initial three-dimensional multimodal language model based on the three-dimensional scene data to obtain a target three-dimensional multimodal language model, where the target three-dimensional multimodal language model is used to process at least one of the following tasks: three-dimensional vision tasks, spatial reasoning tasks, and cross-modal tasks.
[0241] By adopting the embodiments of the present application, by obtaining a two-dimensional image, performing scene depth estimation on the two-dimensional image by adopting a multiple depth estimation strategy to obtain a target depth map, where the multiple depth estimation strategy is used to obtain scene depth information of a given image through different types of depth estimation methods, then generating three-dimensional scene data based on the camera parameters of the two-dimensional image and the target depth map, and finally training an initial three-dimensional multimodal language model based on the three-dimensional scene data to obtain a target three-dimensional multimodal language model, where the target three-dimensional multimodal language model is used to process at least one of the following tasks: three-dimensional vision tasks, spatial reasoning tasks, and cross-modal tasks, the purpose of efficiently and low-costly generating three-dimensional scene data based on two-dimensional images is achieved, thereby realizing low-cost acquisition of large-scale and high-quality three-dimensional scene data, and the generated three-dimensional scene data has both real-world scale and clear details, significantly promoting the development of spatial intelligence technology, and further solving the technical problems of high acquisition cost, insufficient scene diversity and small scale of three-dimensional scene datasets in the related art.
[0242] It should be noted here that the above-mentioned second acquisition module 1001, third estimation module 1002, second generation module 1003, and training module 1004 correspond to steps S41 to S44 in the embodiment. The instances and application scenarios implemented by the four modules and the corresponding steps are the same, but are not limited to the content disclosed in the above embodiments. It should be noted that the above-mentioned modules or units can be hardware components or software components stored in a memory and processed by one or more processors, and the above-mentioned modules can also run in the server 10 provided in the embodiment.
[0243] According to an embodiment of the present application, there is also provided another device embodiment for implementing the above-mentioned image processing method. Figure 11 It is a schematic structural diagram of another image processing device according to an embodiment of the present application, as Figure 11 shown, the device includes:
[0244] A third acquisition module 1101, configured to acquire an image processing request through a first application programming interface, wherein request data carried in the image processing request includes: a two-dimensional image;
[0245] A first return module 1102, configured to return an image processing response through a second application programming interface, wherein response data carried in the image processing response includes: three-dimensional scene data, and the three-dimensional scene data is generated according to the image processing method of any one of the above.
[0246] By adopting the embodiment of the present application, an image processing request is acquired through a first application programming interface, wherein request data carried in the image processing request includes: a two-dimensional image, and then an image processing response is returned through a second application programming interface, wherein response data carried in the image processing response includes: three-dimensional scene data, and the three-dimensional scene data is generated according to the image processing method of any one of the above. Thus, the purpose of efficiently and low-costly generating three-dimensional scene data based on a two-dimensional image is achieved, thereby realizing low-cost acquisition of large-scale and high-quality three-dimensional scene data, and the generated three-dimensional scene data has both real-world scale and clear details, significantly promoting the development of spatial intelligence technology, and further solving the technical problems of high acquisition cost, insufficient scene diversity and small scale in the three-dimensional scene data set in the related art.
[0247] It should be noted here that the above-mentioned third acquisition module 1101 and first return module 1102 correspond to steps S51 and S52 in the embodiment. The instances and application scenarios implemented by the two modules and the corresponding steps are the same, but are not limited to the content disclosed in the above embodiment. It should be noted that the above-mentioned module or unit may be a hardware component or a software component stored in a memory and processed by one or more processors, and the above-mentioned module may also run in the server 10 provided in the embodiment.
[0248] According to the embodiment of the present application, another device embodiment for implementing the above image processing method is further provided. Figure 12 is a schematic structural diagram of another image processing device according to the embodiment of the present application, as Figure 12 shown, the device includes:
[0249] A fourth acquisition module 1201, configured to acquire a currently input image processing dialogue request, wherein request data carried in the image processing dialogue request includes: a two-dimensional image;
[0250] A second return module 1202, configured to return an image processing dialogue reply in response to the image processing dialogue request, wherein information carried in the image processing dialogue reply includes: three-dimensional scene data, and the three-dimensional scene data is generated according to the image processing method of any one of the above;
[0251] A display module 1203 for displaying three-dimensional scene data within a graphical user interface.
[0252] By adopting the embodiment of the present application, a request for an image processing dialogue currently input is obtained. Among them, the request data carried in the request for an image processing dialogue includes: a two-dimensional image. In response to the request for an image processing dialogue, an image processing dialogue reply is returned. The information carried in the image processing dialogue reply includes: three-dimensional scene data. The three-dimensional scene data is generated according to any one of the above-mentioned image processing methods. The three-dimensional scene data is displayed within a graphical user interface, thereby achieving the purpose of efficiently and low-costly generating three-dimensional scene data based on a two-dimensional image, thus realizing the low-cost acquisition of a large-scale and high-quality three-dimensional scene data. Moreover, the generated three-dimensional scene data has both real-world scale and clear details, significantly promoting the development of spatial intelligence technology, and further solving the technical problems in the related art that the acquisition cost of three-dimensional scene data sets is high, the scene diversity is insufficient, and the scale is small.
[0253] It should be noted here that the above-mentioned fourth acquisition module 1201, second return module 1202, and display module 1203 correspond to steps S61 to S63 in the embodiment. The examples and application scenarios implemented by the three modules and the corresponding steps are the same, but are not limited to the content disclosed in the above-mentioned embodiment. It should be noted that the above-mentioned module or unit may be a hardware component or a software component stored in a memory and processed by one or more processors. The above-mentioned module may also run in the server 10 provided in the embodiment.
[0254] According to the embodiment of the present application, another device embodiment for implementing the above-mentioned image processing method is further provided. Figure 13 It is a schematic structural diagram of another image processing device according to the embodiment of the present application, as Figure 13 shown. The device includes:
[0255] A first display module 1301 for responding to an input instruction acting on an operation interface and displaying a two-dimensional image on the operation interface;
[0256] A second display module 1302 for responding to a processing instruction acting on the operation interface and displaying three-dimensional scene data on the operation interface; among them, the three-dimensional scene data is generated according to any one of the above-mentioned image processing methods.
[0257] By adopting the embodiment of the present application, a two-dimensional image is displayed on the operation interface by responding to an input instruction acting on the operation interface, and three-dimensional scene data is displayed on the operation interface by responding to a processing instruction acting on the operation interface; wherein, the three-dimensional scene data is generated according to the image processing method of any one of the above, thereby achieving the purpose of efficiently and low-cost generating three-dimensional scene data based on a two-dimensional image, thus realizing low-cost acquisition of large-scale and high-quality three-dimensional scene data, and the generated three-dimensional scene data has both real-world scale and clear details, significantly promoting the development of spatial intelligence technology, and further solving the technical problems of high acquisition cost, insufficient scene diversity and small scale of the three-dimensional scene dataset in the related art.
[0258] It should be noted here that the above first display module 1301 and second display module 1302 correspond to step 71 and step S72 in the embodiment. The instances and application scenarios realized by the two modules and the corresponding steps are the same, but are not limited to the content disclosed in the above embodiment. It should be noted that the above module or unit can be a hardware component or a software component stored in the memory and processed by one or more processors, and the above module can also run in the server 10 provided in the embodiment.
[0259] It should be noted that the preferred implementation schemes involved in the above embodiments of the present application are the same as the schemes, application scenarios and implementation processes provided in the embodiments, but are not limited to the schemes provided in the embodiments.
[0260] The embodiment of the present application can provide a computing device. Figure 14 It is a structural block diagram of a computing device according to an embodiment of the present application. As Figure 14 shown, the computing device A may include: one or more ( Figure 14 only one is shown in the figure) processors 1402, a memory 1404, a storage controller, and a peripheral interface. The peripheral interface may be connected to a radio frequency module, an audio module, a display screen, etc., which are not limited here.
[0261] The above computing device A can be understood as an integrated intelligent terminal, including but not limited to a server, a desktop computer, a PC (Personal Computer), a model all-in-one machine, etc. And the model described in the above embodiments of the present application may be pre-installed in the computing device.
[0262] Specifically, the computing device A can pre-set various types of models, including but not limited to models in the fields of natural language processing, visual processing, speech processing, code processing, multi-modal task processing, etc., so as to provide diverse model selections. In different product forms, the computing device A can support one or more model usage methods, including but not limited to model training, model invocation, model fine-tuning, model deployment, model inference and application, etc. In some product forms, the computing device A also supports model management, including but not limited to multi-type model management (supporting the management of various types of models such as discriminative and generative models), model version control (supporting the control of different model versions), model evaluation (evaluating the performance and effect of the model based on model evaluation tools), etc. In other product forms, the computing device A can also create applications based on the model, provide API invocation capabilities, and can call the model into the created application through the API interface, and at the same time provide application management tools to realize the management and monitoring of the application.
[0263] Furthermore, the computing device A can also include data management (supporting the creation and management of model tuning data sets), a training center (providing rich training resources to help users learn and master AI technologies), and basic control capabilities (providing enterprise-level basic control capabilities to ensure the security and efficient operation of the system). Through the above functions, a comprehensive and integrated AI development, training, deployment, and application device is provided.
[0264] Among them, the memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the methods and devices in the embodiments of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, implements the methods in the above embodiments. The memory can include high-speed random access memory, and can also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory can further include a memory remotely set relative to the processor, and these remote memories can be connected to the terminal through a network. Examples of the above network include but are not limited to the Internet, enterprise intranet, local area network, mobile communication network, and their combinations.
[0265] The processor can call the executable program stored in the memory through the transmission device to execute the method described in any one of the above embodiments.
[0266] Those of ordinary skill in the art can understand that the structure shown is only schematic, and the computing device A can also be a terminal device such as a smart phone, a tablet computer, a handheld computer, and a Mobile Internet Device (MID), a PAD, etc. The Figure 14 shown structure is only schematic, and the computing device A can also be a terminal device such as a smart phone, a tablet computer, a handheld computer, and a Mobile Internet Device (MID), a PAD, etc. The Figure 14It does not limit the structure of the above computing device. For example, computing device A may also include more or fewer components (such as network interfaces, display devices, etc.) than those shown in the Figure 14 and may have a configuration different from that shown in the Figure 14 .
[0267] Embodiments of the present application may provide an electronic device. Figure 15 is a structural block diagram of an electronic device according to an embodiment of the present application. As Figure 15 shown, the electronic device may include: an input / output device 152; a memory 154, and a processor 156, where the processor 156 is connected to the input / output device 152 and the memory 154 through a bus 158.
[0268] Among them, the memory can be used to store software programs and modules, such as program instructions / modules corresponding to the methods and devices in the embodiments of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, implements the methods in the above embodiments. The memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory may further include a memory remotely disposed relative to the processor, and these remote memories may be connected to the terminal through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0269] The processor may call the executable program stored in the memory through a transmission device to execute the method described in any one of the above embodiments.
[0270] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the relevant hardware of the terminal device through a program, and the program can be stored in a computer-readable storage medium. The storage medium may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc.
[0271] Embodiments of the present application also provide a computer-readable storage medium. Optionally, in this embodiment, the above computer-readable storage medium may be used to store the program code executed by the image processing method provided in the above embodiment.
[0272] Optionally, in this embodiment, the above computer-readable storage medium may be located in any one of the computer terminals in a computer terminal group in a computer network, or in any one of the mobile terminals in a mobile terminal group.
[0273] An embodiment of the present application also provides a computer program product, which includes a computer program that, when executed by a processor, implements any one of the above-mentioned image processing methods.
[0274] In the above embodiments of the present application, the descriptions of the respective embodiments have their own emphases. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0275] In several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the units or modules can be in an electrical or other form.
[0276] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0277] In addition, the functional units in each embodiment of the present application can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0278] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs that can store program codes.
[0279] The above are only the preferred embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.
Claims
1. An image processing method, characterized in that, Including: Obtain a two-dimensional image; Perform scene depth estimation on the two-dimensional image using a multiple depth estimation strategy to obtain a target depth map, where the multiple depth estimation strategy is used to obtain scene depth information of a given image through different types of depth estimation methods; Generate three-dimensional scene data based on the camera parameters of the two-dimensional image and the target depth map, where the three-dimensional scene data is training data applied to the three-dimensional perception and understanding scene.
2. The image processing method according to claim 1, wherein Performing scene depth estimation on the two-dimensional image using the multiple depth estimation strategy to obtain the target depth map includes: Perform relative depth estimation on the two-dimensional image using a relative depth estimation model to obtain a relative depth map; Perform metric depth estimation on the two-dimensional image using a metric depth estimation model to obtain a metric depth map; Fuse the relative depth map and the metric depth map to obtain the target depth map.
3. The image processing method according to claim 2, wherein, Fusing the relative depth map and the metric depth map to obtain the target depth map includes: Determine a scaling factor based on the relative depth map and the metric depth map, where the relative depth map and the metric depth map have the same dimension; Perform scale calibration on the relative depth map based on the scaling factor to obtain the target depth map.
4. The image processing method according to claim 3, wherein Determining the scaling factor based on the relative depth map and the metric depth map includes: Identify a set of valid points from the two-dimensional image; Calculate the average effective relative depth of the set of valid points based on the relative depth map, and calculate the average effective metric depth of the set of valid points based on the metric depth map; Calculate the scaling factor using the average effective relative depth and the average effective metric depth.
5. The image processing method according to claim 1, characterized in that, The camera parameters include: the camera internal parameters of the two-dimensional image and the camera external parameters of the two-dimensional image. The image processing method further includes: Perform internal parameter estimation on the two-dimensional image using a camera internal parameter estimation model to obtain the camera internal parameters, and perform external parameter estimation on the two-dimensional image using a field of view method to obtain the camera external parameters.
6. The image processing method according to claim 5, wherein Generating the three-dimensional scene data based on the camera parameters of the two-dimensional image and the target depth map includes: Generate an initial three-dimensional point cloud representation corresponding to the two-dimensional image based on the camera internal parameters, the camera external parameters, and the target depth map; Perform post-processing on the initial three-dimensional point cloud representation to obtain a target three-dimensional point cloud representation, where the post-processing includes at least some of the following processes: noise filtering processing, noise smoothing processing, missing area filling processing.
7. The image processing method according to claim 6, wherein Generating the initial three-dimensional point cloud representation corresponding to the two-dimensional image based on the camera internal parameters, the camera external parameters, and the target depth map includes: According to the camera internal parameters and the target depth map, project the pixel coordinates of the two-dimensional image into the camera coordinate system to generate a three-dimensional point set in the camera coordinate system; Convert the three-dimensional point set in the camera coordinate system to the initial three-dimensional point cloud representation in the world coordinate system based on the camera external parameters.
8. The image processing method according to claim 6, wherein The image processing method further includes: Determine the maximum depth value and the minimum depth value of the target three-dimensional point cloud representation; Construct a three-dimensional bounding box based on the maximum depth value and the minimum depth value; Convert the three-dimensional bounding box into three-dimensional annotation information based on the two-dimensional annotation information of the two-dimensional image.
9. An image processing method, characterized in that, Including: Obtain a two-dimensional image; Perform scene depth estimation on the two-dimensional image using a multiple depth estimation strategy to obtain a target depth map, where the multiple depth estimation strategy is used to obtain scene depth information of a given image through different types of depth estimation methods; Generate three-dimensional scene data based on the camera parameters of the two-dimensional image and the target depth map; Train an initial three-dimensional multimodal language model based on the three-dimensional scene data to obtain a target three-dimensional multimodal language model, where the target three-dimensional multimodal language model is used to process at least one of the following tasks: three-dimensional vision tasks, spatial reasoning tasks, cross-modal tasks.
10. An image processing method, characterized in that, Including: Obtain an image processing request through a first application programming interface, where the request data carried in the image processing request includes: a two-dimensional image; Return an image processing response through a second application programming interface, where the response data carried in the image processing response includes: three-dimensional scene data, and the three-dimensional scene data is generated according to the image processing method described in any one of claims 1 to 8.
11. An image processing method, characterized in that, Including: Obtain a current input image processing dialogue request, where the request data carried in the image processing dialogue request includes: a two-dimensional image; In response to the image processing dialogue request, return an image processing dialogue reply, where the information carried in the image processing dialogue reply includes: three-dimensional scene data, and the three-dimensional scene data is generated according to the image processing method described in any one of claims 1 to 8; Display the three-dimensional scene data within a graphical user interface.
12. An image processing method, characterized in that, Including: In response to an input instruction acting on an operation interface, display a two-dimensional image on the operation interface; In response to a processing instruction acting on the operation interface, display three-dimensional scene data on the operation interface; Wherein, the three-dimensional scene data is generated according to the image processing method described in any one of claims 1 to 8.
13. An image processing system, characterized in that, Including: A client for sending a two-dimensional image; A server connected to the client, for performing scene depth estimation on the two-dimensional image using a multiple depth estimation strategy to obtain a target depth map, and generating three-dimensional scene data based on the camera parameters of the two-dimensional image and the target depth map, where the multiple depth estimation strategy is used to obtain scene depth information of a given image through different types of depth estimation methods, and the three-dimensional scene data is training data applied to a three-dimensional perception and understanding scene; The client is further configured to output the three-dimensional scene data.
14. An electronic device, characterized in that, Including: A memory storing an executable program; A processor for running the program, where when the program runs, it executes the image processing method described in any one of claims 1 to 8.
15. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored executable program, where when the executable program runs, it controls the device where the computer-readable storage medium is located to execute the image processing method described in any one of claims 1 to 8.
16. A computer program product, characterized in that, Comprising a computer program which, when executed by a processor, implements the image processing method according to any one of claims 1 to 8.
Citation Information
Cited By
Single-image measurement scale human body scene collaborative reconstruction method and system
CN121544813A
Luggage case visual stacking size calculation method based on depth map
CN122156291A
Laser radar point cloud abnormal data processing method and device
CN122222870A