Three-dimensional real scene establishing method and device

By combining radar point cloud and image data to generate accurate depth maps and perform texture mapping, the problems of low efficiency and poor adaptability in the prior art are solved, and high-precision and real-time three-dimensional reconstruction effects are achieved.

CN120339523APending Publication Date: 2025-07-18CHINA COAL RES INST +1
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510796771.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-16
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the prior art, camera-based three-dimensional reconstruction is inefficient in dramatic light changes and weak texture scenarios. Lidar-based three-dimensional reconstruction fails in weak structure scenarios, making it difficult to achieve efficient real-time three-dimensional reconstruction.

Method used

Combining radar point cloud images and acquired images to generate keyframe data, rough depth maps are generated through semi-global matching, and accurate depth maps are generated using depth filtering and optimization neural networks, and combining texture mapping to build a target real-life three-dimensional model.

Benefits of technology

It improves the accuracy and robustness of three-dimensional reconstruction, enhances adaptability in large-scale scenarios, and realizes online real-time three-dimensional reconstruction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339523A_ABST
    Figure CN120339523A_ABST
Patent Text Reader

Abstract

The invention provides a three-dimensional real scene establishment method and device, and relates to the technical field of image processing, and the method comprises the steps: obtaining a point cloud image and a collection image, and the point cloud image and the collection image are collected in a to-be-constructed target; generating key frame data of the to-be-constructed target based on the point cloud image and the collected image; constructing a rough depth map based on the key frame data; performing optimization processing on the rough depth map to generate an accurate depth map; and constructing a candidate live-action three-dimensional model of the to-be-constructed target based on the accurate depth map, and performing texture mapping on the candidate live-action three-dimensional model based on the acquired image to generate a target live-action three-dimensional model. Therefore, the point cloud provides high-precision geometric information, the image provides rich texture details, the fusion of the point cloud and the image can make up for the deficiency of a single sensor, the three-dimensional reconstruction precision is improved, the robustness and adaptability are enhanced, and the combination of the radar and the image can be suitable for large-range scene modeling.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of image processing technologies, and in particular, to a method and apparatus for establishing a three-dimensional real scene. Background Art

[0002] The essence of three-dimensional reconstruction is to use the information of sensors to restore the geometric structure of the environment and calculate the spatial pose of the sensors, including four parts: Structure from motion (SFM), Multiview stereo (MVS), Surface reconstruction (SR), and Texture Mapping (TM). Three-dimensional real scene reconstruction refers to inputting the original data acquired by sensors and outputting a three-dimensional model with textures. Through decades of development, three-dimensional reconstruction has made great progress and has been widely applied in many fields, such as unmanned driving, augmented reality and virtual reality, robot positioning and navigation, three-dimensional digital protection of ancient buildings, and digital twins and digital factories.

[0003] Sensors in current technologies can all be used for three-dimensional reconstruction. Typical sensors include cameras and lidars. Cameras acquire two-dimensional image information. According to the theoretical knowledge related to multi-view geometry, the technical route for three-dimensional reconstruction using cameras has been very mature, but there are still many challenging problems, such as drastic changes in lighting and weak texture scenes. In addition, vision-based three-dimensional real scene reconstruction is often offline and has low efficiency. Lidars can directly acquire three-dimensional information and achieve online real-time three-dimensional reconstruction by matching and registering laser point clouds. However, in weakly structured scenes (such as long straight corridors, tunnels, open areas, etc. with simple and repetitive structures), lidar-based three-dimensional reconstruction often fails. Summary of the Invention

[0004] The present disclosure aims to at least partly solve one of the technical problems in the related technologies.

[0005] To this end, an object of the present disclosure is to propose a method for establishing a three-dimensional real scene.

[0006] A second object of the present disclosure is to propose an apparatus for establishing a three-dimensional real scene.

[0007] A third object of the present disclosure is to propose an electronic device.

[0008] A fourth object of the present disclosure is to propose a non-transitory computer-readable storage medium.

[0009] A fifth object of the present disclosure is to propose a computer program product.

[0010] To achieve the above object, an embodiment of the first aspect of the present disclosure provides a method for establishing a three-dimensional real scene, including: acquiring a point cloud image and a captured image, where the point cloud image and the captured image are captured in a target to be constructed; generating key frame data of the target to be constructed based on the point cloud image and the captured image; constructing a rough depth map based on the key frame data; performing optimization processing on the rough depth map to generate an accurate depth map; constructing a candidate real scene three-dimensional model of the target to be constructed based on the accurate depth map, and performing texture mapping on the candidate real scene three-dimensional model based on the captured image to generate a target real scene three-dimensional model.

[0011] According to an embodiment of the present disclosure, the constructing a rough depth map based on the key frame data includes: processing the key frame data based on a semi-global matching method to generate the rough depth map.

[0012] According to an embodiment of the present disclosure, the performing optimization processing on the rough depth map to generate an accurate depth map includes: performing filtering processing on the rough depth map; inputting the filtered rough depth map into an optimization neural network to generate the accurate depth map.

[0013] According to an embodiment of the present disclosure, before generating the key frame data of the target to be constructed based on the point cloud image and the captured image, it further includes: determining a radar key frame of the target to be constructed based on the point cloud image, and determining an image key frame of the target to be constructed based on the captured image; performing loop closure detection based on the radar key frame and the image key frame; optimizing the point cloud image and the captured image based on the loop closure detection result.

[0014] According to an embodiment of the present disclosure, the performing loop closure detection based on the radar key frame and the image key frame includes: establishing a first map based on the radar key frame, and establishing a second map based on the image key frame; fusing the first map and the second map to generate a three-dimensional space map; performing loop closure detection on the three-dimensional space map.

[0015] According to an embodiment of the present disclosure, the performing filtering processing on the rough depth map includes: acquiring a confidence value of each depth value in the rough depth map; screening the depth values based on the confidence value.

[0016] According to an embodiment of the present disclosure, the method further includes: complementing the filtered depth values based on the semantic information of the image.

[0017] To achieve the above object, an embodiment of the second aspect of the present disclosure provides a three-dimensional real scene establishment device, including: an acquisition module, configured to acquire a point cloud image and a captured image, where the point cloud image and the captured image are captured in a target to be constructed; a construction module, configured to generate key frame data of the target to be constructed based on the point cloud image and the captured image; a generation module, configured to construct a rough depth map based on the key frame data; an optimization module, configured to perform optimization processing on the rough depth map to generate an accurate depth map; an establishment module, configured to construct a candidate real scene three-dimensional model of the target to be constructed based on the accurate depth map, and perform texture mapping on the candidate real scene three-dimensional model based on the captured image to generate a target real scene three-dimensional model.

[0018] To achieve the above object, an embodiment of the third aspect of the present disclosure provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to implement the three-dimensional real scene establishment method as described in the embodiment of the first aspect of the present disclosure.

[0019] To achieve the above object, an embodiment of the fourth aspect of the present disclosure provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to implement the three-dimensional real scene establishment method as described in the embodiment of the first aspect of the present disclosure.

[0020] To achieve the above object, an embodiment of the fifth aspect of the present disclosure provides a computer program product, including a computer program, where the computer program is used to implement the three-dimensional real scene establishment method as described in the embodiment of the first aspect of the present disclosure when executed by a processor. Description of the Drawings

[0021] Figure 1 is a schematic diagram of a three-dimensional real scene establishment method according to an embodiment of the present disclosure; Figure 2 is a schematic diagram of generating a rough depth map based on key frame data through Semi-Global Matching (SGM) according to an embodiment of the present disclosure; Figure 3 is a schematic diagram of another three-dimensional real scene establishment method according to an embodiment of the present disclosure; Figure 4 is a schematic diagram of another three-dimensional real scene establishment method according to an embodiment of the present disclosure; Figure 5 is a schematic diagram of a three-dimensional real scene establishment device according to an embodiment of the present disclosure; Figure 6 is a schematic diagram of an electronic device according to an embodiment of the present disclosure. Detailed implementation manners

[0022] The embodiments of the present disclosure will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the present disclosure, and should not be construed as a limitation of the present disclosure.

[0023] In the technical solution of the present disclosure, the acquisition, storage, use, processing, etc. of data all comply with the relevant regulations of relevant laws and regulations.

[0024] It should be noted that in the embodiments of the present application, some industry-existing solutions such as certain software, components, models, etc. may be mentioned. They should be regarded as exemplary. The purpose is only to illustrate the feasibility in the implementation of the technical solution of the present application, but it does not mean that the applicant has already or necessarily used this solution.

[0025] Figure 1 is a schematic diagram of a three-dimensional real scene establishment method according to an embodiment of the present disclosure. As Figure 1 shown, the three-dimensional real scene establishment method includes the following steps: S101. Obtain a point cloud image and a captured image, where the point cloud image and the captured image are captured in the target to be constructed.

[0026] The three-dimensional real scene establishment method of the embodiments of the present application can be applied to the scenario of establishing a three-dimensional mine map. The execution subject of the three-dimensional real scene establishment of the embodiments of the present application can be the three-dimensional real scene establishment device of the embodiments of the present application, and this three-dimensional real scene establishment device can be set on an electronic device.

[0027] It should be noted that the target to be constructed is an object for which a real scene three-dimensional model needs to be constructed currently. The target to be constructed can be various, and no limitation is made here. For example, the target to be constructed can be an open scene such as a mine or a factory, or a closed scene such as a roadway or a tunnel, etc.

[0028] The point cloud image can be acquired by a radar, and the captured image can be acquired by an image acquisition device.

[0029] S102. Generate key frame data of the target to be constructed based on the point cloud image and the captured image.

[0030] A key frame is a representative frame selected from a video sequence or an image stream. It usually contains rich visual information, and there is a large motion change between frames, which is suitable for feature matching and pose estimation.

[0031] In the embodiments of the present disclosure, there can be various methods for generating key frame data of the target to be constructed based on the point cloud image and the captured image, and no specific limitation is made here.

[0032] In a possible implementation manner, the point cloud image and the captured image can be processed to extract the feature points of the point cloud image and the captured image, and then the feature points of the point cloud image and the captured image are compared to determine whether there are overlapping feature points, and the key frame data is generated based on the overlapping feature points.

[0033] In another possible implementation manner, the point cloud image and the captured image can also be processed according to a preset key frame data generation model to determine the key frame data of the target to be constructed. The key frame data generation model is pre-trained and can be stored in the storage space of the electronic device for convenient retrieval and use when needed.

[0034] S103. Construct a coarse depth map based on the key frame data.

[0035] A depth map is an image that represents the distance information of each point on the surface of a scene or an object relative to a certain viewing angle. In a depth map, each pixel value represents the distance between the pixel point and the camera, that is, the depth value. Such images are usually used in fields such as computer vision, 3D reconstruction, robot navigation, and augmented reality.

[0036] A coarse depth map refers to a depth map generated in image processing and 3D vision tasks that has not been refined. It is usually used as the basis for subsequent optimization or refinement steps and can also be directly used in scenarios with low accuracy requirements.

[0037] In the embodiments of the present disclosure, after obtaining the key frame data, in order to accurately and quickly generate the depth map of the visual key frame, by combining the advantages of traditional geometric methods and the advantages of deep learning methods, the present disclosure proposes a coarse-to-fine depth map generation method.

[0038] In the embodiments of the present disclosure, a coarse depth map can be generated based on the key frame data through semi-global matching (SGM). Semi-global matching (SGM) is an effective method for calculating the disparity between images to achieve stereo vision matching. It approximates global optimization by performing dynamic programming in multiple directions, thereby achieving a better matching effect while maintaining a low computational complexity.

[0039] For example, it can be implemented through the following steps, as Figure 2 shown, the method includes: Step 1: For each pixel , determine search paths , and the search disparity range ; Step 2: Calculate the matching cost of the pixel and each disparity on each path ; Step 3: Through cost aggregation, for each disparity of each path of the pixel , aggregate the costs of neighboring pixels and adjacent disparities along the path to the current pixel-disparity pair to obtain the aggregated cost ; Step 4: Execute cost aggregation multiple times to iteratively aggregate the cost; Step 5: Aggregate the costs on each path for each pixel-disparity to obtain the cost of each pixel for each disparity ; Step 6: Find the optimal cost of different disparities for each pixel , determine the disparity of each pixel to obtain the depth of the pixel, and obtain the rough depth image of the image.

[0040] In implementation, calculating the matching cost of the pixel and each disparity on each path in Step 2 can be implemented through binary descriptor matching. Binary Descriptor Matching is an efficient method for feature point matching in computer vision and is widely used in tasks such as real-time image matching, SLAM, 3D reconstruction, and object recognition. Compared with traditional floating-point descriptors (such as SIFT, SURF), binary descriptors have the advantages of fast calculation speed, small memory footprint, and suitability for hardware acceleration.

[0041] The formula for cost aggregation based on the binary descriptor method is:

[0042] In the formula, p is the pixel, d is the disparity, r is the unit vector along a path direction, P1 is the penalty term for the disparity with a difference of 1 from the current disparity during cost aggregation, P2 is the penalty term for the disparity with a difference greater than or equal to 2 from the current disparity during cost aggregation, represents the currently optimized path; It represents the cost aggregation value when the disparity between adjacent points on this path is d. It means subtracting the minimum cost among different disparities to prevent cost overflow due to excessive values. It means taking the minimum cost among disparities greater than or equal to 2, and then adding the penalty term P2 to this cost.

[0043] The first term of the formula is the original cost of this pixel. .

[0044] , the second term is the aggregation value of the adjacent points on this path with disparities of (d), (d - 1), the aggregation value of the adjacent points on this path with a disparity of (d + 1), the aggregation value with the minimum cost, and the minimum of these four numbers.

[0045] , the third term is the minimum disparity value of the adjacent points on this path, subtracting this value to prevent the cost aggregation value from increasing continuously because it keeps adding.

[0046]

[0047] The final cost aggregation value (S(p, d)) is the sum of the aggregation values of all paths.

[0048] S104, optimize the rough depth map to generate an accurate depth map.

[0049] In the embodiments of the present disclosure, due to errors in matching, the generated rough depth map has a large amount of noise. Therefore, a depth filter can be used to filter the rough depth map to remove depth values with large errors in the rough depth map. During the optimization of the depth network, CNN and residual networks are used for optimization to ensure both the accuracy of the depth map and the efficiency of generating the depth map. During the process of optimizing the depth map using the depth network, emphasis is placed on using the semantic information of the image to further denoise and complete the depth map, especially for scenes such as occlusion, weak texture / no texture, and specular reflection. By using the extracted semantic information to optimize the initially estimated depth map, the accuracy of the estimated depth map is improved.

[0050] S105, construct a candidate real - scene three - dimensional model of the target to be constructed based on the accurate depth map, and perform texture mapping on the candidate real - scene three - dimensional model based on the acquired images to generate the target real - scene three - dimensional model.

[0051] It should be noted that the process of texture mapping is to extract color information, image information, etc. from the captured image, that is, the texture. Specifically, in the process of 3D reconstruction, texture mapping involves accurately pasting the pixel information on the two-dimensional images from multiple perspectives (these images contain the color and texture details of the scene or object surface) onto the surface of the candidate real scene 3D model generated by meshing the dense model.

[0052] In a possible implementation manner, during the meshing process, the candidate real scene 3D model is meshed in an incremental manner, and the meshed candidate real scene 3D model is updated in an incremental update manner. After obtaining the meshed candidate real scene 3D model grid, in combination with the visual key frames and their poses, the optimal perspective is selected for texture mapping. The perspective selection process is the first step of the texture reconstruction algorithm, and the main purpose of this step is to select a suitable perspective for each patch to determine the color expression of the patch. Texture optimization refers to optimizing the seams between different patches. Intuitively, the seams are formed because the perspective images mapped by the left and right patches are different, and the imaging of the same position of the same object under different perspectives will show different colors due to factors such as light, so there will be a discontinuous effect when splicing the images from two perspectives, thus generating seams. Therefore, a very natural solution idea is to adjust the pixel colors corresponding to the left and right image blocks simultaneously to make them closer, so as to fade the discontinuity. In order to achieve the purpose of online real-time real scene 3D reconstruction, the image block method is used to accelerate texture optimization, that is, to efficiently solve the energy function in the texture optimization process, and finally perform texture fusion in real time to obtain the target real scene 3D model with texture.

[0053] In the embodiments of the present disclosure, first, a point cloud image and a captured image are obtained, where the point cloud image and the captured image are captured in the target to be constructed. Then, key frame data of the target to be constructed is generated based on the point cloud image and the captured image. Then, a rough depth map is constructed based on the key frame data. Then, the rough depth map is optimized to generate an accurate depth map. Finally, a candidate real scene 3D model of the target to be constructed is constructed based on the accurate depth map, and texture mapping is performed on the candidate real scene 3D model based on the captured image to generate the target real scene 3D model. Thus, the point cloud provides high-precision geometric information, and the image provides rich texture details. The fusion of the two can make up for the deficiencies of a single sensor, improve the 3D reconstruction accuracy, enhance the robustness and adaptability, and at the same time, the combination of radar and image can be applied to large-scale scene modeling.

[0054] In the above embodiments, to optimize the rough depth map to generate an accurate depth map, it can also be through Figure 3 For further explanation, the method includes: S301, perform filtering processing on the rough depth map.

[0055] Since the rough depth map usually contains a large number of outliers or measurement errors, filtering can effectively reduce these noise points and make the depth values closer to the true distances.

[0056] By filtering the rough depth map, the disparity jump regions of the rough depth map can also be optimized to smooth the transition regions while preserving the edges.

[0057] In a possible implementation, the confidence values of the depth values in the rough depth map can be obtained, and then the depth values can be screened based on the confidence values.

[0058] Due to errors in the matching, the generated Coarse depth map has a large amount of noise. Therefore, a confidence-based depth filter is used to filter the Coarse depth map to remove the depth values with large errors in the Coarse depth map. During the optimization of the depth network, CNN and residual networks are used for optimization to ensure both the accuracy of the depth map and the efficiency of generating the depth map. During the process of optimizing the depth map using the depth network, emphasis is placed on using the semantic information of the image to further denoise and complete the depth map, especially for scenes such as occlusion, weak texture / no texture, and specular reflection. By using the extracted semantic information to optimize the initially estimated depth map, the accuracy of the estimated depth map is improved.

[0059] In another possible implementation, during the process of optimizing the depth map using the depth network, emphasis is placed on using the semantic information of the image to further denoise and complete the depth map, especially for scenes such as occlusion, weak texture / no texture, and specular reflection. By using the extracted semantic information to optimize the initially estimated depth map, the accuracy of the estimated depth map is improved.

[0060] S302, Input the filtered rough depth map into the optimization neural network to generate an accurate depth map.

[0061] In the embodiments of the present disclosure, first, the rough depth map is filtered, and then the filtered rough depth map is input into the optimization neural network to generate an accurate depth map. Thus, this method combining filtering and neural network optimization can make full use of the advantages of both, providing an efficient and accurate solution to generate high-quality depth maps.

[0062] In the above embodiments, before generating the key frame data of the target to be constructed based on the point cloud image and the captured image, the following operations may also be required, such as Figure 4 As shown, including: S401, Determine the radar key frame of the target to be constructed based on the point cloud image, and determine the image key frame of the target to be constructed based on the captured image.

[0063] In the embodiments of the present disclosure, the radar key frames in the point cloud image and the image key frames in the acquired image can be determined by tracking key points.

[0064] S402. Perform loop closure detection based on the radar key frames and the image key frames.

[0065] In the embodiments of the present disclosure, a first map can be established based on the radar key frames, and a second map can be established based on the image key frames. Then, the first map and the second map are fused to generate a three-dimensional space map. Finally, loop closure detection is performed on the three-dimensional space map.

[0066] In a possible implementation manner, the map reconstructed by the lidar and the map reconstructed by the vision can be fused to obtain a complete three-dimensional space map. If loop closure information is detected, the entire map is globally optimized to suppress error drift. Otherwise, the entire map is locally optimized. During the loop closure detection process, since the key frames contain both image information and lidar point cloud information, both the image information and the lidar point cloud information are used simultaneously for loop closure detection to improve the accuracy of loop closure detection.

[0067] For visual information, the bag-of-words model can be used to calculate the bag-of-words vector of the image, and whether there is a visual loop is judged according to the similarity of the bag-of-words vectors; for lidar information, the rotation and translation distances between lidar frames are used, and whether there is a lidar loop is judged according to the distance. If both vision and lidar detect a loop, it is considered that the loop is correct, otherwise it is incorrect.

[0068] S403. Optimize the point cloud image and the acquired image based on the loop closure detection result.

[0069] In the embodiments of the present disclosure, since the lidar point cloud information is accurate, only the poses of the key frames are optimized during the optimization process of the sliding window, and the point cloud information is not optimized. The optimization function is as follows:

[0070] where k represents the number of key frames in the sliding window, and represent the poses of the key frames in the sliding window. rᵢ(2D→2D): represents the reprojection error from two dimensions to two dimensions, which may be used for image feature point matching or optical flow calculation. rᵢ(3D→2D): represents the reprojection error from three dimensions to two dimensions, which is commonly used to project three-dimensional points onto the image plane and compare them with the observed values. rᵢ(3D→3D): represents the error in three-dimensional space, which may be related to three-dimensional point cloud registration or structure alignment.

[0071] Corresponding to the three-dimensional real scene establishment methods provided in the above several embodiments, an embodiment of the present disclosure also provides a three-dimensional real scene establishment device. Since the three-dimensional real scene establishment device provided in the embodiments of the present disclosure corresponds to the three-dimensional real scene establishment methods provided in the above several embodiments, the implementation manners of the above three-dimensional real scene establishment methods are also applicable to the three-dimensional real scene establishment device provided in the embodiments of the present disclosure, and will not be described in detail in the following embodiments.

[0072] Figure 5 FIG. 4 is a schematic diagram of a three-dimensional real scene establishment device according to an embodiment of the present disclosure. As shown in FIG. 4, the three-dimensional real scene establishment device 500 includes: An acquisition module 510, configured to acquire a point cloud image and a captured image, where the point cloud image and the captured image are captured in a target to be constructed.

[0073] A construction module 520, configured to generate key frame data of the target to be constructed based on the point cloud image and the captured image.

[0074] A generation module 530, configured to construct a rough depth map based on the key frame data.

[0075] An optimization module 540, configured to perform optimization processing on the rough depth map to generate an accurate depth map.

[0076] An establishment module 550, configured to construct a candidate real scene three-dimensional model of the target to be constructed based on the accurate depth map, and perform texture mapping on the candidate real scene three-dimensional model based on the captured image to generate a target real scene three-dimensional model.

[0077] According to an embodiment of the present disclosure, constructing a rough depth map based on the key frame data includes: processing the key frame data based on a semi-global matching method to generate a rough depth map.

[0078] According to an embodiment of the present disclosure, performing optimization processing on the rough depth map to generate an accurate depth map includes: performing filtering processing on the rough depth map; inputting the filtered rough depth map into an optimization neural network to generate an accurate depth map.

[0079] According to an embodiment of the present disclosure, before generating the key frame data of the target to be constructed based on the point cloud image and the captured image, it further includes: determining a radar key frame of the target to be constructed based on the point cloud image, so as to determine an image key frame of the target to be constructed based on the captured image; performing loop detection based on the radar key frame and the image key frame; optimizing the point cloud image and the captured image based on the loop detection result.

[0080] According to an embodiment of the present disclosure, closed-loop detection based on radar key frames and image key frames includes: establishing a first map based on radar key frames and establishing a second map based on image key frames; fusing the first map and the second map to generate a three-dimensional space map; and performing closed-loop detection on the three-dimensional space map.

[0081] According to an embodiment of the present disclosure, filtering the rough depth map includes: obtaining confidence values of each depth value in the rough depth map; and screening the depth values based on the confidence values.

[0082] According to an embodiment of the present disclosure, the method further includes: completing the removed depth values based on the semantic information of the image.

[0083] Thus, the point cloud provides high-precision geometric information, and the image provides rich texture details. The fusion of the two can make up for the deficiencies of a single sensor, improve the three-dimensional reconstruction accuracy, enhance the robustness and adaptability. At the same time, the combination of radar and image can be applied to large-scale scene modeling.

[0084] To implement the above embodiments, an electronic device 600 is further proposed in the embodiments of the present disclosure. Figure 6 is a schematic diagram of an electronic device according to an embodiment of the present disclosure, as Figure 6 shown. The electronic device 600 includes: a processor 601 and a memory 602 communicatively connected to the processor. The memory 602 stores instructions executable by at least one processor. The instructions are executed by at least one processor 601 to implement the three-dimensional real-scene establishment method as in the embodiments of the present disclosure. Figures 1-4 embodiments.

[0085] To implement the above embodiments, a non-transitory computer-readable storage medium storing computer instructions is further proposed in the embodiments of the present disclosure, wherein the computer instructions are used to cause a computer to implement the three-dimensional real-scene establishment method as in the embodiments of the present disclosure. Figures 1-4 embodiments.

[0086] To implement the above embodiments, a computer program product is further proposed in the embodiments of the present disclosure, including a computer program, and the computer program realizes the three-dimensional real-scene establishment method as in the embodiments of the present disclosure when executed by a processor. Figures 1-4 embodiments.

[0087] It should be noted that personal information from users should be collected for legal and reasonable purposes and should not be shared or sold outside of such legal uses. Additionally, such collection / sharing should be carried out after obtaining the informed consent of the user, including but not limited to notifying the user to read the user agreement / user notice and sign an agreement / authorization including authorizing relevant user information before the user uses the function. Moreover, any necessary steps should be taken to safeguard and protect access to such personal information data and ensure that others with access to the personal information data comply with their privacy policies and procedures.

[0088] This application is expected to provide embodiments where users can selectively block the use or access to personal information data. That is, this disclosure is expected to provide hardware and / or software to prevent or block access to such personal information data. Once the personal information data is no longer needed, the risk can be minimized by restricting data collection and deleting the data. Additionally, when applicable, personal identifiers are removed from such personal information to protect the privacy of the user.

[0089] In the descriptions of the foregoing embodiments, the descriptions referring to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of this application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. Additionally, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0090] Furthermore, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" can explicitly or implicitly include at least one of such features. In the description of this application, the meaning of "a plurality" is at least two, such as two, three, etc., unless otherwise specifically and clearly defined.

[0091] Any process or method description shown in the flowchart or described in other ways herein can be understood as representing a module, segment, or part of code including one or more executable instructions for implementing a customized logical function or process, and the scope of the preferred embodiments of this application includes additional implementations, where the functions can be executed in a manner that is not in the order shown or discussed, including in a substantially simultaneous manner according to the functions involved or in the reverse order, which should be understood by those skilled in the art to which the embodiments of this application pertain.

[0092] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in conjunction with these instruction execution systems, apparatuses, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that contains, stores, communicates, propagates, or transports a program for use by or in conjunction with an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection portion with one or more wirings (electronic device), a portable computer disk cartridge (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or otherwise processing as appropriate, and then storing it in a computer memory.

[0093] It should be understood that various parts of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), and the like.

[0094] Those of ordinary skill in the art of this technology can understand that all or part of the steps carried by the method of implementing the above embodiments can be completed by a program instructing relevant hardware, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.

[0095] In addition, each functional unit in various embodiments of the present application may be integrated into one processing module, may exist physically alone for each unit, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.

[0096] The above-mentioned storage medium may be a read-only memory, a magnetic disk, an optical disc, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.

Claims

1. A method for establishing a three-dimensional real scene, characterized in that, Including: Obtain a point cloud image and a captured image, where the point cloud image and the captured image are captured in the target to be constructed; Generate key frame data of the target to be constructed based on the point cloud image and the captured image; Construct a rough depth map based on the key frame data; Perform optimization processing on the rough depth map to generate an accurate depth map; Construct a candidate real scene three-dimensional model of the target to be constructed based on the accurate depth map, and perform texture mapping on the candidate real scene three-dimensional model based on the captured image to generate a target real scene three-dimensional model.

2. The method according to claim 1, wherein The constructing the rough depth map based on the key frame data includes: Process the key frame data based on the semi-global matching method to generate the rough depth map.

3. The method according to claim 1 or 2, characterized in that, The performing optimization processing on the rough depth map to generate an accurate depth map includes: Perform filtering processing on the rough depth map; Input the rough depth map after filtering processing into an optimization neural network to generate the accurate depth map.

4. The method according to claim 1, wherein Before generating the key frame data of the target to be constructed based on the point cloud image and the captured image, it further includes: Determine the radar key frame of the target to be constructed based on the point cloud image, and determine the image key frame of the target to be constructed based on the captured image; Perform loop closure detection based on the radar key frame and the image key frame; Optimize the point cloud image and the captured image based on the loop closure detection result.

5. The method according to claim 4, characterized in that The performing loop closure detection based on the radar key frame and the image key frame includes: Establish a first map based on the radar key frame, and establish a second map based on the image key frame; Fuse the first map and the second map to generate a three-dimensional space map; Perform loop closure detection on the three-dimensional space map.

6. The method according to claim 3, characterized in that, The performing filtering processing on the rough depth map includes: Obtain the confidence value of each depth value in the rough depth map; Screen the depth values based on the confidence value.

7. The method according to claim 6, wherein The method further includes: Complement the screened depth values based on the semantic information of the image.

8. A three-dimensional real scene establishment device, characterized in that, Including: An acquisition module, configured to acquire a point cloud image and a captured image, where the point cloud image and the captured image are captured in the target to be constructed; A construction module, configured to generate key frame data of the target to be constructed based on the point cloud image and the captured image; A generation module, configured to construct a rough depth map based on the key frame data; An optimization module, configured to perform optimization processing on the rough depth map to generate an accurate depth map; An establishment module, configured to construct a candidate real scene three-dimensional model of the target to be constructed based on the accurate depth map, and perform texture mapping on the candidate real scene three-dimensional model based on the captured image to generate a target real scene three-dimensional model.

9. An electronic device, characterized in that, Including a memory and a processor; Wherein, the processor runs a program corresponding to the executable program code by reading the executable program code stored in the memory, so as to implement the method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions that, when executed by a processor, are used to implement the method according to any one of claims 1-7.

Citation Information

Patent Citations

  • RGB-D image-based indoor scene three-dimensional reconstruction method

    CN109658449A

  • Method and device for constructing three-dimensional point cloud map by multi-machine cooperation and storage medium

    CN111951397A

  • Laser and image data fused three-dimensional reconstruction method and system

    CN112132972A

  • Method and apparatus for constructing three dimensional model of object

    US20170046868A1

  • Loop closure detection method and system, multi-sensor fusion slam system, robot, and medium

    US20230045796A1