Three-dimensional reconstruction method, device, equipment, medium and program product
By screening key frames and determining the method of supplementing key frames, the problem of balancing the quality and efficiency of 3D reconstruction in the existing technology is solved, and high-quality and efficient 3D reconstruction effects are achieved.
Patent Information
- Application Number
- CN202511149544.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-18
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2045-08-18
AI Technical Summary
It is difficult for existing technologies to achieve both high quality and high efficiency in three-dimensional reconstruction, and existing methods result in low reconstruction quality or low efficiency.
By screening key frames and determining supplementary key frames, a 3D reconstruction model is constructed, and non-key frames with a sufficiently large distance from the key frames in 3D space are screened out as supplementary key frames to assist in building the model.
It achieves both high quality and high efficiency in 3D reconstruction, avoids noise interference caused by image frame duplication, and improves the quality and efficiency of the reconstructed model.
Smart Images

Figure CN120726258A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and in particular to a three-dimensional reconstruction method, apparatus, device, medium and program product. Background Art
[0002] 3D reconstruction is a technique for recovering the 3D structure of a scene from environmental data collected by sensors. For example, 2D images captured by a camera at different locations or times can be used to reconstruct the 3D structure of a scene or objects within it. 3D reconstruction using images is also called visual 3D reconstruction.
[0003] After reconstructing a target object in three dimensions using certain visual 3D reconstruction methods, a 3D reconstructed model of the target object can be obtained. For example, a 2D image of the target object can be captured to obtain multiple image frames. Then, by resolving the target object in these image frames into three-dimensional space, a 3D reconstructed model of the target object can be obtained. However, the 3D reconstructed model reconstructed using existing methods differs significantly from the target object, resulting in low 3D reconstruction quality. Furthermore, when using other existing methods to improve reconstruction quality, additional processing of data from more sensor types is required, resulting in low 3D reconstruction efficiency.
[0004] Therefore, the existing technology has the problem of being difficult to achieve both high 3D reconstruction quality and high 3D reconstruction efficiency. Summary of the Invention
[0005] The embodiments of the present application provide a three-dimensional reconstruction method, apparatus, equipment, medium and program product for efficiently reconstructing a high-quality three-dimensional reconstruction model while taking into account both high three-dimensional reconstruction quality and high three-dimensional reconstruction efficiency when performing visual three-dimensional reconstruction of a target object.
[0006] In a first aspect, an embodiment of the present application provides a three-dimensional reconstruction method, the method comprising: obtaining multiple image frames of a target object to be three-dimensionally reconstructed; screening out key frames from the multiple image frames, the key frames being used to construct a three-dimensional reconstruction model; determining some non-key frames in each non-key frame in the multiple image frames as supplementary key frames, the first distance between any supplementary key frame and its closest key frame in the three-dimensional space of the three-dimensional reconstruction model being greater than a first preset distance, the supplementary key frames being used to assist in constructing the three-dimensional reconstruction model; performing three-dimensional reconstruction based on the supplementary key frames and the key frames to obtain a three-dimensional reconstruction model of the target object.
[0007] In a possible implementation, some non-key frames are determined as supplementary key frames, including: dividing multiple image frames into multiple spatial regions, where the multiple spatial regions are set according to the spatial position of the target object in the three-dimensional space; calculating the second distance between any non-key frame and any key frame in each of the multiple spatial regions; and determining the supplementary key frame in each non-key frame whose second distance is greater than the first preset distance.
[0008] In a possible embodiment, before dividing the multiple image frames into multiple spatial regions, the method also includes: constructing a hemispherical space surrounding the target object based on the point cloud data of the target object; dividing the hemispherical space into multiple hemispherical primitives, and the multiple hemispherical primitives correspond to the multiple spatial regions respectively.
[0009] In one possible implementation, a hemispherical space surrounding the target object is constructed based on the point cloud data of the target object, including: inputting the point cloud data of the target object into a three-dimensional target object detection algorithm module for processing to generate a three-dimensional rectangular frame surrounding the target object; and constructing a hemispherical space surrounding the target object with the center point of the bottom surface of the three-dimensional rectangular frame as the center of the sphere and the bottom surface as the center cross-section of the hemispherical space.
[0010] In one possible implementation, multiple image frames are divided into multiple spatial regions, including: converting the position information in the posture of each image frame into a three-dimensional space to obtain a position value of each image frame in the three-dimensional space; and dividing the image frames whose position values fall into corresponding spatial positions into corresponding spatial regions based on the position values of each image frame and the spatial positions of each spatial region in the three-dimensional space.
[0011] In a possible implementation, the method further includes: when multiple supplementary key frames are determined, constructing a spherical space for each supplementary key frame with a second preset distance as the sphere radius and a position value of each supplementary key frame as the sphere center position; for the spherical space of any supplementary key frame, deleting all non-key frames in the spherical space of any supplementary key frame except itself.
[0012] In a possible implementation, screening out key frames from multiple image frames includes: inputting multiple image frames into a simultaneous positioning and mapping algorithm module for image processing, and outputting processing results, where the processing results include the key frames screened out from the multiple image frames.
[0013] In a possible embodiment, the method also includes: establishing an association relationship between the supplementary key frames and the key frames, the association relationship characterizing the spatial correlation between the image frames in the three-dimensional space; performing three-dimensional reconstruction based on the supplementary key frames and the key frames to obtain a three-dimensional reconstruction model of the target object, including: performing three-dimensional reconstruction based on the association relationship, the supplementary key frames and the key frames to obtain a three-dimensional reconstruction model of the target object.
[0014] In a possible implementation, the multiple image frames are image frames obtained by capturing images by circling the target object with the target object as the center.
[0015] In one possible embodiment, after establishing an association relationship between the supplementary key frame and the key frame, the method further includes: performing image similarity calculation on any two image frames with an established association relationship to obtain an image similarity score; for any two image frames corresponding to an image similarity score greater than a similarity threshold, based on the spatial relationship between the two image frames in three-dimensional space, if the spatial relationship does not meet a preset spatial relationship condition, releasing the established association relationship between the two image frames.
[0016] In one possible embodiment, the spatial relationship is represented by the absolute value of the difference between the index values of the spatial regions corresponding to the two image frames, and the index value is used to represent the position sequence of the corresponding spatial region in multiple spatial regions. The spatial relationship condition includes that the spatial relationship is less than or equal to a preset index threshold.
[0017] In a second aspect, an embodiment of the present application provides a three-dimensional reconstruction device, which includes: an acquisition module for acquiring multiple image frames of a target object to be three-dimensionally reconstructed; a screening module for screening out key frames from the multiple image frames, the key frames being used to construct a three-dimensional reconstruction model; a determination module for determining some non-key frames in each non-key frame in the multiple image frames as supplementary key frames, the first distance between any supplementary key frame and the key frame closest to it in the three-dimensional space of the three-dimensional reconstruction model being greater than a first preset distance, the supplementary key frame being used to assist in constructing the three-dimensional reconstruction model; a reconstruction module for performing three-dimensional reconstruction based on the supplementary key frames and the key frames to obtain a three-dimensional reconstruction model of the target object.
[0018] In one possible embodiment, the determination module is specifically used to: divide multiple image frames into multiple spatial regions, where the multiple spatial regions are set according to the spatial position of the target object in the three-dimensional space; calculate the second distance between any non-key frame and any key frame in each spatial region of the multiple spatial regions; and determine a supplementary key frame in each non-key frame whose second distance is greater than the first preset distance.
[0019] In a possible embodiment, the device also includes a construction module, which is used to: construct a hemispherical space surrounding the target object based on the point cloud data of the target object; divide the hemispherical space into multiple hemispherical primitives, and the multiple hemispherical primitives correspond to multiple spatial regions respectively.
[0020] In one possible embodiment, the construction module is specifically used to: input the point cloud data of the target object into the three-dimensional target object detection algorithm module for processing, and generate a three-dimensional rectangular box surrounding the target object; use the center point of the bottom surface of the three-dimensional rectangular box as the center of the sphere, and use the bottom surface as the center section of the hemispherical space to construct a hemispherical space surrounding the target object.
[0021] In one possible implementation, the determination module is specifically used to: convert the position information in the posture of each image frame into three-dimensional space to obtain the position value of each image frame in the three-dimensional space; and divide the image frames whose position values fall into the corresponding spatial positions into corresponding spatial regions according to the position values of each image frame and the spatial positions of each spatial region in the three-dimensional space.
[0022] In a possible embodiment, the device also includes a deletion module, which is used to: when multiple supplementary key frames are determined, construct a spherical space for each supplementary key frame with a second preset distance as the sphere radius and a position value of each supplementary key frame as the sphere center position; for the spherical space of any supplementary key frame, delete all non-key frames in the spherical space of any supplementary key frame except itself.
[0023] In a possible implementation, the screening module is specifically configured to: input multiple image frames into the simultaneous positioning and mapping algorithm module for image processing, and output a processing result, where the processing result includes key frames screened out from the multiple image frames.
[0024] In one possible embodiment, the device also includes an association module, which is used to: establish an association relationship between the supplementary key frames and the key frames, and the association relationship represents the spatial correlation between the image frames in the three-dimensional space; the reconstruction module is specifically used to: perform three-dimensional reconstruction based on the association relationship, the supplementary key frames and the key frames to obtain a three-dimensional reconstruction model of the target object.
[0025] In a possible implementation, the multiple image frames are image frames obtained by capturing images by circling the target object with the target object as the center.
[0026] In one possible embodiment, the association module is further used to: perform image similarity calculation on any two image frames that have established an association relationship to obtain an image similarity score; for any two image frames corresponding to an image similarity score greater than a similarity threshold, based on the spatial relationship between the two image frames in three-dimensional space, if the spatial relationship does not meet a preset spatial relationship condition, cancel the established association relationship between the two image frames.
[0027] In one possible embodiment, the spatial relationship is represented by the absolute value of the difference between the index values of the spatial regions corresponding to the two image frames, and the index value is used to represent the position sequence of the corresponding spatial region in multiple spatial regions. The spatial relationship condition includes that the spatial relationship is less than or equal to a preset index threshold.
[0028] In a third aspect, an embodiment of the present application provides an electronic device comprising: a memory and a processor; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory, so that the processor executes the first aspect above and / or various possible implementations of the first aspect.
[0029] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, they are used to implement the first aspect above and / or various possible implementation methods of the first aspect.
[0030] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the above first aspect and / or various possible implementation methods of the first aspect.
[0031] The three-dimensional reconstruction method, apparatus, device, medium, and program product provided by the embodiments of the present application, after acquiring multiple image frames of a target object, first selects key frames from the image frames for constructing a three-dimensional reconstruction model. Since the number of key frames is small and the image frames involved in the three-dimensional reconstruction are relatively sparse, the method further identifies some non-key frames from each non-key frame as supplementary key frames to assist in constructing the three-dimensional reconstruction model. These supplementary key frames are non-key frames whose first distance in the three-dimensional space of the three-dimensional reconstruction model is greater than a first preset distance. Therefore, if the distance between these supplementary key frames and the key frames is sufficiently large, the image information provided by the supplementary key frames and the key frames has low duplication, and the supplementary key frames can provide image information that is helpful for the three-dimensional reconstruction. Therefore, when performing three-dimensional reconstruction based on the supplementary key frames and the key frames, the number of image frames involved in the reconstruction is neither too few nor too many, and there is no noise interference introduced by repeated image information between the image frames. Therefore, based on this method, a high-quality three-dimensional reconstruction model can be reconstructed efficiently, achieving both high three-dimensional reconstruction quality and high three-dimensional reconstruction efficiency during the three-dimensional reconstruction. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0033] Figure 1 Schematic diagram of the process of the 3D reconstruction method provided in the embodiment of the present application Figure 1 ;
[0034] Figure 2 A schematic diagram of a target object with multi-faceted similarity provided in an embodiment of the present application;
[0035] Figure 3 A schematic diagram of three-dimensional point cloud data provided in an embodiment of the present application;
[0036] Figure 4 The process of the 3D reconstruction method provided in the embodiment of the present application Figure 2 ;
[0037] Figure 5 A schematic diagram of a hemispherical space provided in an embodiment of the present application;
[0038] Figure 6 A schematic diagram of polar angles and azimuth angles provided in an embodiment of the present application;
[0039] Figure 7 A schematic diagram of cutting a hemispherical space provided in an embodiment of the present application;
[0040] Figure 8 A schematic diagram of a spherical space provided in an embodiment of the present application;
[0041] Figure 9 A schematic diagram of collecting images by circling a target object is provided for an embodiment of the present application;
[0042] Figure 10 A schematic diagram of the structure of a three-dimensional reconstruction device provided in an embodiment of the present application;
[0043] Figure 11 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application.
[0044] The above drawings illustrate specific embodiments of the present application, which will be described in more detail below. These drawings and the textual description are not intended to limit the scope of the present application in any way, but rather to illustrate the concepts of the present application to those skilled in the art by reference to specific embodiments. DETAILED DESCRIPTION
[0045] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with the present application. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present application, as detailed in the appended claims.
[0046] In the technical solutions of the embodiments of this application, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0047] In the embodiments of this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described in this application as "exemplary" or "for example" should not be construed as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.
[0048] In the embodiments of this application, if words such as "first" and "second" are used, they are intended to distinguish between identical or similar items with substantially the same functions and effects. For example, the first electronic device and the second electronic device are merely used to distinguish between different electronic devices and do not limit the order of precedence. Those skilled in the art will understand that words such as "first" and "second" do not limit the quantity or execution order, and words such as "first" and "second" do not necessarily mean that they are different.
[0049] In the embodiments of this application, "at least one" refers to one or more, and "more" refers to two or more. "And / or" describes the relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can mean: A exists alone, A and B exist simultaneously, and B exists alone, where A and B can be singular or plural. The character " / " generally indicates that the associated objects are in an "or" relationship.
[0050] For example, during visual three-dimensional reconstruction, an image acquisition device such as a camera is usually used to capture image frames of the target object, or existing image frames of the target object can be used for three-dimensional reconstruction. After obtaining the image frame, feature points are usually extracted from all image frames of the target object, and each feature point is matched by the correlation of the descriptor of the feature point. Among them, the descriptor can be understood as a vector that is manually defined or output by the end-to-end neural network to describe the regional information of the feature point. When the difference in factors such as viewing angle of an object in the environment is small, the regional information in different image frames is often relatively similar.
[0051] Camera calibration and pose estimation are then used to determine the camera's internal and external parameters. Based on this information, triangulation is used to generate sparse point cloud data. Bundle adjustment and other methods are then used to optimize the camera's relative pose and 3D point positions, thereby recovering the camera's position in the environment and the environment's structural information.
[0052] During 3D reconstruction, among all the image frames of the target object, some are critical for recovering the camera's position and pose information within the environment. These frames can be considered keyframes. Keyframes are not only crucial for recovering position and pose information, but are also crucial for building and optimizing the 3D model.
[0053] In some 3D reconstruction methods, only key frames are used for 3D reconstruction, while other 3D reconstruction methods use all image frames for 3D reconstruction. Since the number of key frames is usually relatively small and their distribution in 3D space may be relatively sparse, the 3D model obtained after reconstruction using only key frames will have large errors, resulting in low 3D reconstruction quality. In other methods, since the target object has many image frames, for example, acquisition devices such as cameras typically acquire images at a frequency of 30 frames per second, a large number of image frames will be generated after long-term acquisition at multiple acquisition points, and many image frames contain a large amount of repeated information. If all image frames are used for 3D reconstruction, not only will the processing of a large amount of image frame data increase the amount of calculation and the running time, greatly reducing the efficiency of 3D reconstruction, but the excessive number of repeated image frames will also introduce some noise interference, affecting the quality of 3D reconstruction. Therefore, none of these existing methods can achieve both high 3D reconstruction quality and 3D reconstruction efficiency.
[0054] In view of this, an embodiment of the present application provides a 3D reconstruction method. After acquiring multiple image frames of a target object, the method first selects key frames for constructing a 3D reconstruction model. Since the number of key frames is small and the image frames involved in the 3D reconstruction are relatively sparse, the method further identifies some non-key frames among the non-key frames as supplementary key frames to assist in constructing the 3D reconstruction model. These supplementary key frames are non-key frames whose first distance in the 3D space of the 3D reconstruction model is greater than a first preset distance. Therefore, if the distance between these supplementary key frames and the key frames is sufficiently large, the image information provided by the supplementary key frames and the key frames has low duplication, and the supplementary key frames can provide image information that is helpful for 3D reconstruction. Therefore, when performing 3D reconstruction based on the supplementary key frames and the key frames, the number of image frames involved in the reconstruction is neither too few nor too many, and there is no noise interference introduced by repeated image information between the image frames. Therefore, based on this method, a high-quality 3D reconstruction model can be reconstructed efficiently, achieving both high 3D reconstruction quality and high 3D reconstruction efficiency during 3D reconstruction.
[0055] The technical solutions of the present application are described in detail below with reference to specific embodiments. The specific embodiments below may be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments. The embodiments of the present application will be described below with reference to the accompanying drawings.
[0056] The methods provided in the embodiments of the present application can be applied to applications, websites, or mini-programs that have 3D reconstruction task processing capabilities. The 3D reconstruction task processing capabilities are implemented on the application, website, or mini-program. For example, a computer equipped with a 3D reconstruction application can implement the 3D reconstruction task processing capabilities by running the 3D reconstruction application. Another example is a terminal electronic device, such as a mobile phone, equipped with a 3D reconstruction mini-program can implement the 3D reconstruction task processing capabilities by running the 3D reconstruction mini-program.
[0057] Figure 1 Schematic diagram of the process of the 3D reconstruction method provided in the embodiment of the present application Figure 1 The execution subject of this method can be an electronic device with corresponding data storage and computing capabilities, such as a computer, mobile phone, server or server cluster. Figure 1 As shown, the method includes:
[0058] S101 , obtaining a plurality of image frames of a target object to be three-dimensionally reconstructed.
[0059] For example, the target object can be any object that needs to be 3D reconstructed, such as a car, building, or other object. The number of target objects is not limited to a single object and can be multiple objects. For example, the target object can be all buildings, pedestrians, trees, and vehicles in an environment scene.
[0060] When obtaining multiple image frames of the target object to be three-dimensionally reconstructed, they can be photographed and obtained through image acquisition devices such as cameras; they can also be read from the database through data transmission; or they can be obtained through other methods, which are not limited in the embodiments of the present application.
[0061] S102, selecting key frames from the plurality of image frames, and using the key frames to construct a three-dimensional reconstruction model.
[0062] For example, a keyframe is an image frame used to construct a 3D reconstruction model among all image frames. The number of keyframes can be one or more. This means that among multiple image frames, those that have a significant impact on 3D reconstruction are selected as keyframes. These selected keyframes can provide sufficient image information during the 3D reconstruction process while also controlling computational complexity and storage requirements.
[0063] There are many ways to filter key frames. For example, multiple image frames can be filtered to obtain key frames using predefined key frame indicators, or they can be filtered using some models or algorithms.
[0064] In a possible implementation, key frames are screened out from multiple image frames. Specifically, the multiple image frames are input into a simultaneous positioning and mapping algorithm module for image processing, and the processing results are output, where the processing results include the key frames screened out from the multiple image frames.
[0065] For example, the simultaneous localization and mapping (SLAM) algorithm can be understood as a robot moving from an unknown location in an unknown environment. During movement, it locates itself based on its location and a map. Simultaneously, it builds an incremental map based on its localization, enabling autonomous positioning and navigation. Visual SLAM involves performing SLAM tasks using images captured by a camera.
[0066] In the Simultaneous Localization and Mapping (SLAM) module, keyframes are image frames selected for map construction and optimization. Multiple image frames are fed into the Simultaneous Localization and Mapping (SLAM) module for image processing, and the SLAM module outputs the selected keyframes. The SLAM module typically considers factors such as motion changes between image frames, parallax changes, scene changes, and feature point coverage when selecting keyframes.
[0067] In an embodiment of the present application, by inputting multiple image frames and performing image processing in the simultaneous positioning and mapping algorithm module, key frames can be quickly screened out, and the screened key frames have high value for three-dimensional reconstruction. Therefore, this method can improve the quality and efficiency of screening key frames.
[0068] S103, among the non-key frames in the multiple image frames, some non-key frames are determined as supplementary key frames, and a first distance between any supplementary key frame and its closest key frame in the three-dimensional space of the three-dimensional reconstructed model is greater than a first preset distance, and the supplementary key frames are used to assist in constructing the three-dimensional reconstructed model.
[0069] For example, in order to increase the density of image frames involved in 3D reconstruction, a portion of non-key frames can be determined in addition to the key frames in multiple image frames to assist in building a 3D reconstruction model. These non-key frames can be understood as supplementary key frames.
[0070] Furthermore, in order to reduce the influence of large duplication between the determined supplementary key frames and the key frames on the reconstruction quality and efficiency, the determination may be made based on the distance between the non-key frames and the key frames.
[0071] For example, for any non-key frame, the distance between the position of the non-key frame and the position of each key frame is calculated; wherein, the position can be understood as the position information in the pose information of the image frame, and the position information can represent the relative position of the camera or other acquisition device when acquiring the image frame; the distance can be understood as the difference between the two positions, and the distance can also be understood as the second distance. In the case of multiple key frames, multiple distances can be obtained for the non-key frame, and the minimum distance among these distances is taken. The minimum distance can be understood as the first distance. By judging whether the first distance is greater than the first preset distance, if it is greater than the first preset distance, the non-key frame can be used as a supplementary key frame, and if it is less than or equal to the first preset distance, the non-key frame cannot be used as a supplementary key frame.
[0072] Based on this, non-key frames are first screened based on the condition that the first distance between any supplementary key frame and the key frame closest to it in the three-dimensional space of the three-dimensional reconstructed model is greater than the first preset distance. Non-key frames that meet the conditions of being supplementary key frames can be screened out, and then one or more supplementary key frames can be determined among these non-key frames by selecting according to preset rules or randomly selecting.
[0073] For example, the three-dimensional space of the three-dimensional reconstructed model can be understood as a virtual three-dimensional space pre-set when constructing the three-dimensional reconstructed model. The three-dimensional space can have a preset coordinate system for determining positions and distances. The first preset distance can be a preset distance of any length and can be set according to actual needs.
[0074] S104: Perform three-dimensional reconstruction based on the supplementary key frames and the key frames to obtain a three-dimensional reconstructed model of the target object.
[0075] For example, after selecting the key frames and determining the supplementary key frames, a three-dimensional reconstruction can be performed based on the supplementary key frames and the key frames to obtain a three-dimensional reconstructed model of the target object. Any software or module capable of three-dimensional reconstruction can be used for the three-dimensional reconstruction, and this application does not limit this. The three-dimensional reconstructed model of the target object can be understood as a stereoscopic model of the target object, which can be a stereoscopic model of the same size as the target object or a stereoscopic model scaled to a certain size.
[0076] The present invention provides a 3D reconstruction method. After acquiring multiple image frames of a target object, the method first selects key frames for constructing a 3D reconstruction model. Since the number of key frames is small and the image frames involved in the 3D reconstruction are sparse, the method further identifies some non-key frames within each non-key frame as supplementary key frames to assist in constructing the 3D reconstruction model. These supplementary key frames are non-key frames whose first distance in the 3D space of the 3D reconstruction model is greater than a first preset distance from the closest key frame. Therefore, if the distance between these supplementary key frames and the key frames is sufficiently large, the image information provided by the supplementary key frames and the key frames has low duplication, and the supplementary key frames can provide image information that is helpful for 3D reconstruction. Therefore, when performing 3D reconstruction based on the supplementary key frames and the key frames, the number of image frames involved in the reconstruction is neither too few nor too many, and there is no noise interference introduced by repeated image information between the image frames. Therefore, based on this method, a high-quality 3D reconstruction model can be reconstructed efficiently, achieving both high 3D reconstruction quality and high 3D reconstruction efficiency during 3D reconstruction.
[0077] For example, some existing visual 3D reconstruction algorithms, such as the open source software COLMAP or Open Multi-View Stereo (OpenMVS), which are based on multi-view stereo reconstruction and motion structure recovery, have a relatively time-consuming part in the overall reconstruction process, such as feature matching of feature points.
[0078] One reason feature matching is time-consuming is that images are disordered in time and space during 3D reconstruction. This disorder causes 3D reconstruction algorithms to commonly use brute-force matching. 3D reconstruction often uses computationally complex feature points and descriptors, such as the Scale-Invariant Feature Transform (SIFT). This is because 3D reconstruction must ensure robust matching of feature points across different scenes. Factors such as exposure or affine transformations cannot affect the majority of matching results, otherwise 3D reconstruction is likely to fail. However, the high complexity of feature points and descriptors also imposes computational pressure, potentially reducing reconstruction efficiency.
[0079] To improve the efficiency of feature point matching, some methods use an incremental matching method based on the temporal order of images. This can be understood as using the temporal order of image acquisition by the camera as prior knowledge to improve feature point matching efficiency. While incremental matching methods can speed up matching to a certain extent, they result in the current image being associated only with images from adjacent time periods, leading to a loss of image matching information. These factors can reduce the quality of 3D reconstruction.
[0080] In addition, some methods also introduce other non-visual sensors to provide prior information. Although this prior information can help improve matching efficiency to a certain extent, the addition of additional sensors will lead to additional steps such as external parameter calibration of sensors, time synchronization, and alignment of different sensor data. These additional steps will consume a lot of time, increase the complexity of hardware and algorithms, and reduce the efficiency of 3D reconstruction to a certain extent.
[0081] In addition to the constraints of 3D reconstruction itself, the scene type of 3D reconstruction will also affect the matching speed and reconstruction accuracy of 3D reconstruction. Some methods meet the reconstruction requirements for small-scale scenes with rich features and high feature point differentiation, but the effect is relatively poor for large scenes or objects with multiple similar faces. An in-depth analysis of the reasons is that large scenes are relatively complex and have large amounts of data. Different places often have similar features. These similar structures can easily lead to structural fractures or misalignment in the reconstruction results. Large scenes also require high computational complexity for algorithms or modules. Even after optimization, many tasks require days or even weeks to reconstruct.
[0082] Although the reconstruction of multi-faceted similar objects is not limited by the amount of scene data, it will encounter the problem of multi-faceted similar mismatching. Figure 2 A schematic diagram of a target object with multi-faceted similarity provided in an embodiment of the present application, such as Figure 2 As shown, the target object is a car. Figure 2 The two image frames in the figure were acquired by circling the target object. As can be seen, the two sides of the car are highly similar. Based on actual testing and other data, the probability of mismatching similar surfaces during circling an object for 3D reconstruction is high, which can lead to a certain degree of degradation in 3D reconstruction quality.
[0083] Figure 3 A schematic diagram of the three-dimensional point cloud data provided in the embodiment of the present application is shown as follows: Figure 3 As shown, the 3D point cloud data is the 3D point cloud data of the target object before the 3D reconstruction model is obtained. Figure 2 Before reconstruction, a camera is used to circle the target object to collect multiple image frames, and the 3D point cloud data is obtained by 3D reconstruction software. Figure 3 As can be clearly seen in the figure, due to the similarities between the left and right sides of the car (especially the tires), the restored 3D point cloud structure of the car exhibits layering. Therefore, when further 3D reconstruction is performed based on this 3D point cloud data, the resulting 3D reconstructed model will also exhibit layering, affecting the reconstruction quality.
[0084] Furthermore, conventional indoor and outdoor 3D reconstructions all follow a collection route that goes from a starting point to an end point and back to the starting point, and the scene must be covered from all angles to obtain relatively complete scene information. However, the collection route for 3D reconstruction of objects that orbit around the object is centered on the object, with the collector circling around the object to collect image frames. This collection method results in repeated information in many image frames. Excessive repeated information reduces the computational efficiency of feature point matching and the bundle adjustment method for back-end optimization. Furthermore, excessive redundant information not only fails to improve the quality of 3D reconstruction, but also reduces reconstruction quality due to the introduction of excessive noise. Therefore, for 3D reconstruction of objects that orbit around the object, it is necessary to screen the data to find the optimal number of image frames that can ensure reconstruction quality and speed.
[0085] Based on the above description, the three-dimensional reconstruction method provided in the embodiment of the present application can be applied to the scenario of three-dimensional reconstruction of orbiting objects, and the reconstruction quality and reconstruction efficiency of three-dimensional reconstruction of orbiting objects can be improved by processing such as screening key frames and determining supplementary key frames.
[0086] In one possible implementation, the method further includes: establishing an association relationship between the supplementary key frames and the key frames, the association relationship characterizing the spatial correlation between the image frames in three-dimensional space; performing three-dimensional reconstruction based on the supplementary key frames and the key frames to obtain a three-dimensional reconstruction model of the target object, including: performing three-dimensional reconstruction based on the association relationship, the supplementary key frames and the key frames to obtain a three-dimensional reconstruction model of the target object.
[0087] For example, the correlation relationship represents the spatial correlation between image frames in three-dimensional space. For example, if the same objects can be observed in two image frames, then the two image frames have a certain common view relationship, which also indicates that the two image frames have a certain spatial correlation in three-dimensional space. These two image frames can be used to reconstruct and optimize a 3D reconstruction model within a certain range.
[0088] One of the purposes of establishing an association relationship between supplementary key frames and key frames is to establish a certain association relationship according to the spatial correlation between these image frames in the early stage of 3D reconstruction. This association relationship can be used as prior knowledge of 3D reconstruction to guide the 3D reconstruction process, thereby helping to speed up the 3D reconstruction and improve the quality.
[0089] Establishing an association relationship between a supplementary keyframe and a keyframe can be achieved in a variety of ways. For example, for any two image frames in the supplementary keyframe and the keyframe, feature points are extracted and represented by descriptors. The extracted feature points can be matched by matching the descriptors. If a certain number of feature points or a certain area of feature points are matched, it can be determined that the two have spatial correlation. The degree of spatial correlation can be further determined by the magnitude of the hits, and then whether to establish an association relationship between the two can be determined based on the degree of spatial correlation.
[0090] Alternatively, an association relationship may be established in other ways. For example, the processing result output by the SLAM module may also include a key frame co-view diagram between key frames. The key frame co-view diagram may be understood as a co-view relationship expression of a graph data structure. By using the key frame co-view diagram and the co-view relationship between the supplementary key frames and the key frames, an association relationship between the supplementary key frames and the key frames is further established. Of course, an association relationship may also be established in other ways, which will not be described in detail here.
[0091] Furthermore, when performing 3D reconstruction based on the supplementary keyframes and the keyframes to obtain the 3D reconstructed model of the target object, the 3D reconstruction can be performed based on the association relationship, the supplementary keyframes and the keyframes to obtain the 3D reconstructed model of the target object.
[0092] In an embodiment of the present application, an association relationship is established between each supplementary key frame and multiple key frames to characterize the spatial correlation between image frames in three-dimensional space. This association relationship can provide effective prior knowledge for three-dimensional reconstruction, which helps to improve the quality and efficiency of three-dimensional reconstruction.
[0093] In a possible implementation, the multiple image frames are image frames obtained by capturing images by circling the target object with the target object as the center.
[0094] For example, if the target object is a car, a camera or other image acquisition device can be used to capture images by circling the car one or more times, with the car as the center. This can generate image frames of the car. Of course, the direction, speed, and frequency of the circling can be freely selected, and image capture can be performed at any location in real space other than the ground. For example, image frames can be captured from the top or side of the car.
[0095] In an embodiment of the present application, the application scenario of three-dimensional reconstruction is a scenario in which images are collected by circling a target object and three-dimensional reconstruction is performed. In this scenario, the three-dimensional reconstruction method of an embodiment of the present application can improve the reconstruction efficiency and reconstruction quality of the three-dimensional reconstruction of the circling object.
[0096] For example, from the above description of the situation where mismatching is prone to occur in the three-dimensional reconstruction of the orbiting object, it can be seen that mismatching may occur when establishing an association relationship between the supplementary key frame and the key frame.
[0097] For example, if an image frame captured from the left front of a car and another from the right front have very similar image content, feature point extraction and feature point descriptor matching might lead to a high degree of spatial correlation between the two, leading to an association between the two. However, in reality, one image was captured from the left side of the car, while the other from the right. The locations of the capture points differ significantly, and while the images are similar, their spatial correlation is low or even non-existent. Therefore, for image frames that have already been associated, it is necessary to identify and delete false matches, thereby dissolving the association between false matches.
[0098] In one possible implementation, after establishing an association relationship between the supplementary key frame and the key frame, the method further includes: performing image similarity calculation on any two image frames with an established association relationship to obtain an image similarity score; and for any two image frames corresponding to an image similarity score greater than a similarity threshold, based on a spatial relationship between the two image frames in three-dimensional space, if the spatial relationship does not meet a preset spatial relationship condition, releasing the established association relationship between the two image frames.
[0099] For example, when calculating the image similarity between two image frames, any algorithm or neural network model that can perform image similarity calculation can be used for calculation, and the obtained image similarity score can be a predicted or evaluated value used to characterize the degree of similarity of the image content of the two image frames.
[0100] The similarity threshold can be any preset threshold used to determine the degree of similarity between image content. For example, for any two image frames that have established a relationship, after inputting them into the similarity neural network model, an image similarity score can be obtained. If the image similarity score is less than or equal to the similarity threshold, the relationship between the two can be retained; if the image similarity score is greater than the similarity threshold, a further determination is made as to whether to disassociate the two images.
[0101] When determining whether to release the association between the two, the established association between the two image frames may be released based on the spatial relationship between the two image frames in the three-dimensional space. If the spatial relationship does not meet a preset spatial relationship condition, the established association between the two image frames may be released.
[0102] Among them, the spatial relationship between the two image frames in the three-dimensional space can be a spatial position relationship or a spatial distance relationship, etc. The spatial relationship can be a relationship that characterizes the acquisition positions of the two image frames in space. The preset spatial relationship condition can be a condition preset for determining the degree of the spatial relationship based on the spatial relationship. For example, if the spatial relationship is the spatial distance relationship between the two image frames in the three-dimensional space, then the spatial distance between the two image frames in the three-dimensional space is calculated. The preset spatial relationship condition can be a preset distance threshold. If the calculated spatial distance is greater than the preset distance threshold, it indicates that the two image frames are extremely similar, but their acquisition positions are very different. In this case, it can be determined that the established association relationship is caused by a mismatch, and the established association relationship between the two image frames can be released at this time.
[0103] After all mismatched associations are resolved, it is equivalent to correcting the established associations, so that the accuracy of the prior knowledge used for 3D reconstruction is improved, which can improve the accuracy of the 3D reconstruction model and improve the quality of 3D reconstruction.
[0104] In an embodiment of the present application, in a scenario where a target object is being circumvented for three-dimensional reconstruction, if the target object has multiple similar surfaces (such as the left and right sides of a car), there may be erroneous associations in the established association relationship due to mismatching. These erroneous associations can cause phenomena such as dislocation during three-dimensional reconstruction, affecting the reconstruction quality. Therefore, this solution calculates image similarity and determines spatial relationship conditions for two image frames with high image similarity. If the spatial relationship between the two image frames does not meet the spatial relationship conditions, it indicates that although the two image frames are very similar, their acquisition positions differ significantly, and they are likely not images acquired at similar positions. Therefore, the established association between the two is likely an erroneous association due to mismatching. At this time, by dissolving the association between the two, phenomena such as dislocation can be avoided when the association relationship is subsequently used for three-dimensional reconstruction, thereby improving the reconstruction quality of the circumvented target object.
[0105] Next, combined with Figure 4 The three-dimensional reconstruction method provided in the embodiments of the present application is further described. Figure 4 The process of the 3D reconstruction method provided in the embodiment of the present application Figure 2 ,like Figure 4 As shown, the method includes the following steps:
[0106] S401, visual SLAM generates image poses and map points. S402, object detection algorithm generates a three-dimensional (3D) rectangular frame. S403, construct a hemispherical space with the reconstructed target object as the center and cut it to obtain multiple hemispherical primitives. S404, after the hemispherical space is cut, the image frames divided into the hemispherical primitives are filtered. S405, the association relationship between image frames is established through the co-viewing image and the location map. S406, image similarity detection is performed on the image frames with established association relationships, and the association relationship is filtered based on the index value of the hemispherical primitive. Of course, Figure 4 The three-dimensional reconstruction method shown can also be understood as some steps performed in the early stage of the three-dimensional reconstruction process. After S406, steps such as S104 can be performed to complete the three-dimensional reconstruction.
[0107] like Figure 4 As shown, a camera device can capture multiple frames by circling the target object multiple times. This image stream is then fed into the visual SLAM algorithm module to generate the pose of each image and a global visual point cloud map. The visual SLAM algorithm module can be understood as the aforementioned simultaneous localization and mapping algorithm module, or SLAM module.
[0108] An example of a visual SLAM algorithm is oriented fast and rotated brief (ORB) SLAM. ORB SLAM can be understood as an algorithm that uses ORB features in images for correlation and matching to perform SLAM tasks. Furthermore, visual SLAM is a real-time algorithm, and its inclusion in 3D reconstruction algorithms has negligible computational overhead.
[0109] The collected image stream is input into the visual SLAM algorithm, which will generate the pose of each image frame (including position information and pose information in the algorithm coordinate system). In addition, it will also filter out key frames, key frame co-visual views, the relationship between non-key frames and key frames, and three-dimensional point cloud data centered on the target object.
[0110] For example, most SLAM algorithms follow the following process: the front-end estimates a rough initial pose for the current image frame through image feature matching, then sends this estimated initial pose to the back-end optimization module. The back-end optimization module constructs a graph optimization problem based on the poses of local map points and other keyframes. Typically, a camera generates 30 image frames per second, and the pose of each image frame must be estimated. Each image frame is a normal image frame, used only to generate its own pose. SLAM back-end optimization relies on the pose graph and map points constructed from the image poses. Map generation is time-consuming, so it is not advisable to include every image frame in map point generation. Furthermore, an overly dense pose graph will result in excessive computational overhead in subsequent graph optimization and bundle adjustment. As a real-time algorithm, the SLAM algorithm filters normal image frames. These frames are then converted into keyframes based on temporal and spatial dimensions and other conditions. These frames can then participate in map point generation and back-end optimization.
[0111] The method provided in the embodiment of the present application first processes the collected image stream with a visual SLAM algorithm, outputs the camera pose trajectory, map points, and keyframe co-view graph, and provides global temporal and spatial prior information. Moreover, the visual SLAM algorithm is an efficient and real-time algorithm that does not affect the overall time efficiency of 3D reconstruction. A visual SLAM algorithm is added to the 3D reconstruction of the circling object, and the co-view graph relationship generated by visual SLAM is used to determine the association relationship of the image frames. The key frame of visual SLAM is the medium for associating global information. It participates in the generation of global map points and the optimization of the global pose graph throughout the entire visual SLAM operation process, so this type of image frame can effectively provide the association relationship of global feature matching.
[0112] Because image frames captured by a single camera lack scale information, the corresponding monocular visual SLAM also lacks scale information. Scale information here refers to the dimensionless nature of the generated 3D information, making it impossible to determine the image position or the distance units of the reconstructed object, such as whether it is in meters or centimeters. While this lack of scale information can be accounted for through relative proportional relationships, after obtaining the visual SLAM results, the dimensions of the objects surrounding the central target can be determined.
[0113] For example, a 3D object detection algorithm can be used to process scale information. This algorithm can be any algorithm for detecting 3D objects, such as VoxelNet (End-to-End Learning for Point Cloud Based 3D Object Detection), PointPillars (Fast Encoders for Object Detection from Point Clouds), and SECOND (Sparsely Embedded Convolutional Detection), or any other 3D object detection algorithm used in autonomous driving technology.
[0114] Using a 3D object detection algorithm, a specific target object, such as a pedestrian, car, or tree, can be detected in a 3D point cloud. By inputting the 3D point cloud map generated by the SLAM algorithm into the 3D object detection algorithm, a corresponding 3D bounding box can be generated.
[0115] Assume that the vertices corresponding to the 3D rectangular box generated by the 3D object detection algorithm are: A, B, C, D, E, F, G, and H. The center point of the bottom rectangle ABCD of the rectangular box can be taken as O. With point O as the center of the circle and the product of the long side BC of the bottom rectangle and the scale factor α as the radius r, a hemispherical space is constructed. Where: , where d bc Indicates the length of side BC.
[0116] Based on this, a hemispherical space (Sphere Space) can be generated. Figure 5 A schematic diagram of a hemispherical space provided in an embodiment of the present application is shown in FIG. Figure 5 As shown, the truck is the target object, and a hemispherical space can be constructed based on the bottom surface of the target object passing through the three-dimensional rectangular box. The value of the scale coefficient α determines the size of the hemispherical space, and the scale coefficient α can be preset according to actual needs. For example, if image frames of different distances are required, the scale coefficient α can take a larger value to make the hemispherical space larger. The scale coefficient α can be, for example, between 1.5 and 2. In addition, the method for generating the hemispherical space uses the relative position distance of the point cloud as the basis, so the corresponding hemispherical space can be generated without obtaining the absolute scale.
[0117] The method provided in the embodiment of the present application adopts a 3D target object detection algorithm to obtain the outer 3D rectangular frame of the target reconstructed object point cloud, and uses the length of the three-dimensional rectangular frame to generate a hemispherical space. The hemispherical space can be divided into multiple hemispherical primitives in a manner that imitates the longitude and latitude of the earth, and the divided spherical primitives are used as units. At the same time, the spherical space is generated with the position center of the key frame, and the excess image frames in the spherical space are deleted. At the same time, several supplementary image frames are randomly selected from the non-key frames to generate a spherical space of the same size. The redundant non-key frames in the spherical space can be deleted, and the image screening can be completed efficiently, while making the filtered image distribution more even.
[0118] After the hemispherical space is generated, it can be segmented in one of a variety of ways. For example, it can be segmented in a random manner, or in a horizontal and vertical manner. This embodiment of the present application provides a segmentation method based on the Earth's longitude and latitude.
[0119] Figure 6 The polar angle and direction angle diagram provided in the embodiment of the present application is as follows: Figure 6 As shown, a polar angle can be defined is the vector between point P1 on the sphere and the origin O The angle with the xy-axis. Alternatively, the polar angle can be defined as the angle with the z-axis. In this embodiment, the polar angle is defined as the angle with the xy-axis, similar to the latitude in longitude and latitude. The polar angle range is [0, π / 2]. is the vector between point P1 on the sphere and the origin O The angle between the projection of the xy-plane and the x-axis is similar to the longitude in latitude and longitude, and is in the range [0, 2π]. The polar angle and the azimuth angle can form a spherical coordinate system. According to the definition rules of the polar angle and azimuth angle, the spherical coordinate system can be converted to the rectangular coordinate system. This can be achieved by the following formula:
[0120]
[0121] in, represents the radius of the hemisphere; x, y, and z are the coordinates of point P1 in the rectangular coordinate system.
[0122] Figure 7 The schematic diagram of cutting the hemispherical space provided in the embodiment of the present application is as follows: Figure 7 As shown, the hemisphere space is divided with reference to the longitude and latitude of the earth, that is, the polar angle resolution (latitude direction), directional angle resolution (longitude direction), the hemisphere space can be divided into indivual, It can be calculated by the following formula:
[0123]
[0124] Each space region after segmentation can be understood as a hemisphere primitive, which is composed of the intersection of the sphere center O, the cutting line in the polar angle direction and the cutting line in the azimuth direction. For example, Figure 7 The intersection points P1, P2, P3, P4 and the center O constitute a spherical primitive V1.
[0125] In the embodiment of the present application, all hemispherical primitives can be sorted by index value in order. For example, a hemispherical primitive is represented as V mn , where m and n represent the index values of the direction angle and polar angle respectively.
[0126] After completing the hemispherical space segmentation, multiple image frames can be classified into corresponding hemispherical primitives.
[0127] You can first determine whether the image frame is in the hemispherical space. You only need to calculate the position of the image frame in the rectangular coordinate system of the hemisphere. The distance to the origin (center of the sphere) is as follows:
[0128]
[0129] in, Indicates location The distance to the origin. When , it means that the image frame is not in the hemispherical space and can be deleted. After that, the polar angle and direction angle of the image frame in the spherical coordinate system can be calculated using the following formula:
[0130]
[0131] in, represents the polar angle; Indicates the direction angle; represents the arccosine function; Represents the inverse tangent function. When the polar angle resolution and azimuthal angle resolution are known, the index value of the hemispherical primitive to which the image frame belongs can be obtained. Specifically, it can be calculated using the following formula:
[0132]
[0133] in, After calculating the index value of the corresponding hemispherical primitive for each image frame, we can get a list of image frames within each hemispherical primitive, which can be expressed as:
[0134]
[0135] in, Indicates V mn a list of image frames for the hemisphere primitive, Represents an image frame ; Represents an image frame .
[0136] Afterwards, the image frames in each hemispherical primitive can be filtered. Since the camera can usually generate 30 image frames per second, the image frames in space are relatively dense. The key frame of visual SLAM is a sparsely distributed image frame that has been strictly screened. Some existing 3D reconstruction methods directly use key frames as input image frames for 3D reconstruction. However, in order to improve optimization efficiency and speed, visual SLAM often sets the key frame images to be very sparsely distributed in time and space, resulting in the number of map points directly used for 3D reconstruction being unable to meet the requirements of 3D reconstruction. Among them, map points can be understood as points on the surface of the 3D reconstruction model.
[0137] Therefore, based on the keyframes, we can appropriately increase the number of image frames involved in 3D reconstruction to make the generated map points meet the requirements, while not adding too many image frames, which will significantly increase the time of 3D reconstruction. All keyframes within the hemisphere primitive can be used for 3D reconstruction, which can be expressed as:
[0138]
[0139] in, Represents a list of keyframes, Indicates key frame keyl; Indicates the key frame keyq.
[0140] Furthermore, a spherical space with a radius of rI can be generated with the position of the key frame as the center. Figure 8 A schematic diagram of a spherical space provided in an embodiment of the present application is shown in FIG. Figure 8 As shown, assuming that the position of the key frame in the three-dimensional space is , the position of non-keyframe is , then the distance between the non-key frame and the key frame The distance can be calculated by the following formula: It can also be understood as the second distance:
[0141]
[0142] A first preset distance d can be set λ , dλ Can be equal to the radius rI. When the current non-key frame I i Within the spherical space of the key frame, the candidate sequence can be eliminated.
[0143] like Figure 8 As shown, the image frame I1 and the image frame I4 are respectively in the key frame I key1 and keyframe I key2 Therefore, these two image frames are not considered as candidate image frames for participating in 3D reconstruction. To count the remaining non-key frames, multiple image frames can be randomly selected in a random manner, requiring that the first distance of these image frames is greater than the first preset distance d λ , then these non-key frames are supplementary key frames.
[0144] For example, these selected supplementary key frames can also generate their own corresponding spherical spaces. Similar methods can be used to count whether other non-key frames are in the spherical spaces of these supplementary key frames, and delete other non-key frames in the spherical space except for themselves. For example, Figure 8 After the corresponding spherical space is generated for the image frame I2, the image frame I3 is deleted through calculation and judgment.
[0145] Exemplarily, the above-mentioned deletion of non-key frames in the spherical space may be repeatedly performed multiple times until all the supplementary image frames generate corresponding spherical spaces.
[0146] The method provided in the embodiment of the present application takes into account the matching of special positions. The special positions here refer to positions of some associated key frames that are far apart, and even the camera shooting angles are different, but the images obtained are similar, such as the front area and other positions. The image information obtained by some image frames that establish association relationships is similar, and this similarity can lead to mismatching. The image similarity detection algorithm in deep learning is used to detect the similarity of images, and the index value of the hemisphere primitive in the direction angle is used to avoid the erroneous elimination of the association relationship of image features with similar positions, so that those image frames that should not be matched can be correctly disassociated.
[0147] After image screening is complete, the next step is to establish a relationship between the supplementary keyframes and the keyframes. Establishing this relationship is to accelerate the feature matching speed of 3D reconstruction and to reduce mismatching of similar features on multiple surfaces.
[0148] For example, the backend of visual SLAM generates a co-visual map of keyframes. This co-visual map is constructed by counting the number of map points observed between keyframes. That is, if a certain number of map points can be observed between keyframes, then two keyframes have a co-visual map and are connected in the graph data structure. Typically, when capturing multiple laps around a target object, a keyframe may be associated with more than a dozen keyframes.
[0149] In addition, the feature matching of the front end during the operation of visual SLAM is often calculated by associating the previous image frame and the local map points to obtain the initial pose, and the local map points are generated and associated by the key frames. Therefore, the ORB SLAM algorithm will select a reference key frame I for the current image frame when performing the initial pose estimation. key_ref The reference keyframe is the keyframe with the greatest degree of co-visibility with the current image frame, i.e., it has the largest number of observed common map points. Therefore, in this embodiment of the present application, the correlation relationship of feature matching can be obtained through the keyframe co-visibility.
[0150] For example, the key frame co-visual map generated by visual SLAM can be obtained with the key frame There are key frames with common viewing relationship, and with key frames The association relationship after feature matching can be expressed as:
[0151]
[0152] in, Represents a keyframe A list of associated relationships, and The key frames are all key frames that establish an associated relationship with it.
[0153] If the reference key frame corresponding to a supplementary key frame belongs to a key frame or A key frame in the list, then it means that the supplementary key frame and the key frame If there is a common view relationship, the supplementary key frame can be added to the list of associated relationships. Repeat this process until all supplementary key frames are traversed, and finally the key frame image can be obtained. The list of associations for feature matching is:
[0154]
[0155] in, and etc. indicate additional keyframes.
[0156] By performing the above operations on all key frames, the association relationship between each supplementary key frame and each key frame can be established.
[0157] Although the association relationship of feature matching obtained by using the key frame co-viewing relationship can avoid mismatching of multi-faceted similar images in most cases, mismatching may still occur in some special cases.
[0158] Figure 9 A schematic diagram of collecting images by circling a target object is provided for the embodiment of the present application, such as Figure 9 As shown, the image frames taken on both sides of the front position of the car, such as image I key1 , I key2 , I key3 and I key4 Images with similar or even identical content but captured at different locations may be associated through the keyframe at the center of the vehicle head. This may result in images being included in the same keyframe's association list, leading to mismatched images and potentially mismatched layers.
[0159] In order to avoid this situation, the image similarity of the image frame and the hemispherical primitive to which it belongs can be comprehensively considered.
[0160] In one possible implementation, the spatial relationship is represented by the absolute value of the difference between the index values of the spatial regions corresponding to the two image frames. The index value is used to represent the position sequence of the corresponding spatial region in multiple spatial regions. The spatial relationship condition includes that the spatial relationship is less than or equal to a preset index threshold.
[0161] For example, the image similarity score can be calculated using an image similarity detection algorithm such as Vision Transformer (VIT) or Self-Distillation with No Labels (DINO). List and calculate the image similarity score S between all image frames ij . You can preset the similarity threshold S λ , when S ij >S λ , it means that the two image frames are highly similar, and it is necessary to consider disassociating them.
[0162] However, some adjacent images are often similar and should be associated with each other. Figure 9 I in key3 and I key4 Therefore, on the one hand, we need to set the similarity threshold S λ Improve, try to avoid accidentally deleting the correct association. On the other hand, it is also necessary to constrain the strength of removing the association through spatial relationships.
[0163] The embodiment of the present application adopts the index value of the hemispherical primitive to which the two image frames belong and The absolute value of the difference between them is used to represent the spatial relationship. You can take the index value corresponding to the direction angle and calculate the absolute value of the difference between the index values. You can use the following formula to calculate:
[0164]
[0165] in, Indicates the absolute value of the difference between index values. Represents an image frame The index value corresponding to the direction angle; Represents an image frame The index value corresponding to the direction angle.
[0166] when Greater than the preset index threshold , it indicates that the positions of the two image frames are far apart, and the established association between the two similar image frames can be released.
[0167] In this embodiment of the present application, the spatial relationship is characterized by the absolute difference between the index values of the spatial regions corresponding to the two image frames. The spatial relationship condition includes the spatial relationship being less than or equal to a preset index threshold. Based on this spatial relationship and the spatial relationship condition, it is possible to effectively determine whether two similar image frames are mismatched or incorrectly associated, thereby improving the accuracy of determining and executing the disassociation.
[0168] Ultimately, the image frames involved in 3D reconstruction and the relationships between them are obtained, enabling image feature matching during 3D reconstruction while ensuring computational efficiency and enabling global correlation with other image frames. Furthermore, this method has no restrictions on the type of target object to be reconstructed within the algorithmic flow. For example, if the target object to be reconstructed is located in the center, the trajectory of the acquired image frames can simply be around it. Therefore, this method is applicable not only to cars but also to other objects.
[0169] The method provided in the embodiment of the present application can solve the problem of mismatching when circumventing multi-faceted similar objects, and improve the speed and quality of three-dimensional reconstruction. Based on the visual SLAM method, corresponding map points and image poses are generated. Through these prior information, the problem of mismatching of similar features in different positions can be avoided. Accelerate the feature matching speed and reduce redundant image information: by taking the target object as the center, cutting the space where it is located into hemispherical voxel units (hemispherical primitives), and screening the image frames within the same voxel unit to reduce the total number of image frames involved in three-dimensional reconstruction. The fewer the number of image frames, the faster the feature matching speed. Appropriately reducing redundant images is also conducive to improving the reconstruction quality. Therefore, this method can provide a more efficient solution for image frame screening, matching association relationships and three-dimensional reconstruction of three-dimensional reconstruction around objects.
[0170] In one possible implementation, before dividing the multiple image frames into multiple spatial regions, the method also includes: constructing a hemispherical space surrounding the target object based on the point cloud data of the target object; dividing the hemispherical space into multiple hemispherical primitives, and the multiple hemispherical primitives correspond to the multiple spatial regions respectively.
[0171] For example, during 3D reconstruction, constructing a hemispherical space centered on an object is more practical than constructing a rectangular or cube-shaped space. For example, the spherical space in the lower half of the target object's plane may not have captured image frames in the real space corresponding to this spherical space. Therefore, a hemispherical space is more practical.
[0172] The advantage of the hemispherical space is that any point on the surface of the hemisphere is at the same distance from the center of the sphere. Therefore, the hemispherical space can filter image frames whose acquisition positions are always within the spherical radius of the hemisphere, and can filter out image frames acquired at acquisition positions outside the spherical radius of the hemisphere. However, it is not easy to achieve the above filtering effect of the hemispherical space in a cube or rectangular space.
[0173] For example, the hemispherical space may be divided into a plurality of hemispherical primitives by simulating the longitude and latitude of the earth as described in the above embodiment, which will not be described in detail here.
[0174] In the embodiments of this application, hemispherical spaces are more suitable for 3D reconstruction scenarios than spaces of other shapes. This solution constructs a hemispherical space surrounding the target object using the point cloud data of the target object and segments the hemispherical space. This can quickly segment multiple hemispherical primitives, achieving the goal of quickly and efficiently setting up multiple spatial regions.
[0175] In one possible implementation, some non-key frames are determined as supplementary key frames, including: dividing multiple image frames into multiple spatial regions, where the multiple spatial regions are set according to the spatial position of the target object in the three-dimensional space; calculating the second distance between any non-key frame and any key frame in each of the multiple spatial regions; and determining the supplementary key frame in each non-key frame whose second distance is greater than the first preset distance.
[0176] Illustratively, the multiple spatial regions may be spatial regions set according to the spatial position of the target object in three-dimensional space. The multiple spatial regions may be, for example, spatial regions of any shape, such as a cube, a cuboid, a cylinder, etc. Illustratively, the multiple spatial regions may also be spatial regions corresponding to the multiple hemispherical primitives described above.
[0177] A second distance is calculated between any non-key frame and any key frame in each of the plurality of spatial regions; and a supplementary key frame is determined for each non-key frame for which the second distance is greater than the first predetermined distance. The second distance is the distance between the non-key frame and the key frame, and may be a distance obtained by calculating the difference in position in three-dimensional space. The second distance may refer to the relevant descriptions in the above embodiments.
[0178] If a non-key frame is determined as a supplementary key frame one by one according to multiple image frames, the calculation process during the determination is limited by a serial method, resulting in low determination efficiency.
[0179] In an embodiment of the present application, all image frames are divided according to spatial regions, and rapid calculations can be performed in multiple spatial regions to determine supplementary key frames, thereby achieving the purpose of parallel processing of multiple image frames and improving the efficiency of determining supplementary key frames.
[0180] In one possible implementation, a hemispherical space surrounding the target object is constructed based on the point cloud data of the target object, including: inputting the point cloud data of the target object into a three-dimensional target object detection algorithm module for processing to generate a three-dimensional rectangular frame surrounding the target object; and constructing a hemispherical space surrounding the target object with the center point of the bottom surface of the three-dimensional rectangular frame as the center of the sphere and the bottom surface as the center cross-section of the hemispherical space.
[0181] For example, the point cloud data of the target object can be obtained through the processing results of the SLAM module, or can also be collected by a distance point cloud data sensor such as radar. The three-dimensional target object detection algorithm module can be a software algorithm module that implements three-dimensional target object detection. The specific implementation of the three-dimensional target object detection algorithm, the three-dimensional rectangular frame, and the construction of a hemispherical space surrounding the target object using the bottom surface of the three-dimensional rectangular frame can refer to the description of the above embodiments.
[0182] In the embodiments of the present application, a 3D object detection algorithm is used to process the point cloud data, assigning scale information to it and generating a 3D rectangular box that conforms to the scale of the point cloud data and surrounds the target object. By constructing a hemispherical space surrounding the target object based on the 3D rectangular box, a hemispherical space can be constructed with greater accuracy and more consistent with the size of the target object.
[0183] In one possible implementation, multiple image frames are divided into multiple spatial regions, including: converting the position information in the posture of each image frame into a three-dimensional space to obtain the position value of each image frame in the three-dimensional space; and dividing the image frames whose position values fall into the corresponding spatial positions into the corresponding spatial regions based on the position values of each image frame and the spatial positions of each spatial region in the three-dimensional space.
[0184] For example, the position information in the pose of each image frame can be obtained by estimating the pose of the image frame, wherein the pose includes position and posture. The position information can be understood as the acquisition position of the acquisition device.
[0185] Using the transformation matrix between the pose coordinate system and the 3D space coordinate system, the position information in the pose of the image frame can be converted into 3D space, obtaining the position value of each image frame in 3D space. Since each spatial region corresponds to a portion of the spatial position in 3D space, the image frames whose position values fall into the corresponding spatial position can be divided into the corresponding spatial region based on the position value of each image frame and the spatial position of each spatial region in 3D space.
[0186] In an embodiment of the present application, when multiple image frames are divided into multiple spatial areas, the position information in the image frame's posture is first converted into three-dimensional space, which can achieve the unification of the coordinate system. Then, according to the position value of each image frame and the spatial position of each spatial area in the three-dimensional space, the multiple image frames can be divided into multiple spatial areas in an orderly manner. The position information of the image frames will not be lost after the division, thereby improving the accuracy of the determined first distances and improving the quality of determining the supplementary key frames.
[0187] In a possible implementation, the method further includes: when multiple supplementary key frames are determined, constructing a spherical space for each supplementary key frame with a second preset distance as the sphere radius and a position value of each supplementary key frame as the sphere center position; for the spherical space of any supplementary key frame, deleting all non-key frames in the spherical space of any supplementary key frame except itself.
[0188] For example, the second preset distance can be understood as a radius value preset for the supplementary key frame, and the radius value can be equal to the first preset distance. For example, the second preset distance is equal to the first preset distance d λ, can also be equal to the radius rI of the spherical space of the key frame, etc.
[0189] For the specific implementation of constructing the spherical space of each supplementary key frame and deleting other non-key frames in the spherical space of any supplementary key frame except itself, reference may be made to the relevant description in the above embodiment, which will not be repeated here.
[0190] In an embodiment of the present application, there may be other non-key frames within the second preset distance range of the supplementary key frame. These non-key frames may also be supplementary key frames. When two supplementary key frames with closer distances are used for three-dimensional reconstruction, the reconstruction efficiency and reconstruction quality may be affected. Therefore, by deleting other non-key frames except the supplementary key frame itself in the spherical space of the supplementary key frame, unnecessary non-key frames can be removed, which helps to further improve the reconstruction quality and efficiency.
[0191] Figure 10 A schematic diagram of the structure of a three-dimensional reconstruction device provided in an embodiment of the present application is shown in FIG. Figure 10 As shown, an embodiment of the present application provides a three-dimensional reconstruction device, which includes: an acquisition module 1001, used to acquire multiple image frames of a target object to be three-dimensionally reconstructed; a screening module 1002, used to screen out key frames from the multiple image frames, the key frames being used to construct a three-dimensional reconstruction model; a determination module 1003, used to determine some non-key frames in each non-key frame in the multiple image frames as supplementary key frames, wherein a first distance between any supplementary key frame and the key frame closest to it in the three-dimensional space of the three-dimensional reconstruction model is greater than a first preset distance, and the supplementary key frame is used to assist in constructing the three-dimensional reconstruction model; a reconstruction module 1004, used to perform three-dimensional reconstruction based on the supplementary key frames and the key frames to obtain a three-dimensional reconstruction model of the target object.
[0192] In a possible implementation, the determination module 1003 is specifically used to: divide multiple image frames into multiple spatial regions, where the multiple spatial regions are set according to the spatial position of the target object in the three-dimensional space; calculate the second distance between any non-key frame and any key frame in each spatial region of the multiple spatial regions; and determine a supplementary key frame in each non-key frame whose second distance is greater than the first preset distance.
[0193] In a possible embodiment, the device also includes a construction module, which is used to: construct a hemispherical space surrounding the target object based on the point cloud data of the target object; divide the hemispherical space into multiple hemispherical primitives, and the multiple hemispherical primitives correspond to multiple spatial regions respectively.
[0194] In one possible embodiment, the construction module is specifically used to: input the point cloud data of the target object into the three-dimensional target object detection algorithm module for processing, and generate a three-dimensional rectangular box surrounding the target object; use the center point of the bottom surface of the three-dimensional rectangular box as the center of the sphere, and use the bottom surface as the center section of the hemispherical space to construct a hemispherical space surrounding the target object.
[0195] In one possible implementation, the determination module 1003 is specifically used to: convert the position information in the posture of each image frame into three-dimensional space to obtain the position value of each image frame in the three-dimensional space; and divide the image frames whose position values fall into the corresponding spatial positions into corresponding spatial regions according to the position values of each image frame and the spatial positions of each spatial region in the three-dimensional space.
[0196] In a possible embodiment, the device also includes a deletion module, which is used to: when multiple supplementary key frames are determined, construct a spherical space for each supplementary key frame with a second preset distance as the sphere radius and a position value of each supplementary key frame as the sphere center position; for the spherical space of any supplementary key frame, delete all non-key frames in the spherical space of any supplementary key frame except itself.
[0197] In a possible implementation, the screening module 1002 is specifically configured to: input multiple image frames into the simultaneous positioning and mapping algorithm module for image processing, and output a processing result, where the processing result includes key frames screened out from the multiple image frames.
[0198] In one possible embodiment, the device also includes an association module, which is used to: establish an association relationship between the supplementary key frames and the key frames, and the association relationship represents the spatial correlation between the image frames in the three-dimensional space; the reconstruction module 1004 is specifically used to: perform three-dimensional reconstruction based on the association relationship, the supplementary key frames and the key frames to obtain a three-dimensional reconstructed model of the target object.
[0199] In a possible implementation, the multiple image frames are image frames obtained by capturing images by circling the target object with the target object as the center.
[0200] In one possible embodiment, the association module is further used to: perform image similarity calculation on any two image frames that have established an association relationship to obtain an image similarity score; for any two image frames corresponding to an image similarity score greater than a similarity threshold, based on the spatial relationship between the two image frames in three-dimensional space, if the spatial relationship does not meet a preset spatial relationship condition, cancel the established association relationship between the two image frames.
[0201] In one possible embodiment, the spatial relationship is represented by the absolute value of the difference between the index values of the spatial regions corresponding to the two image frames, and the index value is used to represent the position sequence of the corresponding spatial region in multiple spatial regions. The spatial relationship condition includes that the spatial relationship is less than or equal to a preset index threshold.
[0202] The three-dimensional reconstruction device provided in the embodiment of the present application can be used to implement the technical solution of the three-dimensional reconstruction method in any of the above embodiments of the present application. Its implementation principle and technical effects are similar, and will not be repeated here in this embodiment.
[0203] Figure 11 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application is shown in FIG. Figure 11 As shown, the electronic device of this embodiment may include: at least one processor 1101; and a memory 1102 communicatively connected to the at least one processor; wherein the memory 1102 stores instructions that can be executed by the at least one processor 1101, and the instructions are executed by the at least one processor 1101 so that the electronic device executes a method as in any of the above embodiments.
[0204] Optionally, the memory 1102 may be independent or integrated with the processor 1101 .
[0205] The implementation principle and technical effects of the electronic device provided in this embodiment can be found in the aforementioned embodiments and will not be described in detail here.
[0206] An embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer-executable instructions. When a processor executes the computer-executable instructions, the method described in any of the above embodiments is implemented.
[0207] An embodiment of the present application further provides a computer program product, including a computer program, which implements the method described in any of the aforementioned embodiments when executed by a processor.
[0208] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules is merely a logical function division. In actual implementation, other division methods may be used. For example, multiple modules may be combined or integrated into another system, or some features may be ignored or not implemented.
[0209] The integrated modules implemented in the form of software functional modules can be stored in a computer-readable storage medium. The software functional modules stored in a storage medium include a number of instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute some of the steps of the methods described in various embodiments of the present application.
[0210] It should be understood that the above-mentioned processor can be a processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), etc. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the application can be directly implemented as being executed by a hardware processor, or can be executed by a combination of hardware and software modules in the processor. The memory may include random access memory (RAM) and may also include non-volatile memory (NVM), such as at least one disk storage, and can also be a USB flash drive, a mobile hard disk, a read-only memory, a magnetic disk, or an optical disk.
[0211] The aforementioned storage medium may be implemented by any type of volatile or nonvolatile memory device, or a combination thereof, such as static random-access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The storage medium may be any available medium that can be accessed by a general-purpose or special-purpose computer.
[0212] An exemplary storage medium is coupled to a processor, such that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be an integral part of the processor. The processor and storage medium can be located in an application-specific integrated circuit. Of course, the processor and storage medium can also exist as discrete components in an electronic device or a host control device.
[0213] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.
[0214] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0215] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better embodiment. Based on this understanding, the technical solution of this application, or the part that contributes to the existing technology, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in each embodiment of this application.
[0216] The above are only preferred embodiments of the present application and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the description and drawings of this application, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.
[0217] Those skilled in the art will readily conceive of other implementations of the embodiments of the present application after considering the specification and practicing the application disclosed herein. This application is intended to cover any variations, uses, or adaptations of the embodiments of the present application, which follow the general principles of the embodiments of the present application and include common knowledge or customary technical means in the art that are not disclosed in the embodiments of the present application.
[0218] It should be understood that the embodiments of the present application are not limited to the precise structures described above and shown in the drawings, and various modifications and changes can be made without departing from the scope thereof. The scope of the embodiments of the present application is limited only by the appended claims.
Claims
1. A three-dimensional reconstruction method, characterized in that: The method comprises: Acquire multiple image frames of a target object to be three-dimensionally reconstructed; Screening out key frames from the multiple image frames, wherein the key frames are used to construct a three-dimensional reconstruction model; Determining, among the non-key frames in the plurality of image frames, some of the non-key frames as supplementary key frames, wherein a first distance between any of the supplementary key frames and the key frame closest to it in the three-dimensional space of the three-dimensional reconstructed model is greater than a first preset distance, and the supplementary key frames are used to assist in constructing the three-dimensional reconstructed model; Three-dimensional reconstruction is performed based on the supplementary key frame and the key frame to obtain a three-dimensional reconstructed model of the target object.
2. The method according to claim 1, characterized in that The determining of some non-key frames as supplementary key frames includes: Dividing the plurality of image frames into a plurality of spatial regions, wherein the plurality of spatial regions are set according to the spatial position of the target object in the three-dimensional space; Calculating a second distance between any of the non-key frames and any of the key frames in each of the plurality of spatial regions; A supplementary key frame is determined in each non-key frame where the second distance is greater than the first preset distance.
3. The method according to claim 2, characterized in that Before dividing the plurality of image frames into a plurality of spatial regions, the method further includes: Constructing a hemispherical space surrounding the target object according to the point cloud data of the target object; The hemispherical space is divided into a plurality of hemispherical primitives, and the plurality of hemispherical primitives correspond to the plurality of space regions respectively.
4. The method according to claim 3, characterized in that The step of constructing a hemispherical space surrounding the target object based on the point cloud data of the target object includes: The point cloud data of the target object is input into a three-dimensional target object detection algorithm module for processing to generate a three-dimensional rectangular frame surrounding the target object; A hemispherical space surrounding the target object is constructed with the center point of the bottom surface of the three-dimensional rectangular frame as the center of the sphere and the bottom surface as the center cross section of the hemispherical space.
5. The method according to claim 2, characterized in that The dividing the plurality of image frames into a plurality of spatial regions comprises: Converting the position information of the pose of each image frame into the three-dimensional space to obtain the position value of each image frame in the three-dimensional space; According to the position value of each image frame and the spatial position of each spatial region in the three-dimensional space, the image frames whose position values fall into the corresponding spatial position are divided into the corresponding spatial regions.
6. The method according to claim 5, characterized in that The method further comprises: When multiple supplementary key frames are determined, a spherical space of each supplementary key frame is constructed with the second preset distance as the sphere radius and the position value of each supplementary key frame as the sphere center position; For the spherical space of any supplementary key frame, other non-key frames in the spherical space of the supplementary key frame except the supplementary key frame itself are deleted.
7. The method according to any one of claims 1 to 6, characterized in that The step of selecting a key frame from the plurality of image frames comprises: The multiple image frames are input into a simultaneous positioning and mapping algorithm module for image processing, and a processing result is output. The processing result includes key frames screened out from the multiple image frames.
8. The method according to any one of claims 2 to 6, characterized in that: The method further comprises: Establishing an association relationship between the supplementary key frame and the key frame, wherein the association relationship represents a spatial correlation between image frames in the three-dimensional space; The performing three-dimensional reconstruction based on the supplementary key frame and the key frame to obtain a three-dimensional reconstructed model of the target object includes: Three-dimensional reconstruction is performed based on the association relationship, the supplementary key frame, and the key frame to obtain a three-dimensional reconstructed model of the target object.
9. The method according to claim 8, characterized in that The multiple image frames are image frames obtained by capturing images by circling the target object with the target object as the center.
10. The method according to claim 9, characterized in that After establishing the association relationship between the supplementary key frame and the key frame, the method further includes: Calculate the image similarity of any two image frames that have established an association relationship to obtain an image similarity score; For any two image frames corresponding to an image similarity score greater than a similarity threshold, based on the spatial relationship of the two image frames in the three-dimensional space, if the spatial relationship does not meet a preset spatial relationship condition, the established association relationship between the two image frames is released.
11. The method according to claim 10, characterized in that The spatial relationship is represented by the absolute value of the difference between the index values of the spatial regions corresponding to the two image frames, and the index value is used to represent the position sequence of the corresponding spatial regions in the multiple spatial regions. The spatial relationship condition includes that the spatial relationship is less than or equal to a preset index threshold.
12. A three-dimensional reconstruction device, characterized in that: The device comprises: An acquisition module, configured to acquire multiple image frames of a target object to be three-dimensionally reconstructed; a screening module, configured to screen out key frames from the plurality of image frames, wherein the key frames are used to construct a three-dimensional reconstruction model; a determination module configured to determine, among the non-key frames in the plurality of image frames, some non-key frames as supplementary key frames, wherein a first distance between any of the supplementary key frames and its closest key frame in the three-dimensional space of the three-dimensional reconstructed model is greater than a first preset distance, and the supplementary key frames are used to assist in constructing the three-dimensional reconstructed model; A reconstruction module is used to perform three-dimensional reconstruction based on the supplementary key frame and the key frame to obtain a three-dimensional reconstructed model of the target object.
13. An electronic device, characterized in that: include: memory and processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory, so that the processor performs the method according to any one of claims 1 to 11.
14. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 11 when executed by a processor.
15. A computer program product, characterized in that The invention comprises a computer program, which implements the method according to any one of claims 1 to 11 when the computer program is executed by a processor.
Citation Information
Patent Citations
Extracting method for disordered image key frame
CN105913096A
Three-dimensional reconstruction method and device, system and storage medium
CN111133477A
Data acquisition method, device and equipment and storage medium
CN111402412A
Scale reduction method and system, three-dimensional reconstruction method and system, storage medium and equipment
CN111402429A
Laser monocular vision fusion positioning mapping method in dynamic scene
CN113345018A