Three-dimensional reconstruction method, device, equipment, medium and program product
By selecting keyframes and determining supplementary keyframes, the problem of balancing quality and efficiency in 3D reconstruction in existing technologies is solved, achieving high-quality and efficient 3D reconstruction results.
Patent Information
- Application Number
- CN202511149544.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-18
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2045-08-18
AI Technical Summary
Existing technologies struggle to balance high quality and high efficiency in 3D reconstruction, resulting in either low reconstruction quality or low efficiency.
By selecting keyframes and identifying supplementary keyframes, a 3D reconstruction model is constructed. Non-keyframes that are sufficiently far from the keyframes in 3D space are selected as supplementary keyframes to assist in model construction.
It achieves both high quality and high efficiency in 3D reconstruction, avoids noise interference caused by repeated image frames, and improves the quality and efficiency of the reconstruction model.
Smart Images

Figure CN120726258B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and in particular to a three-dimensional reconstruction method, apparatus, device, medium, and program product. Background Technology
[0002] 3D reconstruction is a technique for recovering the 3D structure of a scene from environmental data collected by sensors. For example, it involves reconstructing the 3D structure of an environment or objects within it using 2D images captured by a camera at different locations or times. 3D reconstruction based on images can also be called visual 3D reconstruction.
[0003] After reconstructing a target object into a 3D model using various visual 3D reconstruction methods, a 3D reconstruction model of the target object can be obtained. For example, 2D images of the target object can be acquired, resulting in multiple image frames; then, by solving the target object in these image frames into 3D space, a 3D reconstruction model of the target object can be obtained. However, the 3D reconstruction model reconstructed using existing methods often differs significantly from the target object, resulting in low 3D reconstruction quality. Furthermore, when using other existing methods to improve reconstruction quality, additional data from more types of sensors is required, leading to lower efficiency in 3D reconstruction.
[0004] Therefore, existing technologies suffer from the problem of not being able to simultaneously achieve high 3D reconstruction quality and high 3D reconstruction efficiency. Summary of the Invention
[0005] This application provides a three-dimensional reconstruction method, apparatus, device, medium, and program product, which are used to efficiently reconstruct a high-quality three-dimensional reconstruction model by balancing high three-dimensional reconstruction quality and high three-dimensional reconstruction efficiency when performing visual three-dimensional reconstruction of a target object.
[0006] In a first aspect, embodiments of this application provide a three-dimensional reconstruction method, the method comprising: acquiring multiple image frames of a target object to be reconstructed in three dimensions; selecting key frames from the multiple image frames, the key frames being used to construct a three-dimensional reconstruction model; determining some non-key frames as supplementary key frames from among the non-key frames in the multiple image frames, wherein a first distance between any supplementary key frame and its nearest key frame in the three-dimensional space of the three-dimensional reconstruction model is greater than a first preset distance, the supplementary key frames being used to assist in constructing the three-dimensional reconstruction model; and performing three-dimensional reconstruction based on the supplementary key frames and the key frames to obtain a three-dimensional reconstruction model of the target object.
[0007] In one possible implementation, determining some non-keyframes as supplementary keyframes includes: dividing multiple image frames into multiple spatial regions, the multiple spatial regions being set according to the spatial position of the target object in three-dimensional space; calculating a second distance between any non-keyframe and any keyframe in each of the multiple spatial regions; and determining supplementary keyframes among the non-keyframes whose second distance is greater than a first preset distance.
[0008] In one possible implementation, before dividing multiple image frames into multiple spatial regions, the method further includes: constructing a hemispherical space surrounding the target object based on the point cloud data of the target object; dividing the hemispherical space into multiple hemispherical primitives, each of the multiple hemispherical primitives corresponding to multiple spatial regions.
[0009] In one possible implementation, constructing a hemispherical space surrounding the target object based on the point cloud data of the target object includes: inputting the point cloud data of the target object into a three-dimensional target object detection algorithm module for processing to generate a three-dimensional rectangular box surrounding the target object; and constructing a hemispherical space surrounding the target object with the center point of the bottom surface of the three-dimensional rectangular box as the center of the sphere and the bottom surface as the cross-section of the center of the hemispherical space.
[0010] In one possible implementation, dividing multiple image frames into multiple spatial regions includes: converting the position information in the pose of each image frame into three-dimensional space to obtain the position value of each image frame in three-dimensional space; and dividing the image frames whose position values fall into the corresponding spatial regions into the corresponding spatial regions based on the position values of each image frame and the spatial position of each spatial region in three-dimensional space.
[0011] In one possible implementation, the method further includes: when multiple supplementary keyframes are determined, constructing a spherical space for each supplementary keyframe with a second preset distance as the sphere radius and the position value of each supplementary keyframe as the sphere center position; and deleting other non-keyframes except itself from the spherical space of any supplementary keyframe.
[0012] In one possible implementation, selecting keyframes from multiple image frames includes: inputting multiple image frames into a simultaneous positioning and mapping algorithm module for image processing, outputting the processing result, which includes the keyframes selected from the multiple image frames.
[0013] In one possible implementation, the method further includes: establishing a correlation between supplementary keyframes and keyframes, the correlation characterizing the spatial correlation between image frames in three-dimensional space; and performing three-dimensional reconstruction based on the supplementary keyframes and keyframes to obtain a three-dimensional reconstruction model of the target object, including: performing three-dimensional reconstruction based on the correlation, supplementary keyframes and keyframes to obtain a three-dimensional reconstruction model of the target object.
[0014] In one possible implementation, multiple image frames are image frames obtained by capturing images around a target object.
[0015] In one possible implementation, after establishing a correlation between supplementary keyframes and keyframes, the method further includes: calculating image similarity for any two image frames with established correlation to obtain an image similarity score; for any two image frames corresponding to an image similarity score greater than a similarity threshold, based on the spatial relationship between the two image frames in three-dimensional space, if the spatial relationship does not meet a preset spatial relationship condition, deactivating the established correlation between the two image frames.
[0016] In one possible implementation, the spatial relationship is represented by the absolute value of the difference between the index values of the corresponding spatial regions of two image frames. The index value is used to represent the position sequence of its corresponding spatial region in multiple spatial regions. The spatial relationship condition includes that the spatial relationship is less than or equal to a preset index threshold.
[0017] Secondly, embodiments of this application provide a three-dimensional reconstruction apparatus, comprising: an acquisition module for acquiring multiple image frames of a target object to be reconstructed in three dimensions; a filtering module for filtering key frames from the multiple image frames, the key frames being used to construct a three-dimensional reconstruction model; a determination module for determining some non-key frames among the multiple image frames as supplementary key frames, wherein a first distance between any supplementary key frame and its nearest key frame in the three-dimensional space of the three-dimensional reconstruction model is greater than a first preset distance, the supplementary key frames being used to assist in constructing the three-dimensional reconstruction model; and a reconstruction module for performing three-dimensional reconstruction based on the supplementary key frames and the key frames to obtain a three-dimensional reconstruction model of the target object.
[0018] In one possible implementation, the determining module is specifically used to: divide multiple image frames into multiple spatial regions, the multiple spatial regions being set according to the spatial position of the target object in three-dimensional space; calculate a second distance between any non-key frame and any key frame in each of the multiple spatial regions; and determine supplementary key frames among the non-key frames where the second distance is greater than a first preset distance.
[0019] In one possible implementation, the device further includes a construction module for: constructing a hemispherical space surrounding the target object based on point cloud data of the target object; and dividing the hemispherical space into multiple hemispherical primitives, each of which corresponds to a multiple spatial region.
[0020] In one possible implementation, the construction module is specifically used to: process the point cloud data of the target object into the 3D target object detection algorithm module to generate a 3D rectangular box surrounding the target object; and construct a hemispherical space surrounding the target object with the center point of the bottom surface of the 3D rectangular box as the center of the sphere and the bottom surface as the cross-section of the hemispherical space.
[0021] In one possible implementation, the determining module is specifically used to: convert the position information in the pose of each image frame into three-dimensional space to obtain the position value of each image frame in three-dimensional space; and, based on the position value of each image frame and the spatial position of each spatial region in three-dimensional space, divide the image frames whose position values fall into the corresponding spatial positions into the corresponding spatial regions.
[0022] In one possible implementation, the device further includes a deletion module, which is configured to: when multiple supplementary keyframes are determined, construct a spherical space for each supplementary keyframe with a second preset distance as the sphere radius and the position value of each supplementary keyframe as the sphere center position; and delete other non-keyframes in the spherical space of any supplementary keyframe except itself.
[0023] In one possible implementation, the filtering module is specifically used to: input multiple image frames into the simultaneous positioning and mapping algorithm module for image processing, and output the processing results, which include keyframes selected from the multiple image frames.
[0024] In one possible implementation, the device further includes an association module, which is used to: establish an association relationship between supplementary keyframes and keyframes, the association relationship representing the spatial correlation between image frames in three-dimensional space; the reconstruction module is specifically used to: perform three-dimensional reconstruction based on the association relationship, supplementary keyframes and keyframes to obtain a three-dimensional reconstruction model of the target object.
[0025] In one possible implementation, multiple image frames are image frames obtained by capturing images around a target object.
[0026] In one possible implementation, the association module is further configured to: calculate the image similarity between any two image frames that have established an association relationship, and obtain an image similarity score; for any two image frames corresponding to an image similarity score greater than a similarity threshold, based on the spatial relationship between the two image frames in three-dimensional space, if the spatial relationship does not meet the preset spatial relationship conditions, terminate the established association relationship between the two image frames.
[0027] In one possible implementation, the spatial relationship is represented by the absolute value of the difference between the index values of the corresponding spatial regions of two image frames. The index value is used to represent the position sequence of its corresponding spatial region in multiple spatial regions. The spatial relationship condition includes that the spatial relationship is less than or equal to a preset index threshold.
[0028] Thirdly, embodiments of this application provide an electronic device, including: a memory and a processor; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory, causing the processor to perform the first aspect and / or various possible implementations of the first aspect as described above.
[0029] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the first aspect and / or various possible implementations of the first aspect.
[0030] Fifthly, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the first aspect and / or various possible implementations of the first aspect.
[0031] The 3D reconstruction method, apparatus, device, medium, and program products provided in this application embodiment, after acquiring multiple image frames of a target object, first select key frames for constructing a 3D reconstruction model. Since the number of key frames is small and the image frames participating in 3D reconstruction are relatively sparse, this method further identifies some non-key frames as supplementary key frames to assist in constructing the 3D reconstruction model. These supplementary key frames are non-key frames whose first distance in the 3D space of the 3D reconstruction model is greater than a first preset distance from the nearest key frame. Therefore, the distance between these supplementary key frames and the key frames is sufficiently large, resulting in low redundancy of the image information provided by the supplementary key frames and the key frames. The supplementary key frames can provide image information helpful for 3D reconstruction. Therefore, when performing 3D reconstruction based on supplementary key frames and key frames, the number of image frames participating in the reconstruction is neither too few nor too many, and noise interference is not introduced between image frames due to repeated image information. Thus, this method can efficiently reconstruct a high-quality 3D reconstruction model, achieving a balance between high 3D reconstruction quality and high 3D reconstruction efficiency. Attached Figure Description
[0032] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0033] Figure 1 A flowchart illustrating the three-dimensional reconstruction method provided in the embodiments of this application. Figure 1 ;
[0034] Figure 2 A schematic diagram of a target object with multifaceted similarity provided in an embodiment of this application;
[0035] Figure 3 A schematic diagram of the three-dimensional point cloud data provided in the embodiments of this application;
[0036] Figure 4 The flow chart of the three-dimensional reconstruction method provided in the embodiments of this application Figure 2 ;
[0037] Figure 5 A schematic diagram of a hemispherical space provided in an embodiment of this application;
[0038] Figure 6 A schematic diagram of polar angles and orientation angles provided for embodiments of this application;
[0039] Figure 7 A schematic diagram illustrating the cutting of a hemispherical space as provided in an embodiment of this application;
[0040] Figure 8 A schematic diagram of the spherical space provided in an embodiment of this application;
[0041] Figure 9 A schematic diagram illustrating the acquisition of images by bypassing a target object in an embodiment of this application;
[0042] Figure 10 This is a schematic diagram of the structure of the three-dimensional reconstruction device provided in the embodiments of this application;
[0043] Figure 11 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.
[0044] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation
[0045] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0046] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the technical solutions of this application comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0047] In this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design described as "exemplary" or "for example" in this application should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.
[0048] In the embodiments of this application, the use of terms such as "first" and "second" is to distinguish between identical or similar items that have essentially the same function and effect. For example, "first electronic device" and "second electronic device" are merely used to distinguish different electronic devices and do not limit their order of execution. Those skilled in the art will understand that the terms "first" and "second" do not limit the quantity or execution order, and that "first" and "second" do not necessarily imply that they are different.
[0049] In this application embodiment, "at least one" refers to one or more, and "more than one" refers to two or more. "And / or" describes the relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following associated objects have an "or" relationship.
[0050] For example, in visual 3D reconstruction, image frames of the target object are typically acquired using image acquisition devices such as cameras, or existing image frames of the target object can be used for 3D reconstruction. After obtaining the image frames, feature points are usually extracted from all image frames of the target object, and the feature points are matched based on the correlation of their descriptors. A descriptor can be understood as a vector, either manually defined or output end-to-end by a neural network, used to describe the regional information of a feature point. When factors such as viewing angle are similar, the regional information of an object in the environment is often quite similar across different image frames.
[0051] Then, camera calibration and attitude estimation are used to determine the camera's intrinsic and extrinsic parameters. Based on this information, triangulation is used to generate sparse point cloud data. After that, methods such as bundle adjustment are used to optimize the camera's relative pose and 3D point positions, thereby recovering the camera's position information in the environment and the environment's structural information.
[0052] In 3D reconstruction, among all the image frames of the target object, there are some that are crucial for recovering the camera's position and pose information in the environment; these image frames can be understood as keyframes. Besides being crucial for recovering position and pose information, keyframes are also essential for constructing and optimizing the 3D model.
[0053] Some 3D reconstruction methods use only keyframes for reconstruction, while others use all image frames. Since keyframes are typically few in number and sparsely distributed in 3D space, the resulting 3D model using only keyframes has significant errors, leading to low reconstruction quality. Other methods, however, involve numerous image frames of the target object. For example, cameras typically acquire images at a rate of 30 frames per second. Over a long period, this results in a large number of image frames, many of which contain significant repetition. Using all image frames for reconstruction not only increases computational complexity and runtime, significantly reducing efficiency, but also introduces noise due to the excessive repetition, further affecting reconstruction quality. Therefore, current methods cannot simultaneously achieve high 3D reconstruction quality and efficiency.
[0054] In view of this, embodiments of this application provide a 3D reconstruction method. After acquiring multiple image frames of a target object, the method first selects key frames for constructing a 3D reconstruction model. Since the number of key frames is relatively small and the image frames participating in 3D reconstruction are sparse, this method further identifies some non-key frames as supplementary key frames to assist in constructing the 3D reconstruction model. These supplementary key frames are non-key frames whose first distance in the 3D space of the 3D reconstruction model is greater than a first preset distance from the nearest key frame. Therefore, the distance between these supplementary key frames and the key frames is sufficiently large, resulting in low redundancy of image information provided by the supplementary key frames and the key frames. The supplementary key frames can provide image information helpful for 3D reconstruction. Therefore, when performing 3D reconstruction based on supplementary key frames and key frames, the number of image frames participating in the reconstruction is neither too few nor too many, and noise interference is not introduced between image frames due to repeated image information. Thus, this method can efficiently reconstruct a high-quality 3D reconstruction model, achieving a balance between high 3D reconstruction quality and high 3D reconstruction efficiency.
[0055] The technical solutions of this application will be described in detail below with reference to specific embodiments. The specific embodiments described below can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of this application will be described below with reference to the accompanying drawings.
[0056] The method provided in this application can be applied to applications, websites, or mini-programs with 3D reconstruction task processing capabilities. The 3D reconstruction task processing capability is implemented on the application, website, or mini-program. For example, a computer with a 3D reconstruction application deployed can implement the 3D reconstruction task processing capability by running the 3D reconstruction application. Another example is a terminal electronic device, such as a mobile phone, with a 3D reconstruction mini-program deployed, which can implement the 3D reconstruction task processing capability by running the 3D reconstruction mini-program.
[0057] Figure 1 A flowchart illustrating the three-dimensional reconstruction method provided in the embodiments of this application. Figure 1 The execution subject of this method can be an electronic device with corresponding data storage and computing capabilities, such as a computer, mobile phone, server, or server cluster. Figure 1 As shown, the method includes:
[0058] S101: Acquire multiple image frames of the target object to be reconstructed in 3D.
[0059] For example, the target object can be any object that needs to be reconstructed in 3D, such as a car, building, or other object. The number of target objects is not limited to a single object; there can be multiple objects. For example, the target objects can be all the buildings, pedestrians, trees, and vehicles in an environment scene.
[0060] When acquiring multiple image frames of the target object to be reconstructed in 3D, they can be captured by taking pictures with an image acquisition device such as a camera; they can also be read from a database by data transmission; or they can be acquired by other means, which are not limited in this application embodiment.
[0061] S102, select keyframes from multiple image frames, and use the keyframes to build a 3D reconstruction model.
[0062] For example, a keyframe is an image frame used to construct a 3D reconstruction model from all image frames. The number of keyframes can be one or more. It can be understood as selecting those image frames that have a significant impact on 3D reconstruction from multiple image frames as keyframes. The selected keyframes can provide sufficient image information during the 3D reconstruction process, while also controlling computational complexity and storage requirements.
[0063] There are several ways to filter keyframes. For example, keyframes can be obtained by filtering multiple image frames using predefined keyframe metrics, or by using some models or algorithms.
[0064] In one possible implementation, keyframes are selected from multiple image frames. Specifically, multiple image frames are input into the simultaneous positioning and mapping algorithm module for image processing, and the processing results are output, which include the keyframes selected from the multiple image frames.
[0065] For example, the Simultaneous Localization and Mapping (SLAM) algorithm can be understood as follows: a robot starts moving from an unknown location in an unknown environment, performs self-localization based on its position and a map during the movement, and simultaneously builds an incremental map based on its self-localization, thus achieving autonomous localization and navigation. Visual SLAM refers to performing SLAM tasks using images captured by a camera.
[0066] In the Simultaneous Localization and Mapping (SLAM) algorithm module, keyframes refer to the image frames selected for map construction and optimization. By inputting multiple image frames into the SLAM module for image processing, the module can output the selected keyframes. When selecting keyframes using the SLAM module, selection criteria typically include motion variations between image frames, disparity variations, scene variations, and feature point coverage.
[0067] In this embodiment of the application, by simultaneously inputting multiple image frames into the positioning and mapping algorithm module for image processing, key frames can be quickly selected, and the selected key frames have high value for 3D reconstruction. Therefore, this method can improve the quality and efficiency of key frame selection.
[0068] S103, among the non-key frames in multiple image frames, some non-key frames are determined as supplementary key frames. The first distance between any supplementary key frame and its nearest key frame in the three-dimensional space of the three-dimensional reconstruction model is greater than a first preset distance. The supplementary key frames are used to assist in the construction of the three-dimensional reconstruction model.
[0069] For example, in order to improve the density of image frames involved in 3D reconstruction, a portion of non-key frames can be identified in addition to the key frames in multiple image frames to assist in building the 3D reconstruction model. These non-key frames can be understood as supplementary key frames.
[0070] Furthermore, in order to reduce the impact on reconstruction quality and efficiency caused by the large repetition between the determined supplementary keyframes and keyframes, the keyframes can be determined based on the distance between the non-keyframes and keyframes.
[0071] For example, for any non-keyframe, calculate the distance between the position of the non-keyframe and the positions of each keyframe. Here, position can be understood as the positional information in the pose information of the image frame, which represents the relative position of the acquisition device such as the camera when acquiring the image frame. Distance can be understood as the difference between two positions, or a second distance. When there are multiple keyframes, multiple distances can be obtained for the non-keyframe. The smallest distance among these distances is taken as the first distance. It is then determined whether this first distance is greater than a first preset distance. If it is greater, the non-keyframe can be used as a supplementary keyframe; if it is less than or equal to the first preset distance, the non-keyframe cannot be used as a supplementary keyframe.
[0072] Based on this, non-keyframes are first filtered based on the condition that the first distance between any supplementary keyframe and its nearest keyframe in the three-dimensional space of the three-dimensional reconstruction model is greater than a first preset distance. Non-keyframes that meet the conditions for being supplementary keyframes can be filtered out. Then, one or more supplementary keyframes can be determined from these non-keyframes by selecting them through preset rules or by random selection.
[0073] For example, the three-dimensional space of a 3D reconstruction model can be understood as a virtual three-dimensional space pre-set when constructing the 3D reconstruction model. This three-dimensional space may have a preset coordinate system for determining position and distance, etc. The first preset distance can be a preset distance of arbitrary length, which can be set according to actual needs.
[0074] S104, based on the supplementary keyframes and the keyframes, perform 3D reconstruction to obtain the 3D reconstruction model of the target object.
[0075] For example, after selecting keyframes and determining supplementary keyframes, 3D reconstruction can be performed based on the supplementary keyframes and the keyframes to obtain a 3D reconstructed model of the target object. Any software or module capable of 3D reconstruction can be used for the 3D reconstruction; this application does not limit this. The 3D reconstructed model of the target object can be understood as a three-dimensional model of the target object, which can be a three-dimensional model of the same size as the target object or a three-dimensional model scaled down.
[0076] The 3D reconstruction method provided in this application, after acquiring multiple image frames of the target object, first filters out key frames for constructing a 3D reconstruction model. Since the number of key frames is small and the image frames participating in 3D reconstruction are sparse, this method further identifies some non-key frames as supplementary key frames to assist in constructing the 3D reconstruction model. These supplementary key frames are non-key frames whose first distance in the 3D space of the nearest key frame is greater than a first preset distance. Therefore, the distance between these supplementary key frames and the key frames is sufficiently large, resulting in low redundancy of the image information provided by the supplementary key frames and the key frames. The supplementary key frames can provide image information helpful for 3D reconstruction. Therefore, when performing 3D reconstruction based on supplementary key frames and key frames, the number of image frames participating in the reconstruction is neither too few nor too many, and noise interference is not introduced between image frames due to repeated image information. Thus, this method can efficiently reconstruct a high-quality 3D reconstruction model, achieving a balance between high 3D reconstruction quality and high 3D reconstruction efficiency.
[0077] For example, some existing visual 3D reconstruction algorithms, such as the open-source software COLMAP or Open Multi-View Stereo (OpenMVS) based on multi-view stereo reconstruction and motion structure recovery, have a relatively time-consuming part in the overall reconstruction process, such as feature matching of feature points.
[0078] One reason for the time-consuming nature of feature matching is that images are disordered in both time and space during 3D reconstruction. This disorder leads 3D reconstruction algorithms to commonly employ brute-force matching. 3D reconstruction often uses computationally complex feature points and descriptors, such as Scale-Invariant Feature Transform (SIFT). This is because 3D reconstruction needs to ensure the robustness of feature point matching across different scenes; factors such as exposure or affine transformations should not affect most of the matching results, otherwise 3D reconstruction is prone to failure. However, the high complexity of feature points and descriptors also brings computational pressure, potentially reducing reconstruction efficiency.
[0079] To improve feature point matching efficiency, some methods employ incremental matching based on the temporal order of images. This can be understood as using the temporal order of image acquisition by the camera as prior knowledge to help improve feature point matching efficiency. While incremental matching can speed up the matching process to some extent, it limits the association of the current image to images within adjacent time periods, leading to the loss of matching information. These factors degrade the quality of 3D reconstruction.
[0080] In addition, some methods also introduce other non-visual type sensors to provide prior information. Although this prior information can help improve matching efficiency to some extent, it will lead to additional steps such as extrinsic parameter calibration, time synchronization, and data alignment of different sensors due to the addition of additional sensors. These additional steps will consume a lot of time, increase the complexity of hardware and algorithms, and reduce the efficiency of 3D reconstruction to some extent.
[0081] Besides the inherent limitations of 3D reconstruction itself, the type of scene in 3D reconstruction also affects the matching speed and reconstruction accuracy. Some methods perform well for small-scale scenes with rich features and high feature point discrimination, but perform poorly for large scenes or scenes involving multiple similar objects. In-depth analysis reveals that large scenes are complex and involve large amounts of data, often with different areas sharing similar features. These similar structures can easily lead to structural breaks or misalignments in the reconstruction results. Large scenes also place higher demands on the computational complexity of algorithms or modules; even after optimization, many tasks still require several days or even weeks to complete reconstruction.
[0082] While reconstructing multi-faceted similar objects does not have the limitation of large scene data volume, it will encounter the problem of mismatch of multi-faceted similarity. Figure 2 A schematic diagram of a target object with multifaceted similarity provided in the embodiments of this application, such as... Figure 2 As shown, the target object is a car. Figure 2 The two image frames in the image were acquired by driving around the target object. It can be seen that the two sides of the car are highly similar. Based on actual tests and other data, the probability of mismatching similar surfaces during 3D reconstruction after driving around the object is relatively high, which will lead to a certain degree of degradation in the quality of the 3D reconstruction.
[0083] Figure 3 This is a schematic diagram of the three-dimensional point cloud data provided in the embodiments of this application, such as... Figure 3 As shown, this 3D point cloud data is the 3D point cloud data of the target object before obtaining the 3D reconstruction model. This 3D point cloud data is... Figure 2 The image shows the point cloud data obtained during the reconstruction of the target object. Before reconstruction, multiple image frames were captured by circling the target object with a camera, and this data was then used in 3D reconstruction software to obtain the 3D point cloud data. Figure 3 It is evident from the data that, due to the similarity between the left and right sides of the vehicle (especially the tires), the reconstructed 3D point cloud structure of the car exhibits layering. Therefore, when further 3D reconstruction is performed based on this point cloud data, the resulting 3D reconstruction model will also show layering, affecting the reconstruction quality.
[0084] Furthermore, conventional indoor and outdoor 3D reconstruction typically involves a data acquisition route that goes from the starting point to the end point and back again, requiring coverage of the scene from all angles to obtain relatively complete scene information. However, 3D reconstruction of objects around a central point involves the data acquisition route centered on the object, with the data acquisition personnel circling around it to collect image frames. This acquisition method results in a lot of information duplication in image frames. Excessive duplication reduces the efficiency of feature point matching and backend optimization using bundle adjustment. Moreover, excessive redundant information not only fails to improve the reconstruction quality but also introduces excessive noise, further degrading the reconstruction quality. Therefore, for 3D reconstruction of objects around a central point, it is necessary to filter the data to find an optimal number of image frames that ensures both reconstruction quality and speed.
[0085] Based on the above description, the three-dimensional reconstruction method provided in this application embodiment can be applied to scenarios of three-dimensional reconstruction of bypassed objects, and the reconstruction quality and efficiency of three-dimensional reconstruction of bypassed objects can be improved by processing such as filtering key frames and determining supplementary key frames.
[0086] In one possible implementation, the method further includes: establishing a correlation between supplementary keyframes and keyframes, wherein the correlation characterizes the spatial correlation between image frames in three-dimensional space; and performing three-dimensional reconstruction based on the supplementary keyframes and keyframes to obtain a three-dimensional reconstruction model of the target object, including: performing three-dimensional reconstruction based on the correlation, supplementary keyframes and keyframes to obtain a three-dimensional reconstruction model of the target object.
[0087] For example, the correlation relationship characterizes the spatial correlation between image frames in three-dimensional space. For instance, if the same objects can be observed through two image frames, then there is a certain co-view relationship between the two image frames, which also indicates that the two image frames have a certain spatial correlation in three-dimensional space. Using these two image frames, a three-dimensional reconstruction model within a certain range can be reconstructed and optimized.
[0088] One of the purposes of supplementing keyframes and establishing relationships between keyframes is to establish certain relationships based on the spatial correlation between these image frames in the early stages of 3D reconstruction. These relationships can serve as prior knowledge for 3D reconstruction to guide the 3D reconstruction process, thereby helping to speed up the 3D reconstruction and improve its quality.
[0089] Establishing a correlation between supplementary keyframes and keyframes can be achieved in various ways. For example, for any two image frames in the supplementary keyframe and keyframes, feature points are extracted and descriptors are represented. By matching the descriptors, the extracted feature points can be matched. If a certain number or region of feature points are matched, it can be determined that the two have spatial correlation. Furthermore, the degree of spatial correlation can be determined by the magnitude of the match, and then it can be determined whether to establish a correlation between the two based on the degree of spatial correlation.
[0090] Alternatively, relationships can be established in other ways. For example, the processing results output by the SLAM module can include a keyframe co-view between keyframes. This keyframe co-view can be understood as an expression of the common-view relationships within a graph data structure. Through the keyframe co-view and the common-view relationships between supplementary keyframes, further relationships can be established between supplementary keyframes and keyframes. Of course, other methods can also be used to establish relationships, which will not be elaborated upon here.
[0091] Furthermore, when performing 3D reconstruction based on supplementary keyframes and keyframes to obtain a 3D reconstruction model of the target object, it can be based on the correlation relationship, supplementary keyframes, and keyframes to obtain a 3D reconstruction model of the target object.
[0092] In the embodiments of this application, a correlation relationship representing the spatial correlation between image frames in three-dimensional space is established between each supplementary keyframe and multiple keyframes. This correlation relationship can provide effective prior knowledge for three-dimensional reconstruction, which helps to improve the quality and efficiency of three-dimensional reconstruction.
[0093] In one possible implementation, multiple image frames are obtained by capturing images around a target object.
[0094] For example, when the target object is a car, an image acquisition device such as a camera can be used to capture images of the car by circling it one or more times. Of course, the direction, speed, and image acquisition frequency are not limited, and image acquisition can be performed at any location in real space other than the ground. For example, images can be captured from the top or side of the car to obtain image frames.
[0095] In this application embodiment, the application scenario of 3D reconstruction is a scenario in which images are collected around a target object and 3D reconstruction is performed. In this scenario, the 3D reconstruction method of this application embodiment can improve the reconstruction efficiency and reconstruction quality of the 3D reconstruction of the bypassed object.
[0096] For example, as can be seen from the above description of the situation where mismatches are prone to occur in the 3D reconstruction of bypassed objects, mismatches may occur when establishing the association between supplementary keyframes and keyframes.
[0097] For example, if an image frame captured from the left side of a car's front and another captured from the right side of the car's front are extremely similar in content, feature point extraction and feature point descriptor matching might lead to the conclusion that the two images have a high degree of spatial correlation, potentially establishing an association between them. However, in reality, one image was captured from the left side of the car's front, and the other from the right side; their capture point locations differ significantly. They are merely similar in appearance, and their spatial correlation is actually low or nonexistent. Therefore, for image frames that have already established an association, it is necessary to identify and remove false matches, thus resolving the mismatched association.
[0098] In one possible implementation, after supplementing keyframes and establishing associations between keyframes, the method further includes: calculating image similarity for any two image frames with established associations to obtain image similarity scores; for any two image frames corresponding to image similarity scores greater than a similarity threshold, based on the spatial relationship between the two image frames in three-dimensional space, if the spatial relationship does not meet preset spatial relationship conditions, deactivating the established association between the two image frames.
[0099] For example, when calculating the image similarity between two image frames, any algorithm or neural network model capable of image similarity calculation can be used. The resulting image similarity score can be a predicted or evaluated value used to characterize the degree of similarity between the image content of the two image frames.
[0100] The similarity threshold can be any preset threshold used to determine the degree of similarity between image content. For example, for any two image frames that have established a relationship, after inputting them into the similarity neural network model, an image similarity score can be obtained. If the image similarity score is less than or equal to the similarity threshold, the relationship between the two can be retained; if the image similarity score is greater than the similarity threshold, it is further determined whether to remove the relationship between the two.
[0101] When determining whether to sever the association between two images, the spatial relationship between the two images in three-dimensional space can be used as a basis. If the spatial relationship does not meet the preset spatial relationship conditions, the established association between the two images can be severed.
[0102] The spatial relationship between two image frames in three-dimensional space can be a spatial positional relationship or a spatial distance relationship. A spatial relationship can characterize the spatial relationship between the acquisition locations of the two image frames. Preset spatial relationship conditions can be conditions that determine the degree of spatial relationship. For example, if the spatial relationship is the spatial distance between two image frames in three-dimensional space, then the spatial distance between the two image frames in three-dimensional space is calculated. The preset spatial relationship condition can be a preset distance threshold. If the calculated spatial distance is greater than the preset distance threshold, it indicates that the two image frames have a very high similarity, but their acquisition locations are far apart. In this case, it can be determined that the established association is due to a mismatch, and the established association between the two image frames can be terminated.
[0103] After removing all mismatched associations, the established associations are corrected, which improves the accuracy of the prior knowledge used for 3D reconstruction. This, in turn, improves the accuracy of the 3D reconstruction model and enhances the quality of 3D reconstruction.
[0104] In this embodiment, in a scenario where 3D reconstruction is performed around a target object, if the target object has multiple similar surfaces (such as the left and right sides of a car), the established association relationships may contain mismatches leading to erroneous associations. These erroneous associations can cause phenomena such as misalignment during 3D reconstruction, affecting the reconstruction quality. Therefore, this solution calculates image similarity and determines spatial relationship conditions for two image frames with high image similarity. If the spatial relationship between the two image frames does not meet the spatial relationship conditions, it indicates that although the two image frames are very similar, their acquisition positions differ significantly, and they are likely not images acquired from similar locations. Therefore, the established association between them is likely an erroneous association due to mismatch. In this case, removing the association between the two can avoid phenomena such as misalignment during subsequent 3D reconstruction using the association relationship, thereby improving the reconstruction quality when performing 3D reconstruction on a target object that is being bypassed.
[0105] Next, let's combine Figure 4 The three-dimensional reconstruction method provided in the embodiments of this application will be further described. Figure 4 The flow chart of the three-dimensional reconstruction method provided in the embodiments of this application Figure 2 ,like Figure 4 As shown, the method includes the following steps:
[0106] S401, Visual SLAM generates image poses and map points. S402, Object detection algorithm generates 3D bounding boxes. S403, Constructing a hemispherical space centered on the reconstructed target object and cutting it into multiple hemispherical primitives. S404, After cutting the hemispherical space, filtering the image frames assigned to the hemispherical primitives. S405, Establishing associations between image frames using a common view and location map. S406, Performing image similarity detection on the established associations of image frames and filtering associations based on the index values of the hemispherical primitives. Of course, Figure 4 The 3D reconstruction method shown can also be understood as some steps performed in the early stage of the 3D reconstruction process. After S406, steps such as S104 can be executed to complete the 3D reconstruction.
[0107] like Figure 4 As shown, an image stream consisting of multiple image frames can be acquired by circling the target object multiple times using a camera device. This image stream is then input into the visual SLAM algorithm module to generate the pose of each image and a global visual point cloud map. The visual SLAM algorithm module can be understood as the simultaneous localization and mapping algorithm module or SLAM module mentioned above.
[0108] Visual SLAM algorithms can include, for example, oriented fast and rotated brief (ORB) SLAM. ORB SLAM can be understood as an algorithm that uses image ORB features for association matching to perform SLAM tasks. Furthermore, visual SLAM is a real-time algorithm, and its computational cost for 3D reconstruction algorithms is negligible.
[0109] The acquired image stream is input into the visual SLAM algorithm, which generates the pose of each image frame (including position and orientation information in the algorithm coordinate system). In addition, it also filters out keyframes, keyframe co-views, non-keyframes and keyframe relationships, and 3D point cloud data centered on the target object.
[0110] For example, most SLAM algorithms follow this process: the front end estimates a rough initial pose for the current image frame through image feature matching, and then sends this estimated initial pose to the back end for optimization. The back end optimization module constructs a graph optimization problem using local map points and the poses of other keyframes. Typically, a camera generates 30 image frames per second, and each image frame needs its pose estimated. Each image frame is a regular image frame, used only to generate its own pose. The optimization of the SLAM back end depends on the pose graph formed by the image poses and the map points. Map generation is time-consuming, so it is not advisable to include every image frame in map point generation. Moreover, an overly dense pose graph would lead to excessive computational costs for subsequent graph optimization and bundle adjustment. As a real-time algorithm, SLAM filters regular image frames. These frames are selected as keyframes under time, space, and other conditions, and can be used to participate in map point generation and back end optimization.
[0111] The method provided in this application first processes the acquired image stream using a visual SLAM algorithm, outputting a co-view of camera pose trajectory, map points, and keyframes. This provides global temporal and spatial prior information. Moreover, the visual SLAM algorithm is a highly efficient real-time algorithm, and it does not affect the overall 3D reconstruction time efficiency. A visual SLAM algorithm is incorporated into the 3D reconstruction of bypassed objects. Visual SLAM generates co-view relationships to determine the association relationships of image frames. Keyframes in visual SLAM serve as the medium for associating global information; they participate in the generation of global map points and the optimization of the global pose map throughout the entire visual SLAM process. Therefore, this type of image frame can effectively provide global feature matching association relationships.
[0112] Since image frames captured by a single camera lack scale information, the corresponding monocular visual SLAM also lacks scale information. Here, scale information refers to the fact that the generated 3D information has no dimensions, making it impossible to determine the image location or the distance unit of the reconstructed object (e.g., whether it's meters or centimeters). Although scale information is absent, it can be addressed through relative proportions. Therefore, after obtaining the visual SLAM results, the dimensions surrounding the central target object can be determined first.
[0113] For example, 3D object detection algorithms can be used to process scale information. These algorithms can be any type of algorithm used to detect 3D objects, such as at least one of the following: End-to-End Learning for Point Cloud Based 3D Object Detection (VoxelNet), Fast Encoders for Object Detection from Point Clouds (PointPillars), and Sparsely Embedded Convolutional Detection (SECOND), or other 3D object detection algorithms used in autonomous driving technology.
[0114] 3D object detection algorithms can detect specific target objects, such as pedestrians, cars, or trees, within a 3D point cloud. Inputting the 3D point cloud map generated by the SLAM algorithm into the 3D object detection algorithm will generate corresponding 3D bounding boxes.
[0115] Assume the vertices of the 3D bounding box generated by the 3D object detection algorithm are A, B, C, D, E, F, G, and H. We can take the center point of the base rectangle ABCD as O. Using O as the center, and the product of the longer side BC of the base rectangle and the scaling factor α as the radius r, we construct a hemispherical space. Where: , where d bc This represents the side length of BC.
[0116] Based on this, a hemispherical space can be generated. Figure 5 This is a schematic diagram of the hemispherical space provided in the embodiments of this application, as shown below. Figure 5 As shown, the truck is the target object. A hemispherical space can be constructed based on the target object through the base of a 3D rectangular frame. The value of the scaling factor α determines the size of the hemispherical space, and the scaling factor α can be preset according to actual needs. For example, if image frames at different distances are needed, the scaling factor α can be taken as a larger value, making the hemispherical space larger. The scaling factor α can be, for example, between 1.5 and 2. In addition, this hemispherical space generation method uses the relative position distance of point clouds as a basis, so it is not necessary to obtain the absolute scale to complete the generation of the corresponding hemispherical space.
[0117] The method provided in this application uses a 3D target object detection algorithm to obtain the outer 3D bounding box of the point cloud of the target reconstructed object, and uses the length of the 3D bounding box to generate a hemispherical space. It can divide the hemispherical space into multiple hemispherical primitives in a way that mimics the latitude and longitude of the earth. Using the divided spherical primitives as units, a spherical space is generated at the position center of the key frame. Excessive image frames in the spherical space are deleted. At the same time, several supplementary image frames are randomly selected from the non-key frames to generate a spherical space of the same size. This can delete redundant non-key frames in the spherical space, which can efficiently complete the image filtering and make the distribution of the filtered images more uniform.
[0118] After generating the hemispherical space, it can be divided in several ways. For example, it can be divided randomly, or horizontally and vertically. This application provides a division method based on Earth's latitude and longitude.
[0119] Figure 6 A schematic diagram of polar angles and orientation angles provided for embodiments of this application, such as... Figure 6 As shown, a polar angle can be defined. Let P1 be a vector from the origin O on the sphere. The angle between the polar angle and the xy-axis plane. Alternatively, the polar angle can be defined as the angle between the polar angle and the z-axis. In this embodiment, the hemispherical space is divided using the Earth's latitude and longitude system, so the polar angle is defined as the angle between the polar angle and the xy-axis plane, similar to latitude in latitude and longitude. The polar angle ranges from [0, π / 2]. Defining the direction angle. Let P1 be a vector from the origin O on the sphere. The angle between the projection of the coordinate axes onto the xy-plane and the x-axis is similar to longitude in latitude and longitude, ranging from [0, 2π]. The polar angle and direction angle can form a spherical coordinate system. According to the definition rules of polar angle and direction angle, the spherical coordinate system can be converted to a rectangular coordinate system using the following formula:
[0120]
[0121] in, Let x represent the radius of the sphere in the hemispherical space; x, y, and z are the coordinates of point P1 in the rectangular coordinate system, respectively.
[0122] Figure 7 This is a schematic diagram of cutting a hemispherical space according to an embodiment of this application, as shown below. Figure 7 As shown, the hemispherical space is divided using a method referencing Earth's latitude and longitude, i.e., polar angular resolution. (Latitude direction), directional angular resolution (Longitude direction) can divide the hemispherical space along the polar angle and direction angle directions into indivual, It can be calculated using the following formula:
[0123]
[0124] Each segmented spatial region can be understood as a hemispherical primitive, composed of the intersection of the sphere's center O, the tangent lines in the polar direction, and the tangent lines in the directional direction. For example, Figure 7 The intersection points P1, P2, P3, P4 and the center O of the sphere together constitute a spherical element V1.
[0125] In this embodiment, all hemispherical primitives can be sorted by their index values in sequence. For example, a certain hemispherical primitive can be represented as V. mn , where m and n represent the index values of the direction angle and polar angle, respectively.
[0126] After completing the hemispherical spatial segmentation, multiple image frames can be classified into the corresponding hemispherical primitives.
[0127] We can first determine whether the image frame is within the hemispherical space. This only requires calculating the image frame's position in the hemispherical Cartesian coordinate system. The distance to the origin (center of the ball) is sufficient, as follows:
[0128]
[0129] in, Indicates position The distance to the origin. When If the image frame is not within the hemispherical space, it can be deleted. Then, the polar angle and orientation angle of the image frame in the spherical coordinate system can be calculated using the following formula:
[0130]
[0131] in, Indicates the polar angle; Indicates the direction angle; Represents the inverse cosine function; This represents the arctangent function. Given the polar angle resolution and orientation angle resolution, the index value of the hemispherical primitive to which the image frame belongs can be obtained, specifically calculated using the following formula:
[0132]
[0133] in, This indicates the rounding down sign. After calculating the index value of the corresponding hemispherical primitive for each image frame, a list of image frames within each hemispherical primitive can be obtained, which can be represented as:
[0134]
[0135] in, V represents mn A list of image frames for the hemispherical primitive. Represents an image frame ; Represents an image frame .
[0136] Next, the image frames within each hemispherical primitive can be filtered. Since a camera typically generates 30 image frames per second, the image frames in space are relatively dense. Visual SLAM keyframes are sparsely distributed image frames that have undergone rigorous filtering. Some existing 3D reconstruction methods directly use keyframes as input image frames for 3D reconstruction. However, to improve optimization efficiency and speed, visual SLAM often sets the keyframe images to be very sparse in both time and space, resulting in a insufficient number of map points directly used for 3D reconstruction to meet the requirements. Map points can be understood as points on the surface of the 3D reconstruction model.
[0137] Therefore, based on the keyframes, the number of image frames participating in 3D reconstruction can be appropriately increased to ensure that the generated map points meet the requirements, while avoiding excessive increases in image frames that would significantly lengthen the 3D reconstruction time. All keyframes within the hemispherical primitive can be used for 3D reconstruction, and can be represented as:
[0138]
[0139] in, This represents a list of keyframes. Indicates keyframe keyl; This represents the keyframe keyq.
[0140] Furthermore, a spherical space with radius rI can be generated centered on the keyframe position. Figure 8 A schematic diagram of the spherical space provided in the embodiments of this application, as shown below. Figure 8 As shown, assume the keyframe's position in three-dimensional space is... The position of the non-keyframe is So the distance between non-keyframes and keyframes The distance can be calculated using the following formula. This can also be understood as the second distance:
[0141]
[0142] You can set the first preset distance d λ dλ It can be equal to a radius of rI. When At this time, it indicates the current non-keyframe I. i Within the spherical space of the keyframe, candidate sequences can be eliminated.
[0143] like Figure 8 As shown, image frame I1 and image frame I4 are respectively in keyframe I key1 and keyframe I key2 Since these two image frames are located in a spherical space, they are not considered as candidate image frames for 3D reconstruction. To count the remaining non-key frames, multiple image frames can be randomly selected, requiring that the first distance between these image frames is greater than a first preset distance d. λ These non-keyframes are then supplementary keyframes.
[0144] For example, these selected supplementary keyframes can also generate their respective sphere spaces. A similar method can be used to check whether other non-keyframes are within the sphere spaces of these supplementary keyframes, and then delete other non-keyframes from the sphere spaces except for themselves. For example, Figure 8 After generating the corresponding spherical space for image frame I2, image frame I3 was deleted through calculation and judgment.
[0145] For example, the above process of deleting non-keyframes in the sphere space can be repeated multiple times until all supplementary image frames generate corresponding sphere spaces.
[0146] The method provided in this application considers matching in special locations. These special locations refer to situations where some associated keyframes are located very far apart, even at different camera angles, but the acquired images are similar, such as the area around the front of a vehicle. Image frames with some established associations may have similar image information, which can lead to mismatches. The method utilizes image similarity detection algorithms from deep learning to detect image similarity, while simultaneously using the index value of hemispherical primitives in the orientation angle to avoid erroneous removal of image feature associations that are close in location, ensuring that image frames that should not be matched are correctly de-associated.
[0147] After image selection is complete, the next step is to establish relationships between supplementary keyframes and keyframes. Establishing relationships serves two purposes: firstly, it accelerates feature matching in 3D reconstruction, and secondly, it reduces mismatches of features with multiple similar faces.
[0148] For example, the backend of visual SLAM generates a shared view of keyframes. This shared view is constructed based on the number of map points observed between keyframes. That is, if a certain number of map points can be observed between keyframes, then two keyframes are shared and interconnected in the graph data structure. Typically, when shooting multiple times around a target object, one keyframe may be associated with more than a dozen other keyframes.
[0149] Furthermore, in the visual SLAM process, feature matching at the front end often calculates the initial pose by associating the previous image frame with local map points. These local map points are generated and associated from keyframes. Therefore, algorithms like ORB SLAM select a reference keyframe I when performing initial pose estimation for the current image frame. key_ref The reference keyframe is the keyframe with the highest degree of co-view with the current image frame, meaning it has the largest number of observed common map points. Therefore, in this embodiment, the correlation of feature matching can be obtained through the keyframe co-view.
[0150] For example, keyframe co-views generated by visual SLAM can obtain keyframes. Keyframes with shared viewing relationships, and keyframes The associations formed after feature matching can be represented as follows:
[0151]
[0152] in, Keyframe A list of relationships, and All keyframes are keyframes that establish a relationship with them.
[0153] If a supplementary keyframe corresponds to a reference keyframe that is a keyframe or If a keyframe is listed, then that supplementary keyframe is related to the keyframe. If there is a shared view relationship, the supplementary keyframe can be added to the list of related keyframes. This process is repeated until all supplementary keyframes are traversed, ultimately yielding the keyframe image. The list of feature matching associations is as follows:
[0154]
[0155] in, and The symbols "etc." indicate supplementary keyframes.
[0156] By performing the above operations on all keyframes, the relationships between each supplementary keyframe and each keyframe can be established.
[0157] While using the co-view relationship of keyframes to obtain the feature matching association can avoid mismatches of multi-faceted similar images in most cases, mismatches can still occur in some special cases.
[0158] Figure 9 A schematic diagram illustrating the acquisition of images by bypassing a target object is provided for an embodiment of this application, such as... Figure 9 As shown, image frames were taken from both sides of the front of the car, such as image I. key1 I key2 I key3 and I key4 Images with similar or even identical content, but captured from different locations, might be associated using a keyframe located at the center of the vehicle's front. Following this method, they might end up in the same keyframe's association list, leading to mismatched images and potentially causing misalignment.
[0159] To avoid this situation, the image similarity of the image frames and the hemispherical primitives to which they belong can be considered together.
[0160] In one possible implementation, spatial relationships are represented by the absolute value of the difference between the index values of the corresponding spatial regions of two image frames. The index value is used to represent the position sequence of its corresponding spatial region in multiple spatial regions. The spatial relationship condition includes that the spatial relationship is less than or equal to a preset index threshold.
[0161] For example, image similarity scores can be calculated using image similarity detection algorithms such as Vision Transformer (VIT) or Self-Distillation with No Labels (DINO). For each... Calculate the image similarity score S between all pairs of image frames in the list. ij A similarity threshold S can be preset. λ When S ij >S λ If two image frames are highly similar, it is necessary to consider removing their association.
[0162] However, some adjacent images are often highly similar, and these images should inherently be associated with each other, such as... Figure 9 I in key3 and I key4 Therefore, on the one hand, the similarity threshold S needs to be set... λ To improve efficiency, we should strive to avoid mistakenly deleting correct associations. On the other hand, we also need to use spatial relationships to constrain the force of de-association.
[0163] This application embodiment uses the index values of the hemispherical primitives to which the two image frames belong. and Spatial relationships are characterized by the absolute value of the difference between them. The index value corresponding to the direction angle can be taken, and the absolute value of the difference between the index values can be calculated using the following formula:
[0164]
[0165] in, This represents the absolute value of the difference between index values. Represents an image frame The index value corresponding to the direction angle; Represents an image frame The index value corresponding to the direction angle.
[0166] when Greater than the preset index threshold When the positions of two image frames are far apart, the established association between these two similar image frames can be terminated.
[0167] In this embodiment, spatial relationship is represented by the absolute value of the difference between the index values of the corresponding spatial regions of two image frames. The spatial relationship condition includes that the spatial relationship is less than or equal to a preset index threshold. Based on this spatial relationship and spatial relationship condition, it is possible to effectively determine whether there is a mismatched association between two relatively similar image frames, thereby improving the accuracy of determining and executing the deassociation.
[0168] The final result is the image frames involved in the 3D reconstruction and the relationships between them. This makes it possible to globally correlate other image frames while ensuring computational efficiency during the 3D reconstruction process. Furthermore, this method does not restrict the type of the target object to be reconstructed. For example, if the target object is located in the center, the trajectory of the acquired image frames only needs to be around that target object. Therefore, it is applicable not only to cars but also to other objects.
[0169] The method provided in this application can solve the problem of mismatching when circumventing similar objects on multiple surfaces, thereby improving the speed and quality of 3D reconstruction. Based on visual SLAM, corresponding map points and image poses are generated. This prior information avoids mismatching of similar features at different locations. It accelerates feature matching and reduces redundant image information: by dividing the space around the target object into hemispherical voxel units (hemispherical primitives), image frames within the same voxel unit are filtered to reduce the total number of image frames participating in 3D reconstruction. Fewer image frames result in faster feature matching, and appropriately reducing redundant images also improves reconstruction quality. Therefore, this method provides a more efficient scheme for image frame filtering, matching relationships, and 3D reconstruction around objects.
[0170] In one possible implementation, before dividing multiple image frames into multiple spatial regions, the method further includes: constructing a hemispherical space surrounding the target object based on the point cloud data of the target object; dividing the hemispherical space into multiple hemispherical primitives, each of which corresponds to a multiple spatial region.
[0171] For example, in 3D reconstruction, constructing a hemispherical space centered on the object is more consistent with the specific implementation process than constructing a rectangular or cubic space. For instance, the lower half of the spherical space located on the plane where the target object is located may not have image frames captured in the corresponding real-world space; therefore, a hemisphere is more practical.
[0172] The advantage of hemispherical space is that any point on the surface of the hemisphere is equidistant from the center of the sphere. Therefore, hemispherical space can filter image frames whose acquisition positions are always within the radius of the hemisphere, and can filter out image frames acquired from acquisition positions outside the radius of the hemisphere. Cube or rectangular spaces cannot easily achieve the same filtering effect as hemispherical space.
[0173] For example, dividing the hemispherical space into multiple hemispherical primitives can be done by simulating the Earth's latitude and longitude in the above embodiments, which will not be elaborated here.
[0174] In this embodiment, a hemispherical space is more suitable for 3D reconstruction scenarios compared to spaces of other shapes. This solution constructs a hemispherical space surrounding the target object using point cloud data and then segments the hemispherical space to quickly divide it into multiple hemispherical primitives, achieving the goal of quickly and effectively setting up multiple spatial regions.
[0175] In one possible implementation, determining some non-keyframes as supplementary keyframes includes: dividing multiple image frames into multiple spatial regions, the multiple spatial regions being set according to the spatial position of the target object in three-dimensional space; calculating a second distance between any non-keyframe and any keyframe in each of the multiple spatial regions; and determining supplementary keyframes among the non-keyframes whose second distance is greater than a first preset distance.
[0176] For example, the multiple spatial regions can be spatial regions set according to the spatial position of the target object in three-dimensional space. The multiple spatial regions can be spatial regions of arbitrary shape, such as cubes, cuboids, cylinders, etc. For example, the multiple spatial regions can also be the spatial regions corresponding to the aforementioned multiple hemispherical primitives.
[0177] Within each of the multiple spatial regions, a second distance is calculated between any non-keyframe and any keyframe. Supplementary keyframes are then identified from among the non-keyframes whose second distance is greater than a first preset distance. This second distance is the distance between a non-keyframe and a keyframe, and can be obtained by calculating the difference in position in three-dimensional space. The second distance can be referenced to the relevant descriptions in the above embodiments.
[0178] If the determination of whether to designate a non-key frame as a supplementary key frame is made frame by frame according to multiple image frames, the calculation process during the determination is limited by the serial method, resulting in low determination efficiency.
[0179] In this embodiment, all image frames are divided into spatial regions, which can be quickly processed in multiple spatial regions to determine supplementary keyframes, thereby achieving the goal of parallel processing of multiple image frames and improving the efficiency of determining supplementary keyframes.
[0180] In one possible implementation, constructing a hemispherical space surrounding the target object based on the point cloud data of the target object includes: inputting the point cloud data of the target object into a 3D target object detection algorithm module for processing to generate a 3D bounding box surrounding the target object; and constructing a hemispherical space surrounding the target object with the center point of the bottom surface of the 3D bounding box as the center of the sphere and the bottom surface as the cross-section of the hemispherical space.
[0181] For example, the point cloud data of the target object can be obtained through the processing results of the SLAM module, or it can be acquired through a radar equidistant point cloud data sensor. The 3D target object detection algorithm module can be a software algorithm module that implements 3D target object detection. The specific implementation of the 3D target object detection algorithm, the 3D bounding box, and the hemispherical space surrounding the target object constructed through the bottom surface of the 3D bounding box can be referred to the descriptions in the above embodiments.
[0182] In this embodiment, a 3D target object detection algorithm is used to process the point cloud data, which can assign scale information and generate a 3D bounding box that conforms to the scale of the point cloud data and surrounds the target object. Constructing a hemispherical space surrounding the target object based on this 3D bounding box allows for the creation of a more accurate hemispherical space that better matches the size of the target object.
[0183] In one possible implementation, multiple image frames are divided into multiple spatial regions, including: converting the position information in the pose of each image frame into three-dimensional space to obtain the position value of each image frame in three-dimensional space; and dividing the image frames whose position values fall into the corresponding spatial regions into the corresponding spatial regions based on the position values of each image frame and the spatial position of each spatial region in three-dimensional space.
[0184] For example, the positional information in the pose of each image frame can be obtained by estimating the pose of the image frame, where the pose includes position and orientation. This positional information can be understood as the acquisition position of the acquisition device.
[0185] By using the transformation matrix between the pose coordinate system and the 3D spatial coordinate system, the position information in the pose of an image frame can be transformed into 3D space, obtaining the position value of each image frame in 3D space. Since each spatial region corresponds to a part of the spatial position in 3D space, the image frames whose position values fall into the corresponding spatial positions can be divided into the corresponding spatial regions based on the position values of each image frame and the spatial positions of each spatial region in 3D space.
[0186] In this embodiment, when multiple image frames are divided into multiple spatial regions, the position information in the pose of the image frames is first converted to three-dimensional space to achieve coordinate system unification. Then, based on the position value of each image frame and the spatial position of each spatial region in three-dimensional space, the multiple image frames can be orderly divided into multiple spatial regions. After division, the position information of the image frames will not be lost, thereby improving the accuracy of the determined first distance and improving the quality of determining supplementary keyframes.
[0187] In one possible implementation, the method further includes: when multiple supplementary keyframes are determined, constructing a spherical space for each supplementary keyframe with a second preset distance as the sphere radius and the position value of each supplementary keyframe as the sphere center position; and deleting other non-keyframes except itself from the spherical space of any supplementary keyframe.
[0188] For example, the second preset distance can be understood as a radius value preset for supplementary keyframes, which can be equal to the first preset distance. For instance, the second preset distance is equal to the first preset distance d. λIt can also be equal to the radius rI of the sphere space of the keyframe, etc.
[0189] The specific implementation of constructing the sphere space of each supplementary keyframe and deleting other non-keyframes in the sphere space of any supplementary keyframe (excluding itself) can be found in the relevant descriptions in the above embodiments, and will not be repeated here.
[0190] In this embodiment of the application, within the second preset distance range of the supplementary keyframe, there may be other non-keyframes. These non-keyframes may also be supplementary keyframes. If two closely spaced supplementary keyframes are used for 3D reconstruction, it may affect the reconstruction efficiency and reconstruction quality. Therefore, by deleting other non-keyframes in the sphere space of the supplementary keyframe except itself, unnecessary non-keyframes can be removed, which helps to further improve the reconstruction quality and efficiency.
[0191] Figure 10 This is a schematic diagram of the structure of the three-dimensional reconstruction device provided in the embodiments of this application, such as... Figure 10 As shown in the figure, this application embodiment provides a three-dimensional reconstruction device, which includes: an acquisition module 1001, used to acquire multiple image frames of a target object to be three-dimensionally reconstructed; a filtering module 1002, used to filter out key frames from the multiple image frames, the key frames being used to construct a three-dimensional reconstruction model; a determination module 1003, used to determine some non-key frames as supplementary key frames from each of the multiple image frames, wherein the first distance between any supplementary key frame and its nearest key frame in the three-dimensional space of the three-dimensional reconstruction model is greater than a first preset distance, the supplementary key frames being used to assist in constructing the three-dimensional reconstruction model; and a reconstruction module 1004, used to perform three-dimensional reconstruction based on the supplementary key frames and the key frames to obtain a three-dimensional reconstruction model of the target object.
[0192] In one possible implementation, the determining module 1003 is specifically used to: divide multiple image frames into multiple spatial regions, the multiple spatial regions being set according to the spatial position of the target object in three-dimensional space; calculate a second distance between any non-key frame and any key frame in each of the multiple spatial regions; and determine supplementary key frames in each non-key frame where the second distance is greater than a first preset distance.
[0193] In one possible implementation, the device further includes a construction module for: constructing a hemispherical space surrounding the target object based on point cloud data of the target object; and dividing the hemispherical space into multiple hemispherical primitives, each of which corresponds to a multiple spatial region.
[0194] In one possible implementation, the construction module is specifically used to: process the point cloud data of the target object into the 3D target object detection algorithm module to generate a 3D rectangular box surrounding the target object; and construct a hemispherical space surrounding the target object with the center point of the bottom surface of the 3D rectangular box as the center of the sphere and the bottom surface as the cross-section of the hemispherical space.
[0195] In one possible implementation, the determining module 1003 is specifically used to: convert the position information in the pose of each image frame into three-dimensional space to obtain the position value of each image frame in three-dimensional space; and classify the image frames whose position values fall into the corresponding spatial positions into the corresponding spatial regions according to the position values of each image frame and the spatial positions of each spatial region in three-dimensional space.
[0196] In one possible implementation, the device further includes a deletion module, which is configured to: when multiple supplementary keyframes are determined, construct a spherical space for each supplementary keyframe with a second preset distance as the sphere radius and the position value of each supplementary keyframe as the sphere center position; and delete other non-keyframes in the spherical space of any supplementary keyframe except itself.
[0197] In one possible implementation, the filtering module 1002 is specifically used to: input multiple image frames into the simultaneous positioning and mapping algorithm module for image processing, and output the processing results, which include keyframes selected from the multiple image frames.
[0198] In one possible implementation, the device further includes an association module, which is used to: establish an association relationship between supplementary keyframes and keyframes, the association relationship representing the spatial correlation between image frames in three-dimensional space; the reconstruction module 1004 is specifically used to: perform three-dimensional reconstruction based on the association relationship, supplementary keyframes and keyframes to obtain a three-dimensional reconstruction model of the target object.
[0199] In one possible implementation, multiple image frames are image frames obtained by capturing images around a target object.
[0200] In one possible implementation, the association module is further configured to: calculate the image similarity between any two image frames that have established an association relationship, and obtain an image similarity score; for any two image frames corresponding to an image similarity score greater than a similarity threshold, based on the spatial relationship between the two image frames in three-dimensional space, if the spatial relationship does not meet the preset spatial relationship conditions, terminate the established association relationship between the two image frames.
[0201] In one possible implementation, the spatial relationship is represented by the absolute value of the difference between the index values of the corresponding spatial regions of two image frames. The index value is used to represent the position sequence of its corresponding spatial region in multiple spatial regions. The spatial relationship condition includes that the spatial relationship is less than or equal to a preset index threshold.
[0202] The three-dimensional reconstruction apparatus provided in this application embodiment can be used to execute the technical solution of the three-dimensional reconstruction method in any of the above embodiments of this application. Its implementation principle and technical effect are similar, and will not be described again in this embodiment.
[0203] Figure 11 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application, such as... Figure 11 As shown, the electronic device of this embodiment may include: at least one processor 1101; and a memory 1102 communicatively connected to the at least one processor; wherein the memory 1102 stores instructions that can be executed by the at least one processor 1101, and the instructions are executed by the at least one processor 1101 to cause the electronic device to perform the method as described in any of the above embodiments.
[0204] Optionally, the memory 1102 can be either standalone or integrated with the processor 1101.
[0205] The implementation principle and technical effects of the electronic device provided in this embodiment can be found in the foregoing embodiments, and will not be repeated here.
[0206] This application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the method described in any of the foregoing embodiments.
[0207] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the methods described in any of the foregoing embodiments.
[0208] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules may be combined or integrated into another system, or some features may be ignored or not executed.
[0209] The integrated modules implemented as software functional modules described above can be stored in a computer-readable storage medium. These software functional modules, stored in a storage medium, include several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute some steps of the methods described in the various embodiments of this application.
[0210] It should be understood that the aforementioned processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. A general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the application can be directly manifested as being executed by a hardware processor, or executed by a combination of hardware and software modules within the processor. The memory may include random access memory (RAM), and may also include non-volatile memory (NVM), such as at least one disk storage device, and may also be a USB flash drive, external hard drive, read-only memory, disk, or optical disc, etc.
[0211] The aforementioned storage media can be implemented from any type of volatile or non-volatile storage device or a combination thereof, such as Static Random-Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The storage media can be any available medium accessible to general-purpose or special-purpose computers.
[0212] An exemplary storage medium is coupled to a processor, enabling the processor to read information from and write information to the storage medium. Alternatively, the storage medium can be an integral part of the processor. The processor and storage medium can reside within an application-specific integrated circuit (ASIC). Alternatively, the processor and storage medium can exist as discrete components within an electronic device or host device.
[0213] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0214] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0215] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0216] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.
[0217] Other embodiments of the present application will readily occur to those skilled in the art upon consideration of the specification and practice of the application disclosed herein. This application is intended to cover any variations, uses, or adaptations of the embodiments of this application that follow the general principles of the embodiments of this application and include common knowledge or customary techniques in the art not disclosed in the embodiments of this application.
[0218] It should be understood that the embodiments of this application are not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from their scope. The scope of the embodiments of this application is limited only by the appended claims.
Claims
1. A three-dimensional reconstruction method, characterized in that, The method includes: Acquire multiple image frames of the target object to be reconstructed in 3D; Keyframes are selected from the plurality of image frames, and the keyframes are used to construct a three-dimensional reconstruction model. In each of the non-key frames in the plurality of image frames, some non-key frames are determined as supplementary key frames. The first distance between any supplementary key frame and its nearest key frame in the three-dimensional space of the three-dimensional reconstruction model is greater than a first preset distance. The supplementary key frame is used to assist in the construction of the three-dimensional reconstruction model. Based on the supplementary keyframes and the keyframes, a three-dimensional reconstruction is performed to obtain a three-dimensional reconstruction model of the target object; The step of identifying some non-keyframes as supplementary keyframes includes: The plurality of image frames are divided into a plurality of spatial regions, which are set according to the spatial position of the target object in the three-dimensional space; In each of the plurality of spatial regions, calculate the second distance between any non-keyframe and any keyframe; In each non-keyframe where the second distance is greater than the first preset distance, a supplementary keyframe is determined; Before dividing the plurality of image frames into the plurality of spatial regions, the method further includes: Construct a hemispherical space surrounding the target object based on the point cloud data of the target object; The hemispherical space is divided into multiple hemispherical primitives, each of which corresponds to a different spatial region.
2. The method according to claim 1, characterized in that, The step of constructing a hemispherical space surrounding the target object based on the point cloud data of the target object includes: The point cloud data of the target object is input into the 3D target object detection algorithm module for processing to generate a 3D bounding box surrounding the target object. Using the center point of the bottom surface of the three-dimensional rectangular frame as the center of the sphere, and the bottom surface as the central section of the hemispherical space, a hemispherical space is constructed to surround the target object.
3. The method according to claim 1, characterized in that, The step of dividing the multiple image frames into multiple spatial regions includes: The position information in the pose of each image frame is converted into the three-dimensional space to obtain the position value of each image frame in the three-dimensional space; Based on the position value of each image frame and the spatial position of each spatial region in the three-dimensional space, the image frames whose position values fall into the corresponding spatial positions are divided into the corresponding spatial regions.
4. The method according to claim 3, characterized in that, The method further includes: Given multiple supplementary keyframes, a spherical space for each supplementary keyframe is constructed with the second preset distance as the sphere radius and the position value of each supplementary keyframe as the sphere center position. For any supplementary keyframe in the sphere space, delete all other non-keyframes in the sphere space of that supplementary keyframe except itself.
5. The method according to any one of claims 1-4, characterized in that, The step of filtering keyframes from the plurality of image frames includes: The multiple image frames are input into the simultaneous positioning and mapping algorithm module for image processing, and the processing results are output, including keyframes selected from the multiple image frames.
6. The method according to any one of claims 1-4, characterized in that, The method further includes: A correlation is established between the supplementary keyframe and the keyframe, and the correlation characterizes the spatial correlation between image frames in the three-dimensional space; The step of performing 3D reconstruction based on the supplementary keyframes and the keyframes to obtain a 3D reconstructed model of the target object includes: Based on the aforementioned correlation, the supplementary keyframes, and the keyframes, a 3D reconstruction is performed to obtain a 3D reconstruction model of the target object.
7. The method according to claim 6, characterized in that, The multiple image frames are image frames obtained by taking the target object as the center and capturing images by circling around the target object.
8. The method according to claim 7, characterized in that, After establishing the association between the supplementary keyframe and the keyframe, the method further includes: For any two image frames that have established a relationship, calculate the image similarity to obtain an image similarity score; For any two image frames corresponding to an image similarity score greater than a similarity threshold, if the spatial relationship between the two image frames in the three-dimensional space does not meet a preset spatial relationship condition, the established association between the two image frames is terminated.
9. The method according to claim 8, characterized in that, The spatial relationship is characterized by the absolute value of the difference between the index values of the spatial regions corresponding to the two image frames. The index value is used to represent the position sequence of the corresponding spatial region in the plurality of spatial regions. The spatial relationship condition includes that the spatial relationship is less than or equal to a preset index threshold.
10. A three-dimensional reconstruction device, characterized in that, The device includes: The acquisition module is used to acquire multiple image frames of the target object to be reconstructed in 3D. A filtering module is used to filter out key frames from the plurality of image frames, the key frames being used to construct a three-dimensional reconstruction model; The determination module is used to determine some non-key frames as supplementary key frames in each of the multiple image frames. The first distance between any supplementary key frame and its nearest key frame in the three-dimensional space of the three-dimensional reconstruction model is greater than a first preset distance. The supplementary key frame is used to assist in the construction of the three-dimensional reconstruction model. The reconstruction module is used to perform three-dimensional reconstruction based on the supplementary keyframes and the keyframes to obtain a three-dimensional reconstruction model of the target object; The determining module is specifically used to: divide the plurality of image frames into a plurality of spatial regions, wherein the plurality of spatial regions are set according to the spatial position of the target object in the three-dimensional space; calculate a second distance between any non-key frame and any key frame in each of the plurality of spatial regions; and determine supplementary key frames in each of the non-key frames in which the second distance is greater than the first preset distance. The device further includes a construction module, which is used to: construct a hemispherical space surrounding the target object based on the point cloud data of the target object; and divide the hemispherical space into multiple hemispherical primitives, each of which corresponds to a multiple spatial region.
11. An electronic device, characterized in that, include: Memory and processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory, causing the processor to perform the method as described in any one of claims 1-9.
12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method as described in any one of claims 1-9.
13. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the method according to any one of claims 1-9.
Citation Information
Patent Citations
Data acquisition method, device and equipment and storage medium
CN111402412A