Three-dimensional reconstruction method, device, system, equipment, medium, product and vehicle cleaning system
By establishing a common relationship between sensing devices and screening three-dimensional spatial points, the existing three-dimensional reconstruction methods are solved, and the problem of low efficiency and accuracy in complex scenarios and large-scale data processing is achieved, achieving a more efficient and robust three-dimensional reconstruction process.
Patent Information
- Application Number
- CN202510526197.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-04-25
AI Technical Summary
The existing three-dimensional reconstruction methods have problems with high computational complexity and low registration accuracy when dealing with complex scenarios and large-scale data. Especially when environmental changes are severe or severe occlusions are severe, efficiency and accuracy are difficult to guarantee.
By collecting multiple two-dimensional pictures and three-dimensional point clouds, a common-visual relationship between sensing devices is established, observable three-dimensional spatial points are selected, and a mapping relationship between two-dimensional pictures and three-dimensional spatial points is established, thereby correlating different two-dimensional pictures and improving the robustness and efficiency of three-dimensional reconstruction.
It realizes improving image matching efficiency in complex scenarios and large-scale data processing, enhancing the robustness of the three-dimensional reconstruction process, reducing mismatch, improving the accuracy of cross-view angle matching and the redundancy and reliability of three-dimensional point cloud data.
Smart Images

Figure CN120047628A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of three-dimensional reconstruction technology, and in particular, to a three-dimensional reconstruction method, device, system, equipment, medium, product, and vehicle cleaning system. Background Art
[0002] In the current scenario of big data processing, during the three-dimensional reconstruction process, images collected from multiple perspectives need to be effectively associated to extract corresponding features and perform matching, and then an accurate three-dimensional model is generated.
[0003] Currently, the commonly used image association techniques mainly include brute-force matching and incremental matching in time series. However, the brute-force matching method has too high computational complexity. When processing a large number of images, matching all feature points of each pair of images will result in significant time consumption and resource waste; while the incremental matching in time series can gradually update the model when processing dynamic scenes and has relatively high efficiency, but this method depends on the matching results of the previous frame and may have large errors. For example, when the environment changes violently or the occlusion is severe, the incremental method is easily severely affected, resulting in a decrease in registration accuracy. At the same time, this method usually requires strong sequence consistency, increasing the requirements for the imaging device and the shooting order, and limiting the flexibility of applications.
[0004] Therefore, the existing three-dimensional reconstruction methods still face many challenges when dealing with complex scenes and large-scale data, and there is an urgent need for a more efficient and more robust matching strategy. Summary of the Invention
[0005] Embodiments of the present application provide a three-dimensional reconstruction method, device, system, equipment, medium, product, and vehicle cleaning system, so as to achieve the effect of accelerating image matching and having better robustness during three-dimensional reconstruction.
[0006] In a first aspect, an embodiment of the present application provides a three-dimensional reconstruction method, which is characterized by including:
[0007] In response to a processing instruction, collect multiple two-dimensional pictures and three-dimensional point clouds of a target object; the three-dimensional point clouds are collected by a first sensing device at multiple acquisition times with different poses, the two-dimensional pictures are collected by a second sensing device at multiple acquisition times with different poses, and the first sensing device and the second sensing device are relatively fixedly arranged;
[0008] Based on the acquisition times corresponding to each three-dimensional space point in the three-dimensional point cloud, establish a co-visibility relationship between the acquisition times of the first sensing device;
[0009] For any two-dimensional image corresponding to a collection moment, in combination with the co-visibility relationship between the collection moments, three-dimensional space points that can be observed by the second sensing device in the pose corresponding to the current collection moment are screened to establish a mapping relationship between the two-dimensional image and the three-dimensional space points;
[0010] Based on the mapping relationship, an association relationship between the two-dimensional images is established, and three-dimensional reconstruction of at least part of the target object is performed based on the association relationship.
[0011] In a possible implementation manner, the establishing the co-visibility relationship between the collection moments of the first sensing device based on the collection moments corresponding to the three-dimensional space points in the three-dimensional point cloud includes:
[0012] Based on the collection moments corresponding to the three-dimensional space points in the three-dimensional point cloud, a first value of the same three-dimensional space points collected at each collection moment and any other collection moment is determined;
[0013] When the first value reaches a preset threshold, it is determined that the two collection moments are co-visible, thereby establishing the co-visibility relationship between the collection moments.
[0014] In a possible implementation manner, the screening the three-dimensional space points that can be observed by the second sensing device in the pose corresponding to the current collection moment for any two-dimensional image corresponding to a collection moment, in combination with the co-visibility relationship between the collection moments, to establish the mapping relationship between the two-dimensional image and the three-dimensional space points includes:
[0015] For any two-dimensional image corresponding to a collection moment, in combination with the co-visibility relationship between the collection moments, at least one collection moment co-visible with the current collection moment is determined;
[0016] The three-dimensional space points corresponding to the current collection moment and the three-dimensional space points corresponding to at least one collection moment co-visible with the current collection moment are screened to obtain the three-dimensional space points that can be observed by the second sensing device in the pose corresponding to the current collection moment, so as to establish the mapping relationship between the two-dimensional image and the three-dimensional space points.
[0017] In a possible implementation manner, the screening the three-dimensional space points corresponding to the current collection moment and the three-dimensional space points corresponding to at least one collection moment co-visible with the current collection moment to obtain the three-dimensional space points that can be observed by the second sensing device in the pose corresponding to the current collection moment includes:
[0018] The three-dimensional space points corresponding to the current collection moment and the three-dimensional space points corresponding to at least one collection moment co-visible with the current collection moment are subjected to three-dimensional conversion to obtain the image three-dimensional points corresponding to the three-dimensional space points;
[0019] Filter the three-dimensional image points based on the depth information of the three-dimensional image points.
[0020] Perform two-dimensional conversion on the filtered three-dimensional image points to obtain the image position points corresponding to each of the filtered three-dimensional image points.
[0021] Obtain the size information of the two-dimensional picture corresponding to the current acquisition moment, and filter the image position points based on the size information.
[0022] Take the three-dimensional space points corresponding to the filtered image position points as the three-dimensional space points that the second sensing device can observe in the pose corresponding to the current acquisition moment.
[0023] In a possible implementation manner, the filtering of the three-dimensional image points based on the depth information of the three-dimensional image points includes:
[0024] Based on a pre-set depth range, retain the three-dimensional image points that meet the depth range, and remove the three-dimensional image points that do not meet the depth range to filter the three-dimensional image points.
[0025] In a possible implementation manner, the size information includes the size range corresponding to the two-dimensional picture.
[0026] The filtering of the image position points based on the size information includes:
[0027] Retain the image position points that meet the size range, and remove the image position points that do not meet the size range to filter the image position points.
[0028] In a possible implementation manner, after obtaining the size information of the two-dimensional picture corresponding to the current acquisition moment, it further includes:
[0029] Divide the two-dimensional picture into multiple image regions according to the size range corresponding to the two-dimensional picture.
[0030] The retaining of the image position points that meet the size range and the removing of the image position points that do not meet the size range to filter the image position points includes:
[0031] Retain the image position points that meet the size range, and remove the image position points that do not meet the size range.
[0032] For the retained image position points, determine the second value of the image position points falling in each of the image regions.
[0033] For each of the image regions, when the second value reaches a preset threshold, randomly remove the image position points that fall within the current image region so that the total number of image position points within the current image region is less than or equal to the preset threshold, in order to screen the image position points.
[0034] In a possible implementation manner, establishing the association relationships between the two-dimensional pictures based on the mapping relationship includes:
[0035] Traverse each of the two-dimensional pictures, and based on the mapping relationship between the two-dimensional picture and the three-dimensional space point, determine the third value of the same three-dimensional space point corresponding to each of the two-dimensional pictures and any other two-dimensional picture;
[0036] For each of the two-dimensional pictures, perform an association degree ranking on the other two-dimensional pictures based on the third value corresponding to any other two-dimensional picture;
[0037] Based on the result of the association degree ranking, establish the association relationships between the two-dimensional pictures.
[0038] In a second aspect, an embodiment of the present application provides a three-dimensional reconstruction device, which is characterized by including:
[0039] An acquisition module, configured to acquire multiple two-dimensional pictures and three-dimensional point clouds of a target object in response to a processing instruction; the three-dimensional point clouds are acquired by a first sensing device at multiple acquisition times with different poses, the two-dimensional pictures are acquired by a second sensing device at multiple acquisition times with different poses, and the first sensing device and the second sensing device are relatively fixedly arranged;
[0040] A establishment module, configured to establish a co-visibility relationship between the acquisition times of the first sensing device based on the acquisition times corresponding to the three-dimensional space points in the three-dimensional point cloud;
[0041] A screening module, configured to, for the two-dimensional picture corresponding to any acquisition time, combine the co-visibility relationships between the acquisition times, screen out the three-dimensional space points that the second sensing device can observe at the pose corresponding to the current acquisition time, in order to establish the mapping relationship between the two-dimensional picture and the three-dimensional space point;
[0042] A reconstruction module, configured to establish the association relationships between the two-dimensional pictures based on the mapping relationship, and perform at least partial three-dimensional reconstruction of the target object based on the association relationships.
[0043] In a third aspect, an embodiment of the present application provides a three-dimensional reconstruction system, including a first sensing device, a second sensing device, and a processing device; the processing device is respectively connected to the first sensing device and the second sensing device;
[0044] The first sensing device and the second sensing device are relatively fixedly arranged;
[0045] The processing device is configured to control the movement of the first sensing device and the second sensing device, and adopt the three-dimensional reconstruction method in the first aspect above and / or various possible implementation manners of the first aspect, control the first sensing device to collect the three-dimensional point cloud corresponding to the target object parked in the target area, control the second sensing device to collect the two-dimensional picture of the target object, and perform three-dimensional reconstruction of at least part of the target object based on the three-dimensional point cloud and the two-dimensional picture.
[0046] In a fourth aspect, an embodiment of the present application provides a vehicle washing system, including the three-dimensional reconstruction system, a cleaning head and a controller in the second aspect above and / or various possible implementation manners of the second aspect;
[0047] The controller is respectively connected to the processing device in the three-dimensional reconstruction system and the cleaning head;
[0048] The cleaning head is connected to at least one receiving groove;
[0049] The controller is configured to control the movement of the cleaning head and control the cleaning head to spray the cleaning liquid in the receiving groove onto the vehicle body based on the result of three-dimensional reconstruction of at least part of the vehicle in the target area by the processing device.
[0050] In a fifth aspect, an embodiment of the present application provides an electronic device, including: a memory, a processor;
[0051] The memory stores computer execution instructions;
[0052] The processor executes the computer execution instructions stored in the memory, so that the processor executes the first aspect above and / or various possible implementation manners of the first aspect.
[0053] In a sixth aspect, an embodiment of the present application provides a computer-readable storage medium, in which computer execution instructions are stored, and when the computer execution instructions are executed by a processor, they are used to implement the first aspect above and / or various possible implementation manners of the first aspect.
[0054] In a seventh aspect, an embodiment of the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the first aspect above and / or various possible implementation manners of the first aspect.
[0055] The 3D reconstruction method, device, system, equipment, medium, product, and vehicle washing system provided by the embodiments of the present application can associate the 3D space points observed from multiple co-viewing perspectives of the first sensing device at each acquisition time to the same 2D image by establishing the co-viewing relationship of the perspectives corresponding to the first sensing device, thereby increasing the number of 3D space points corresponding to a single 2D image. When facing occluded areas or low reflectivity areas, more complete and more 3D space points corresponding to the current perspective can be filtered out. Moreover, by using the mapping relationship between 3D space points and 2D images to establish the association relationship between different 2D images, false matches can be filtered out and the accuracy of cross-perspective matching can be improved. Through the association between different 2D images, shared 3D space points can be found in multiple 2D images, thereby improving the redundancy and reliability of 3D point cloud data and enhancing the robustness of the 3D reconstruction process. Description of the Drawings
[0056] The drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present application and used together with the specification to explain the principles of the present application.
[0057] Figure 1 Schematic structural diagram of the 3D reconstruction system provided by the present application;
[0058] Figure 2 Schematic diagram of the relative setting positions of the first sensing device and the second sensing device in the 3D reconstruction system provided by the present application;
[0059] Figure 3 Schematic flowchart of the 3D reconstruction method provided by the present application;
[0060] Figure 4 Schematic diagram of the structure of a voxel unit in the voxel structure for storing 3D space points provided by the present application;
[0061] Figure 5 Schematic diagram of the relationship between the lidar coordinate system corresponding to the first sensing device, the camera coordinate system corresponding to the second sensing device, and the world coordinate system provided by the present application;
[0062] Figure 6 Schematic diagram of the conversion for converting the filtered image 3D points in the camera coordinate system to the corresponding image position points in the plane of the 2D image in the 3D reconstruction method provided by the present application;
[0063] Figure 7 Schematic diagram of the segmentation of a 2D image divided into multiple image regions in the 3D reconstruction method provided by the present application;
[0064] Figure 8 Schematic structural diagram of the 3D reconstruction device provided by the present application;
[0065] Figure 9 Schematic structural diagram of the vehicle washing system provided for this application;
[0066] Figure 10 Schematic structural diagram of the electronic device provided for this application.
[0067] Through the above-mentioned drawings, specific embodiments of this application have been shown, and there will be more detailed descriptions hereinafter. These drawings and textual descriptions are not intended to limit the scope of the concept of this application in any way, but to illustrate the concept of this application to those skilled in the art by referring to specific embodiments. Detailed implementation manners
[0068] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with this application. On the contrary, they are merely examples of devices and methods consistent with some aspects of this application as detailed in the appended claims.
[0069] The three-dimensional reconstruction method provided by the embodiments of this application can be applied to, for example, Figure 1 the three-dimensional reconstruction system 100 shown as follows. The three-dimensional reconstruction system 100 includes a first sensing device 110, a second sensing device 120, and a processing device 130; the processing device 130 is respectively connected to the first sensing device 110 and the second sensing device 120.
[0070] The first sensing device 110 and the second sensing device 120 are relatively fixedly arranged;
[0071] When receiving a processing instruction input by a user, the processing device 130 can generate corresponding acquisition signals and send them to the first sensing device 110 and the second sensing device 120, so as to control the first sensing device 110 and the second sensing device 120 to respectively collect three-dimensional point clouds and two-dimensional pictures of a target object placed in a target area according to a preset acquisition frequency.
[0072] As an example, when the three-dimensional reconstruction system 100 is applied to an automated vehicle washing scenario, the target object can be a vehicle. Correspondingly, the three-dimensional point cloud of the target object refers to the body point cloud of the vehicle, and the two-dimensional picture refers to the body picture of the vehicle.
[0073] As shown in Figure 2 the following, it should be noted that the first sensing device 110 and the second sensing device 120 can be fixedly arranged on a mobile base. The first sensing device 110 can be a laser sensor, and the second sensing device 120 can be a camera.
[0074] The mobile base can be arranged on an orbit set around the target area, and the orbit can be a sliding orbit set on the ground or a sliding orbit hoisted on the ceiling or set on the wall surface.
[0075] While the processing device 130 generates an acquisition signal and sends it to the first sensing device 110 and the second sensing device 120, the processing device 130 can generate a movement signal and send it to the mobile base to control the mobile base to slide on the orbit, so as to move the first sensing device 110 and the second sensing device 120 arranged on the mobile base around the target area to collect two-dimensional pictures of the target object at various angles and a complete three-dimensional point cloud of the target object in the target area.
[0076] Furthermore, after the processing device 130 obtains two-dimensional pictures and three-dimensional point clouds of the target object at various angles, the processing device 130 can perform at least partial three-dimensional reconstruction on the target object in the target area based on the two-dimensional pictures and three-dimensional point clouds, so that the processing device 130 can accurately identify the shape and size of the target object. As an example, in the scenario of automatic vehicle washing, the processing device 130 reconstructs the vehicle body model, which can provide a reference for the customized car washing process, so that during the automatic washing process, the cleaning head can accurately approach the vehicle body, and appropriate cleaning intensity and methods can be adopted for different parts of the vehicle body to avoid damaging the vehicle surface.
[0077] Among them, the processing device 130 can be a terminal or a server. The terminal includes but is not limited to various personal computers, laptop computers, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, etc. The portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The server can be implemented by an independent server or a server cluster composed of multiple servers. The terminal and the server can be directly or indirectly connected through wired or wireless communication methods, such as through a network connection.
[0078] In one embodiment, a three-dimensional reconstruction method is provided. In this embodiment, an example is given in which the three-dimensional reconstruction method is applied to the processing device in the above three-dimensional reconstruction system. As Figure 2 shown, the three-dimensional reconstruction method includes:
[0079] Step 302, in response to a processing instruction, collect multiple two-dimensional pictures and three-dimensional point clouds of the target object; the three-dimensional point cloud is collected by the first sensing device at multiple acquisition times with different poses, and the two-dimensional pictures are collected by the second sensing device at multiple acquisition times with different poses, and the first sensing device and the second sensing device are relatively fixedly arranged.
[0080] A processing instruction refers to an instruction for three-dimensional reconstruction of a target object parked in a target area.
[0081] Among them, the processing instruction can be issued by a user through an intelligent terminal communicatively connected to a processing device in a three-dimensional reconstruction system. As an example, the human-computer interaction interface of the intelligent terminal can specifically display a platform interface pre-specified by a service provider for vehicle washing, and the user can issue a processing instruction by clicking on a specific component in the platform interface. The user's intelligent terminal can also be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, in-vehicle intelligent devices, etc., and the portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The present embodiment does not limit the setting form of the intelligent terminal for the user to issue a processing instruction, and any intelligent terminal capable of wired or wireless communication connection with the processing device in the three-dimensional reconstruction system should be included in the protection scope of this application.
[0082] In addition, the processing instruction can also be issued by the user through the human-computer interaction interface of the processing device in the three-dimensional reconstruction system. Among them, the processing device can also be a human-computer interaction device fixedly connected by wire to a first sensing device and a second sensing device near the target area, such as a touch screen set on a wall or a touch screen set on a separate column.
[0083] As an example, the first sensing device can be a laser sensor, and the second sensing device can be a camera.
[0084] The first sensing device and the second sensing device are relatively fixedly arranged to indicate that the relative position formed by the spatial arrangement of the first sensing device and the second sensing device is fixed. Each pose of the first sensing device has a corresponding unique pose of the second sensing device. The relative fixed arrangement of the first sensing device and the second sensing device helps to ensure data consistency and facilitate subsequent data processing and fusion.
[0085] Step 304: Based on the acquisition times corresponding to each three-dimensional spatial point in the three-dimensional point cloud, establish a co-visibility relationship between the acquisition times of the first sensing device.
[0086] Specifically, the three-dimensional point cloud data corresponding to the target object is generated by the first sensing device at different acquisition times. These three-dimensional point cloud data contain the geometric information of the surface of the target object and the surrounding environment. Each three-dimensional spatial point contains position information and reflection intensity information, and each three-dimensional spatial point has at least one corresponding acquisition time, which is used to indicate the viewing angles corresponding to different acquisition times at which the first sensing device can observe the same position.
[0087] The co-visibility relationship is used to indicate the coincidence relationship between the 3D point cloud data collected at a certain acquisition moment and the 3D point cloud data collected at other acquisition moments. Specifically, when the first sensing device observes a target object in a target area at different acquisition moments, it can identify which 3D spatial points in the 3D point cloud data collected at different acquisition moments represent the same position.
[0088] By establishing the co-visibility relationship between each acquisition moment, not only can the spatial information of the 3D point cloud data be enhanced, but also a basis for 3D reconstruction in a dynamic scene can be provided.
[0089] Step 306: For the 2D image corresponding to any acquisition moment, in combination with the co-visibility relationship between each acquisition moment, filter out the 3D spatial points that the second sensing device can observe with the pose corresponding to the current acquisition moment, so as to establish the mapping relationship between the 2D image and the 3D spatial points.
[0090] The processing device can, for any 2D image, according to the acquisition moment of the 2D image and the co-visibility relationship between each acquisition moment, regard the 3D spatial points corresponding to the acquisition moment of the 2D image and the 3D spatial points corresponding to the remaining acquisition moments that have a co-visibility relationship with the acquisition moment of the 2D image as the 3D spatial points corresponding to the 2D image, and traverse all 2D images to establish the mapping relationship between each 2D image and several 3D spatial points in the 3D point cloud.
[0091] Step 308: Based on the mapping relationship, establish the association relationship between each 2D image, and perform at least partial 3D reconstruction of the target object based on the association relationship.
[0092] As an example, based on the 3D spatial points corresponding to each 2D image, for any 2D image, the processing device determines the number of coincidences between the 3D spatial points corresponding to each of the remaining 2D images and the 3D spatial points corresponding to the current 2D image, and sorts the remaining 2D images in descending order according to the number of coincident 3D spatial points. The sorting result is used to indicate the level of association between the remaining 2D images and the current 2D image. The processing device can, for example, determine that several of the 2D images ranked in the front are associated with the current 2D image, so as to determine the association relationship between each 2D image.
[0093] The above three-dimensional reconstruction method can collect high-precision and high-density three-dimensional point cloud data using the first sensing device, and further establish the co-visibility relationship of the viewpoints corresponding to the first sensing device at each acquisition time. Thus, the three-dimensional space points observed from multiple co-visible viewpoints of the first sensing device can be associated with the same two-dimensional image, thereby increasing the number of three-dimensional space points corresponding to a single two-dimensional image. When facing occluded areas (in the vehicle automatic washing scenario, the occluded area is, for example, the vehicle bottom) or low reflectivity areas (in the vehicle automatic washing scenario, the low reflectivity area is, for example, a black vehicle body), more complete and more three-dimensional space points corresponding to the current viewpoint can be screened out; moreover, by using the mapping relationship between the three-dimensional space points and the two-dimensional images to establish the association relationship between different two-dimensional images, false matches can be filtered out and the accuracy of cross-viewpoint matching can be improved; through the association between different two-dimensional images, shared three-dimensional space points can be found in multiple two-dimensional images, thereby improving the redundancy and reliability of the three-dimensional point cloud data, and thus enhancing the robustness of the three-dimensional reconstruction process.
[0094] In some alternative embodiments, step 304 includes:
[0095] Based on the acquisition times corresponding to the three-dimensional space points in the three-dimensional point cloud, determine the first value of the same three-dimensional space points acquired at each acquisition time and any other acquisition time;
[0096] When the first value reaches a preset threshold, determine that the two acquisition times are co-visible, thereby establishing the co-visibility relationship between each acquisition time.
[0097] As Figure 4 shown, in one embodiment, the multiple three-dimensional space points included in the three-dimensional point cloud can be stored using a voxel structure. Storing the multiple three-dimensional space points included in the three-dimensional point cloud in a voxel structure means dividing the three-dimensional space into uniform or non-uniform small cubic voxel units. Each voxel unit contains several three-dimensional space points within the spatial region, and the size of each voxel unit can be adjusted according to the actual scenario. When the target scenario is large, larger voxel units can be used as the minimum storage unit, and when the target scenario is small, the volume of the voxel unit can be correspondingly reduced. This storage method can use separate voxel units to store at least part of the three-dimensional point cloud in a region of the three-dimensional space.
[0098] Furthermore, the processing device can generate a corresponding hash value for each voxel unit and use an octree structure to store the voxel units, thereby realizing hierarchical management of the three-dimensional space. It should be noted that this application does not limit the specific steps for generating the hash value corresponding to each voxel unit, as long as it can generate a unique hash value for each voxel unit to achieve accurate differentiation of each voxel unit.
[0099] The process of constructing an octree may include: First, initialize the octree nodes, including initializing the bounding box, grading index, and voxel data of each octree node. The bounding box refers to the bounding box of the three-dimensional region represented by the current octree node. The grading index is used to record the depth or level of the current octree node. The voxel data points to the voxel unit or the sub-octree nodes of the next layer. Second, recursively construct the octree. In this step, first, calculate the bounding box of the entire voxel structure and create the root node of the octree. The bounding box of the root node needs to cover all voxel units. Subsequently, for each voxel unit, check whether the position of the voxel unit is within the bounding box of the current octree node. If not, skip this voxel unit. If it is within the bounding box of the current octree node, further determine whether the current octree node is a leaf node (i.e., has no child nodes). If it is a leaf node and the number of voxel units within the current octree node does not exceed the number of voxel units that a single octree node can store, add the voxel unit to the current octree node. If the current octree node is already a leaf node and the number of voxel units within the current octree node exceeds the number of voxel units that a single octree node can store, then the octree node needs to be split: create eight sub-octree nodes of the current octree node, each sub-octree node representing one-eighth of the space of the current octree node, redistribute the voxel units within the current octree node to each sub-octree node, and delete the voxel units of the current octree node. Continue to recursively call the process of inserting voxel units for the newly created sub-octree nodes, so as to add voxel units in the correct hierarchical structure. When all voxel units are inserted into the octree structure, the construction of the octree is completed.
[0100] As a hierarchical data structure, the octree structure can effectively manage three-dimensional space, and the hash value corresponding to each voxel unit can be used for fast indexing and positioning of specific voxel units. In the hash table, by calculating the hash value, the position of the voxel unit-related data in memory can be quickly determined, avoiding layer-by-layer search in the octree structure.
[0101] In this embodiment, after the processing device collects the three-dimensional point cloud in step 302 and stores the three-dimensional point cloud using the voxel structure and constructs the octree structure, it can match the three-dimensional space points collected at each acquisition moment to the voxel units where these three-dimensional space points are located using the octree structure, and label all the matched voxel units with the labels corresponding to the acquisition moments, so as to indicate that all the three-dimensional space points within the current voxel unit can be collected by the first sensing device at this acquisition moment.
[0102] When the processing device executes the step of determining the first value of the same three-dimensional spatial points collected at each acquisition moment and any other acquisition moment based on the acquisition moments corresponding to the three-dimensional spatial points in the three-dimensional point cloud, it can match the voxel units corresponding to each acquisition moment according to the tags carried by each voxel unit, and regard all the three-dimensional spatial points included in these voxel units as the three-dimensional spatial points corresponding to the current acquisition moment, so as to realize the expansion of the three-dimensional spatial points corresponding to each acquisition moment. In actual operation, due to the sparsity of the three-dimensional point cloud data or the hardware limitations of the first sensing device, there may not be enough data points in some areas of the target object. By expanding the three-dimensional spatial points actually collected by the first sensing device at any acquisition moment, the blanks in these areas can be effectively filled, and the integrity of the data can be improved.
[0103] Further, based on the three-dimensional spatial points corresponding to each acquisition moment, determine the number of repeated three-dimensional spatial points among the three-dimensional spatial points corresponding to any two acquisition moments to obtain the first value. When the first value reaches a preset threshold, determine that the two acquisition moments are co-visible, so as to establish the co-visibility relationship between each acquisition moment.
[0104] The above three-dimensional reconstruction method can effectively establish the association between different acquisition moments based on the repeatability of the three-dimensional spatial points corresponding to the acquisition moments, introduce the time dimension so that the three-dimensional point cloud is not just a set of static spatial points, but a dynamic representation that includes the changes of the first sensing device in the time dimension. By establishing the co-visibility relationship between different acquisition moments, the dynamic change features can be extracted using the time series, which helps to predict and reflect the changes in the perspective.
[0105] In some optional embodiments, step 306 includes:
[0106] For the two-dimensional picture corresponding to any acquisition moment, in combination with the co-visibility relationship between each acquisition moment, determine at least one acquisition moment that is co-visible with the current acquisition moment;
[0107] Screen the three-dimensional spatial points corresponding to the current acquisition moment and the three-dimensional spatial points corresponding to at least one acquisition moment that is co-visible with the current acquisition moment, and obtain the three-dimensional spatial points that the second sensing device can observe in the pose corresponding to the current acquisition moment, so as to establish the mapping relationship between the two-dimensional picture and the three-dimensional spatial points.
[0108] Optionally, the step of screening the three-dimensional spatial points corresponding to the current acquisition moment and the three-dimensional spatial points corresponding to at least one acquisition moment that is co-visible with the current acquisition moment, and obtaining the three-dimensional spatial points that the second sensing device can observe in the pose corresponding to the current acquisition moment includes:
[0109] Perform a three-dimensional transformation on the three-dimensional space points corresponding to the current acquisition moment and the three-dimensional space points corresponding to at least one acquisition moment co-visible with the current acquisition moment to obtain the image three-dimensional points corresponding to each three-dimensional space point;
[0110] Based on the depth information of the image three-dimensional points, filter the image three-dimensional points;
[0111] Perform a two-dimensional transformation on the filtered image three-dimensional points to obtain the image position points corresponding to each of the filtered image three-dimensional points;
[0112] Obtain the size information of the two-dimensional picture corresponding to the current acquisition moment, and filter the image position points based on the size information;
[0113] Take the three-dimensional space points corresponding to the filtered image position points as the three-dimensional space points that the second sensing device can observe in the pose corresponding to the current acquisition moment.
[0114] As Figure 5 shown, specifically, in this embodiment, three coordinate systems are set, including the lidar coordinate system corresponding to the first sensing device 、the camera coordinate system corresponding to the second sensing device and the world coordinate system . Since the first sensing device and the second sensing device are relatively fixedly arranged, the transformation matrix for converting the camera coordinate system to the lidar coordinate system can be fixedly defined as . This transformation matrix is used to indicate the external parameter transformation between the first sensing device and the second sensing device. The world coordinate system 、usually the lidar coordinate system corresponding to the first sensing device in the original position, during the process of the first sensing device collecting the three-dimensional point cloud, the pose change of the subsequent first sensing device relative to the first sensing device in the original position can be marked as .
[0115] The step of performing a three-dimensional transformation on the three-dimensional space points corresponding to the current acquisition moment and the three-dimensional space points corresponding to at least one acquisition moment co-visible with the current acquisition moment, which is used to convert the three-dimensional space points in the lidar coordinate system to the image three-dimensional points in the camera coordinate system , can be determined by using the transformation matrix for converting the camera coordinate system to the lidar coordinate system , and the pose change of the first sensing device corresponding to the current acquisition moment relative to the first sensing device in the original position.
[0116] As an example, for the three-dimensional space points corresponding to the current acquisition moment and the three-dimensional space points corresponding to at least one acquisition moment that is co-visible with the current acquisition moment, the step of performing three-dimensional conversion can adopt the following formula to obtain the image three-dimensional points corresponding to each three-dimensional space point:
[0117]
[0118] Among them, represents the coordinate position of the image three-dimensional point, ; represents the rotation matrix in the transformation matrix of the extrinsic parameters of the first sensing device and the second sensing device ; represents the translation vector in the transformation matrix of the extrinsic parameters of the first sensing device and the second sensing device ; represents the rotation matrix of the inverse matrix of the pose of the first sensing device in the world coordinate system ; represents the translation vector of the inverse matrix of the pose of the first sensing device in the world coordinate system ; represents the translation vector of the center point of the voxel unit where each three-dimensional space point is located in the world coordinate system ; represents the translation vector of the center point of the voxel unit where each three-dimensional space point is located in the world coordinate system ;
[0119] Furthermore, in the step of screening image three-dimensional points based on the depth information of the image three-dimensional points, the Z-axis in the camera coordinate system can be understood as emitting outward along the optical axis from the center of the second sensing device. When the depth information of the image three-dimensional point is less than 0, it can be considered that the three-dimensional space point corresponding to this image three-dimensional point is on the back of the second sensing device, that is, it is not within the field of view of the second sensing device. Therefore, based on the depth information of each image three-dimensional point, the three-dimensional space points corresponding to the image three-dimensional points with a depth less than 0 are removed.
[0120] In addition, points that are too far away from the center of the second sensing device, that is, points with a large depth information , often have larger errors and will cause certain interference to subsequent projections. Therefore, in this step, a depth upper limit value can be preset, and the three-dimensional space points corresponding to the image three-dimensional points with a depth greater than the depth upper limit value are removed.
[0121] Optionally, screening image three-dimensional points based on the depth information of the image three-dimensional points includes:
[0122] Based on a preset depth range, retain the three-dimensional image points that meet the depth range and remove the three-dimensional image points that do not meet the depth range to screen the three-dimensional image points.
[0123] Among them, the depth range can be [0, dmax], where dmax represents the upper limit value of the depth.
[0124] Furthermore, if it is necessary to convert the three-dimensional image points after screening in the camera coordinate system to the corresponding image position points in the uv plane where the two-dimensional picture is located, it is necessary to first convert the three-dimensional image points in the camera coordinate system to the two-dimensional image points in the image coordinate system, and then convert the two-dimensional image points in the image coordinate system to the image position points in the uv plane where the two-dimensional picture is located. As Figure 6 shown, convert the three-dimensional image point P after screening in the camera coordinate system to the corresponding image position point p in the uv plane where the two-dimensional picture is located. This involves the conversion of the origin and measurement units as well as the conversion of ratios and sums. This application does not elaborate on the conversion process between the three-dimensional image points in the camera coordinate system and the corresponding image position points in the uv plane where the two-dimensional picture is located. However, it should be noted that any method that can satisfy the conversion of three-dimensional points in one coordinate system to the position points in the plane where a picture is located should be included in the protection scope of this application.
[0125] In the uv plane where the two-dimensional picture is located, the size information corresponding to the two-dimensional picture at the current acquisition moment can, for example, indicate the size range of the two-dimensional picture. For example, the length of the two-dimensional picture is (0, h), and the width of the two-dimensional picture is (0, w).
[0126] Optionally, screen the image position points based on the size information, including:
[0127] Retain the image position points that meet the size range and remove the image position points that do not meet the size range to screen the image position points.
[0128] The above three-dimensional reconstruction method can ensure the removal of three-dimensional space points that do not conform to the imaging physical principle, have high noise, and poor accuracy by screening out three-dimensional image points with depth information less than 0 or greater than the upper limit value of the depth. By screening out the image position points that do not meet the size range, it is possible to retain only the points that are actually visible in the two-dimensional picture. Through the combination of depth screening and plane screening, it can ensure that the finally obtained three-dimensional space points are actually valid, can be accurately mapped to the two-dimensional picture, and at the same time improve the reliability and calculation efficiency of data analysis. Overall, it optimizes the correlation between the three-dimensional point cloud data and the two-dimensional picture, which is helpful for subsequent environmental perception, object recognition, and decision support.
[0129] In an alternative embodiment, after obtaining the size information of the two-dimensional image corresponding to the current acquisition moment, the following steps are further included:
[0130] Divide the two-dimensional image into multiple image regions according to the size range corresponding to the two-dimensional image;
[0131] The step of retaining the image position points that meet the size range and removing the image position points that do not meet the size range to screen the image position points includes:
[0132] Retain the image position points that meet the size range and remove the image position points that do not meet the size range;
[0133] For the retained image position points, determine the second value of the image position points falling in each image region;
[0134] For each image region, when the second value reaches a preset threshold, randomly remove the image position points falling within the current image region so that the total number of image position points within the current image region is less than or equal to the preset threshold to screen the image position points.
[0135] As Figure 7 shown, the processing device can divide the two-dimensional image into multiple image regions according to the size information corresponding to the two-dimensional image.
[0136] In a specific image region, if the number of image position points is too large, it may lead to data redundancy. The 3D reconstruction method in this embodiment controls the number of image position points in each image region by randomly removing redundant image position points, which can reduce data redundancy while retaining the representative and information-valued image position points in the image region and more accurately reflecting the characteristics of the scene.
[0137] In some alternative embodiments, step 308 includes:
[0138] Traverse each two-dimensional image, and based on the mapping relationship between the two-dimensional image and the three-dimensional space points, determine the third value of the same three-dimensional space points corresponding to each two-dimensional image and any other two-dimensional image;
[0139] For each two-dimensional image, perform a correlation degree sorting on the other two-dimensional images based on the third value corresponding to any other two-dimensional image;
[0140] Based on the result of the correlation degree sorting, establish the correlation relationship between each two-dimensional image.
[0141] The third value is used to indicate the number of three-dimensional space points shared between two two-dimensional images.
[0142] As an example, there are four two-dimensional images. The set of three-dimensional space points corresponding to the two-dimensional image 1 is {P1, P2, P3, P4}, the set of three-dimensional space points corresponding to the two-dimensional image 2 is {P2, P3, P4, P5}, the set of three-dimensional space points corresponding to the two-dimensional image 3 is {P3, P5, P6, P7}, and the set of three-dimensional space points corresponding to the two-dimensional image 4 is {P1, P2, P6, P7}. Further calculate the number of shared three-dimensional points between every two two-dimensional images (i.e., the third value). It can be determined that the two-dimensional image 1 and the two-dimensional image 2 share three-dimensional space points P2, P3, P4, that is, the third value is 3; the two-dimensional image 1 and the two-dimensional image 3 share the three-dimensional space point P3, that is, the third value is 1; the two-dimensional image 1 and the two-dimensional image 4 share the three-dimensional space points P1, P2, that is, the third value is 2; the two-dimensional image 2 and the two-dimensional image 3 share the three-dimensional space points P3, P5, that is, the third value is 2; the two-dimensional image 2 and the two-dimensional image 4 share the three-dimensional space point P2, that is, the third value is 1; the two-dimensional image 3 and the two-dimensional image 4 share the three-dimensional space points P6, P7, that is, the third value is 2.
[0143] For the two-dimensional image 1, based on the third values 3, 1, 2 corresponding to the two-dimensional images 2, 3, 4, the correlation degree sorting result of the two-dimensional images 2, 3, 4 is the two-dimensional image 2, the two-dimensional image 4, and the two-dimensional image 3.
[0144] For the two-dimensional image 2, based on the third values 3, 2, 1 corresponding to the two-dimensional images 1, 3, 4, the correlation degree sorting result of the two-dimensional images 2, 3, 4 is the two-dimensional image 1, the two-dimensional image 3, and the two-dimensional image 4.
[0145] For the two-dimensional image 3, based on the third values 2, 2, 2 corresponding to the two-dimensional images 1, 2, 4, the correlation degree sorting result of the two-dimensional images 2, 3, 4 is a tie among the two-dimensional image 1, the two-dimensional image 2, and the two-dimensional image 4.
[0146] For the two-dimensional image 4, based on the third values 2, 1, 2 corresponding to the two-dimensional images 1, 2, 3, the correlation degree sorting result of the two-dimensional images 1, 2, 3 is a tie between the two-dimensional image 1 and the two-dimensional image 3, and the two-dimensional image 2.
[0147] In this embodiment, the processing device may, based on the correlation degree sorting result, regard the top n two-dimensional images in the sorting as having a correlation relationship with the current two-dimensional image. Or, it may directly, based on the third value, regard the two-dimensional images with the third value greater than the preset value threshold as having a correlation relationship with the current two-dimensional image.
[0148] The more the number of shared three-dimensional space points, the higher the overlap degree of the target scenes observed by the two two-dimensional images, and the stronger the matching reliability, thereby avoiding incorrect matching caused by the dynamic environment.
[0149] The above three-dimensional reconstruction method establishes the correlation relationship between different two-dimensional images through the co-visibility of three-dimensional space points, which is more stable than traditional pure visual matching (such as feature point matching). It can comprehensively utilize the observation data at different acquisition times and from different perspectives to form a multi-view understanding of a specific scene, so as to obtain richer information in a complex environment.
[0150] It should be understood that although the steps in the flowcharts involved in the above embodiments are sequentially shown according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.
[0151] Based on the same inventive concept, the embodiments of the present application also provide a three-dimensional reconstruction device for implementing the above-mentioned three-dimensional reconstruction method. The solution provided by the three-dimensional reconstruction device to solve the problem is similar to the solution described in the above three-dimensional reconstruction method. Therefore, the specific limitations in one or more of the following device embodiments can refer to the limitations on the three-dimensional reconstruction method in the above text, and will not be repeated here.
[0152] In one embodiment, as Figure 8 shown, a three-dimensional reconstruction device 800 is provided, including:
[0153] An acquisition module 802, configured to acquire multiple two-dimensional images and three-dimensional point clouds of a target object in response to a processing instruction; the three-dimensional point cloud is acquired by a first sensing device at different poses at multiple acquisition times, and the two-dimensional images are acquired by a second sensing device at different poses at multiple acquisition times, and the first sensing device and the second sensing device are relatively fixedly arranged;
[0154] A establishing module 804, configured to establish a co-visibility relationship between the acquisition times of the first sensing device based on the acquisition times corresponding to the three-dimensional space points in the three-dimensional point cloud;
[0155] A screening module 806, configured to, for the two-dimensional image corresponding to any acquisition time, combine the co-visibility relationship between the acquisition times to screen out the three-dimensional space points that the second sensing device can observe in the pose corresponding to the current acquisition time, so as to establish a mapping relationship between the two-dimensional image and the three-dimensional space points;
[0156] A reconstruction module 808, configured to establish an association relationship between each two-dimensional image based on a mapping relationship, and perform at least partial three-dimensional reconstruction of the target object based on the association relationship.
[0157] In some alternative embodiments, the establishing module 804 is further configured to:
[0158] Based on the acquisition time corresponding to each three-dimensional spatial point in the three-dimensional point cloud, determine a first value of the same three-dimensional spatial point acquired at each acquisition time and any other acquisition time;
[0159] When the first value reaches a preset threshold, determine that the two acquisition times are co-visible, thereby establishing a co-visibility relationship between each acquisition time.
[0160] In some alternative embodiments, the screening module 806 is further configured to:
[0161] For the two-dimensional image corresponding to any acquisition time, in combination with the co-visibility relationship between each acquisition time, determine at least one acquisition time that is co-visible with the current acquisition time;
[0162] Screen the three-dimensional spatial points corresponding to the current acquisition time and the three-dimensional spatial points corresponding to at least one acquisition time that is co-visible with the current acquisition time, to obtain the three-dimensional spatial points that the second sensing device can observe in the pose corresponding to the current acquisition time, so as to establish a mapping relationship between the two-dimensional image and the three-dimensional spatial points.
[0163] In some alternative embodiments, the screening module 806 is further configured to:
[0164] Perform three-dimensional conversion on the three-dimensional spatial points corresponding to the current acquisition time and the three-dimensional spatial points corresponding to at least one acquisition time that is co-visible with the current acquisition time, to obtain image three-dimensional points corresponding to each three-dimensional spatial point;
[0165] Based on the depth information of the image three-dimensional points, screen the image three-dimensional points;
[0166] Perform two-dimensional conversion on the screened image three-dimensional points, to obtain image position points corresponding to each screened image three-dimensional point;
[0167] Obtain the size information of the two-dimensional image corresponding to the current acquisition time, and screen the image position points based on the size information;
[0168] Use the three-dimensional spatial points corresponding to the screened image position points as the three-dimensional spatial points that the second sensing device can observe in the pose corresponding to the current acquisition time.
[0169] In some alternative embodiments, the screening module 806 is further configured to:
[0170] Based on a preset depth range, retain the three-dimensional image points that meet the depth range and remove the three-dimensional image points that do not meet the depth range to filter the three-dimensional image points.
[0171] In some alternative embodiments, the size information includes the size range corresponding to the two-dimensional picture;
[0172] The screening module 806 is further configured to:
[0173] Retain the image position points that meet the size range and remove the image position points that do not meet the size range to filter the image position points.
[0174] In some alternative embodiments, the screening module 806 is further configured to:
[0175] Divide the two-dimensional picture into multiple image regions according to the size range corresponding to the two-dimensional picture;
[0176] Retain the image position points that meet the size range and remove the image position points that do not meet the size range;
[0177] For the retained image position points, determine the second value of the image position points falling in each image region;
[0178] For each image region, when the second value reaches a preset threshold, randomly remove the image position points falling within the current image region so that the total number of image position points within the current image region is less than or equal to the preset threshold to filter the image position points.
[0179] In some alternative embodiments, the reconstruction module 808 is further configured to:
[0180] Traverse each two-dimensional picture, and based on the mapping relationship between the two-dimensional picture and the three-dimensional space points, determine the third value of the same three-dimensional space points corresponding to each two-dimensional picture and any other two-dimensional picture;
[0181] For each two-dimensional picture, based on the third value corresponding to any other two-dimensional picture, sort the correlation degrees of the other two-dimensional pictures;
[0182] Based on the result of the correlation degree sorting, establish the correlation relationship between each two-dimensional picture.
[0183] Each module in the above device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form so that the processor can call and execute the operations corresponding to the above respective modules.
[0184] In one embodiment, as Figure 9As shown, a vehicle cleaning system 900 is provided, including the above three-dimensional reconstruction system 100, a cleaning head 910, and a controller 920;
[0185] The controller is respectively connected to a processing device 130 in the three-dimensional reconstruction system 100 and the cleaning head;
[0186] The cleaning head 910 is connected to at least one containing tank;
[0187] The controller 920 is configured to control the movement of the cleaning head 910 and control the cleaning head 910 to spray the cleaning liquid in the containing tank onto the vehicle body based on the result of at least partial three-dimensional reconstruction of the vehicle in the target area by the processing device 130.
[0188] The cleaning head 910 is connected to at least one containing tank through a pipeline. The cleaning head 910 can be arranged on another movable base or a sliding track, and the cleaning head 910 can move along another preset path or along the sliding track to clean the vehicle body of the vehicle parked in the target area.
[0189] Figure 10 It is a schematic structural diagram of the electronic device provided in this application. As Figure 10 shown, the electronic device 1000 provided in this embodiment includes: at least one processor 1001 and a memory 1002. Optionally, the device 1000 further includes a communication component 1003. Among them, the processor 1001, the memory 1002, and the communication component 1003 are connected through a bus 1004.
[0190] In a specific implementation process, at least one processor 1001 executes computer-executable instructions stored in the memory 1002, so that at least one processor 1001 executes the above method.
[0191] For the specific implementation process of the processor 1001, reference can be made to the above method embodiment, and its implementation principle and technical effects are similar, which will not be elaborated here in this embodiment.
[0192] In the above embodiment, it should be understood that the processor can be a central processing unit (English: Central Processing Unit, abbreviated as: CPU), and can also be other general-purpose processors, digital signal processors (English: Digital Signal Processor, abbreviated as: DSP), application specific integrated circuits (English: Application Specific Integrated Circuit, abbreviated as: ASIC), etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in combination with the invention can be directly embodied as being executed by a hardware processor, or executed by a combination of hardware and software modules in the processor.
[0193] The memory may include a random access memory (RAM), and may also include a non-volatile memory (NVM), such as at least one disk memory.
[0194] The bus may be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, the buses in the drawings of this application are not limited to only one bus or one type of bus.
[0195] This application also provides a computer program product, including a computer program which, when executed by a processor, implements the above method.
[0196] This application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the above method.
[0197] The above-readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disc. The readable storage medium can be any available medium accessible by a general-purpose or special-purpose computer.
[0198] An exemplary readable storage medium is coupled to the processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can be located in an application specific integrated circuit (ASIC). Of course, the processor and the readable storage medium can also exist as discrete components in a device.
[0199] The division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings or direct couplings or communication connections shown or discussed among each other can be indirect couplings or communication connections through some interfaces, devices or units, and can be in electrical, mechanical or other forms.
[0200] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0201] In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.
[0202] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or a part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks or optical discs that can store program codes.
[0203] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When this program is executed, it executes the steps including the above method embodiments; and the foregoing storage medium includes: various media such as ROM, RAM, magnetic disks or optical discs that can store program codes.
[0204] Finally, it should be noted that those skilled in the art will readily conceive of other embodiments of the present invention after considering the specification and practicing the invention disclosed herein. The present invention is intended to cover any variations, uses, or adaptations of the present invention, which follow the general principles of the present invention and include well-known common general knowledge or conventional technical means in the technical field not disclosed in the present invention. It is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present invention is only limited by the appended claims.
Claims
1. A three-dimensional reconstruction method, characterized in that: include: In response to the processing instruction, collecting a plurality of two-dimensional images and a three-dimensional point cloud of the target object; The three-dimensional point cloud is acquired by a first sensing device at different positions and postures at multiple acquisition times, the two-dimensional image is acquired by a second sensing device at different positions and postures at multiple acquisition times, and the first sensing device and the second sensing device are relatively fixedly arranged; Based on the acquisition time corresponding to each three-dimensional spatial point in the three-dimensional point cloud, establishing a common view relationship between each acquisition time for the first sensing device; For a two-dimensional image corresponding to any acquisition moment, combined with the common view relationship between the acquisition moments, the three-dimensional space points that can be observed by the second sensing device at the posture corresponding to the current acquisition moment are screened to establish a mapping relationship between the two-dimensional image and the three-dimensional space points; Based on the mapping relationship, an association relationship is established between the two-dimensional images, and based on the association relationship, at least a portion of the target object is reconstructed in three dimensions.
2. The method according to claim 1, characterized in that The establishing of a common view relationship between the acquisition moments of the first sensing device based on the acquisition moments corresponding to the three-dimensional spatial points in the three-dimensional point cloud includes: Based on the collection time corresponding to each three-dimensional space point in the three-dimensional point cloud, determine a first value of the same three-dimensional space point collected at each collection time and at any other collection time; When the first value reaches a preset threshold, it is determined that the two acquisition moments are in common view, thereby establishing a common view relationship between the acquisition moments.
3. The method according to claim 1, characterized in that The method of screening the two-dimensional image corresponding to any acquisition moment and combining the common view relationship between the acquisition moments to obtain the three-dimensional space point that can be observed by the second sensing device at the posture corresponding to the current acquisition moment to establish a mapping relationship between the two-dimensional image and the three-dimensional space point includes: For a two-dimensional image corresponding to any acquisition moment, combining the co-viewing relationship between the acquisition moments, determining at least one acquisition moment that is co-viewing with the current acquisition moment; The three-dimensional space points corresponding to the current acquisition moment and the three-dimensional space points corresponding to at least one acquisition moment that is co-visible with the current acquisition moment are screened to obtain the three-dimensional space points that can be observed by the second sensing device with the posture corresponding to the current acquisition moment, so as to establish a mapping relationship between the two-dimensional image and the three-dimensional space points.
4. The method according to claim 3, characterized in that The three-dimensional space point corresponding to the current acquisition moment and the three-dimensional space point corresponding to at least one acquisition moment that is co-visual with the current acquisition moment are screened to obtain the three-dimensional space point that can be observed by the second sensing device at the posture corresponding to the current acquisition moment, including: Performing three-dimensional conversion on the three-dimensional space point corresponding to the current acquisition moment and the three-dimensional space point corresponding to at least one acquisition moment that is co-visual with the current acquisition moment, to obtain the three-dimensional image point corresponding to each of the three-dimensional space points; Based on the depth information of the three-dimensional points of the image, screening the three-dimensional points of the image; Performing two-dimensional transformation on the filtered three-dimensional image points to obtain image position points corresponding to each filtered three-dimensional image point; Obtaining size information of the two-dimensional image corresponding to the current acquisition moment, and filtering the image position points based on the size information; The three-dimensional space point corresponding to the filtered image position point is used as the three-dimensional space point that can be observed by the second sensing device at the posture corresponding to the current acquisition time.
5. The method according to claim 4, characterized in that The screening of the three-dimensional points of the image based on the depth information of the three-dimensional points of the image includes: Based on a preset depth range, the image 3D points that meet the depth range are retained, and the image 3D points that do not meet the depth range are removed, so as to screen the image 3D points.
6. The method according to claim 4, characterized in that The size information includes a size range corresponding to the two-dimensional image; The screening of the image position points based on the size information includes: The image position points that meet the size range are retained, and the image position points that do not meet the size range are removed, so as to screen the image position points.
7. The method according to claim 6, characterized in that After obtaining the size information of the two-dimensional image corresponding to the current acquisition time, the method further includes: Dividing the two-dimensional image into a plurality of image regions according to a size range corresponding to the two-dimensional image; The step of retaining the image position points that meet the size range and removing the image position points that do not meet the size range to screen the image position points includes: The image position points that meet the size range are retained, and the image position points that do not meet the size range are removed; For the reserved image position points, determining second values of the image position points falling in each of the image regions; For each of the image areas, when the second value reaches a preset threshold, the image position points falling within the current image area are randomly removed so that the total number of image position points in the current image area is less than or equal to the preset threshold, so as to screen the image position points.
8. The method according to claim 1, characterized in that: The establishing of an association relationship between the two-dimensional images based on the mapping relationship includes: Traversing each of the two-dimensional images, and determining, based on a mapping relationship between the two-dimensional images and the three-dimensional space points, a third value of a same three-dimensional space point corresponding to each of the two-dimensional images and any other two-dimensional images; For each of the two-dimensional pictures, based on a third value corresponding to any of the other two-dimensional pictures, sorting the other two-dimensional pictures by relevance; Based on the result of the association degree sorting, an association relationship between the two-dimensional pictures is established.
9. A three-dimensional reconstruction device, characterized in that: include: A collection module, for collecting a plurality of two-dimensional images and a three-dimensional point cloud of a target object in response to a processing instruction; The three-dimensional point cloud is acquired by a first sensing device at different positions and postures at multiple acquisition times, the two-dimensional image is acquired by a second sensing device at different positions and postures at multiple acquisition times, and the first sensing device and the second sensing device are relatively fixedly arranged; An establishing module, configured to establish a common view relationship between each acquisition moment of the first sensing device based on the acquisition moment corresponding to each three-dimensional spatial point in the three-dimensional point cloud; A screening module, for screening the two-dimensional image corresponding to any acquisition moment, in combination with the common view relationship between the acquisition moments, to obtain the three-dimensional space points that can be observed by the second sensing device at the posture corresponding to the current acquisition moment, so as to establish a mapping relationship between the two-dimensional image and the three-dimensional space point; A reconstruction module is used to establish an association relationship between the two-dimensional images based on the mapping relationship, and to perform at least a partial three-dimensional reconstruction of the target object based on the association relationship.
10. A three-dimensional reconstruction system, characterized in that: It comprises a first sensing device, a second sensing device and a processing device; the processing device is connected to the first sensing device and the second sensing device respectively; The first sensing device and the second sensing device are relatively fixedly arranged; The processing device is used to control the movement of the first sensing device and the second sensing device, and adopt the three-dimensional reconstruction method described in any one of claims 1 to 8, control the first sensing device to collect the three-dimensional point cloud corresponding to the target object parked in the target area, control the second sensing device to collect the two-dimensional image of the target object, and perform three-dimensional reconstruction of at least part of the target object based on the three-dimensional point cloud and the two-dimensional image.
11. A vehicle cleaning system, characterized in that: comprising the three-dimensional reconstruction system, a cleaning head and a controller as claimed in claim 10; The controller is respectively connected to the processing device and the cleaning head in the three-dimensional reconstruction system; The cleaning head is connected to at least one receiving groove; The controller is used to control the movement of the cleaning head based on the result of the three-dimensional reconstruction of the vehicle in the target area by the processing device, and to control the cleaning head to spray the cleaning liquid in the containing tank onto the body of the vehicle.
12. An electronic device, characterized in that: include: Memory, processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory, so that the processor performs the method according to any one of claims 1 to 8.
13. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 8 when executed by a processor.
14. A computer program product, characterized in that The invention comprises a computer program, which implements the method according to any one of claims 1 to 8 when being executed by a processor.
Citation Information
Patent Citations
Multi-sensor data fusion method, device and system
CN115908578A
Three-dimensional color modeling method and system based on multi-sensor fusion
CN119418006A
Open type three-dimensional reconstruction method, automatic depth positioning method, equipment and robot
CN119832151A
Modeling method and apparatus using three-dimensional (3D) point cloud
US20180211399A1
Three-dimensional reconstruction method, device, system, and storage medium
WO2023164845A1