Multi-modal Image Fusion Method, Device, Computer Equipment and Storage Medium

By determining and removing three-dimensional feature points with excessive fusion errors in the fusion process between three-dimensional images and two-dimensional images, the problem of low image fusion quality in the prior art is solved, and higher image fusion accuracy and quality are achieved.

CN120070210BActive Publication Date: 2025-06-27FOSHAN POWER SUPPLY BUREAU GUANGDONG POWER GRID
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510550289.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-06-27
Estimated Expiration
2045-04-29

AI Technical Summary

Technical Problem

The problem of low image fusion quality in the prior art, especially when there are many three-dimensional feature points with excessive fusion errors in the process of fusion between three-dimensional images and two-dimensional images, affecting the accuracy and quality of image fusion.

Method used

By acquiring three-dimensional images and two-dimensional images and dividing regions according to the same preset shape, the similarity between the three-dimensional feature points of the same depth distance and the two-dimensional feature points are determined, and the three-dimensional feature points whose similarity is less than the preset similarity threshold are eliminated, so as to perform image fusion.

Benefits of technology

The fusion accuracy of three-dimensional features and two-dimensional feature points is improved, the image fusion quality is improved, and the fusion error is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070210B_ABST
    Figure CN120070210B_ABST
Patent Text Reader

Abstract

A method, apparatus, computer device, and storage medium for fusing multimodal images provided by the present application. The method includes: obtaining a three-dimensional image and a two-dimensional image to be subjected to image fusion, dividing regions according to the same preset shape to obtain a three-dimensional image region corresponding to the three-dimensional image and a two-dimensional image region corresponding to the two-dimensional image; determining three-dimensional feature points located at the same depth distance in the three-dimensional image region; determining two-dimensional feature points corresponding to the three-dimensional feature points in the two-dimensional image region; obtaining the similarity degree between the three-dimensional feature points and the two-dimensional feature points; in the same depth distance of the three-dimensional image region, determining the number of three-dimensional feature points corresponding to a similarity degree less than a preset similarity degree threshold; if the number is greater than a preset number threshold, removing the three-dimensional feature points with a similarity degree less than the preset similarity degree threshold; and obtaining a fused image based on the unremoved three-dimensional feature points and the corresponding two-dimensional feature points. Using this method can improve the quality of image fusion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and in particular to a multimodal image fusion method, device, computer equipment and storage medium. Background Art

[0002] With the rapid development of operating robots in various industries, multi-sensor fusion technology provides a perception basis for operating robots. For example, LiDAR and camera fusion is widely used in power operation scenarios due to its wide perception range and high positioning accuracy.

[0003] At present, the fusion method of three-dimensional images and two-dimensional images mainly uses the PnP algorithm to solve the two-dimensional feature points and three-dimensional feature points to obtain the external parameter matrix, so as to realize the direct fusion between multimodal images.

[0004] However, the applicant discovered during implementation that the related technology at least has the problem of low image fusion quality. Summary of the invention

[0005] Based on this, the purpose of this application is to solve at least one of the above-mentioned technical defects, especially the technical defect of low image fusion quality in the prior art. This application provides a multimodal image fusion method, device, computer equipment and storage medium.

[0006] In a first aspect, the present application provides a multimodal image fusion method, the method comprising:

[0007] Acquire a three-dimensional image and a two-dimensional image to be fused, and divide the regions according to the same preset shape to obtain a three-dimensional image region corresponding to the three-dimensional image and a two-dimensional image region corresponding to the two-dimensional image;

[0008] Determine the three-dimensional feature points located at the same depth distance in the three-dimensional image area; and determine the two-dimensional feature points corresponding to the three-dimensional feature points in the two-dimensional image area;

[0009] Obtain the similarity between the three-dimensional feature points and the two-dimensional feature points;

[0010] In the same depth distance of the three-dimensional image area, determine the number of three-dimensional feature points corresponding to the degree of similarity less than a preset similarity threshold; if the number is greater than the preset number threshold, remove the three-dimensional feature points whose similarity is less than the preset similarity threshold;

[0011] Based on the uneliminated three-dimensional feature points and the corresponding two-dimensional feature points, image fusion is performed to obtain a fused image.

[0012] In one embodiment, obtaining the similarity between the three-dimensional feature point and the two-dimensional feature point includes:

[0013] Obtain the projected pixel coordinates corresponding to the three-dimensional feature points and determine the two-dimensional pixel coordinates of the two-dimensional feature points;

[0014] Based on the projected pixel coordinates and the two-dimensional pixel coordinates, obtain the similarity degree.

[0015] In one embodiment, obtaining the projected pixel coordinates corresponding to the three-dimensional feature points includes:

[0016] Based on the three-dimensional feature points at the current depth and the two-dimensional feature points corresponding to the three-dimensional feature points, obtain the external parameter calibration parameters corresponding to the current depth; wherein, the current depth is any depth;

[0017] Use the external parameter calibration parameters to project the three-dimensional feature points at the current depth to obtain the projected pixel coordinates of the three-dimensional feature points.

[0018] In one embodiment, using the external parameter calibration parameters to project the three-dimensional feature points at the current depth to obtain the projected pixel coordinates of the three-dimensional feature points includes:

[0019] Obtain the internal camera parameters corresponding to the camera device of the two-dimensional image;

[0020] Use the external parameter calibration parameters and the internal camera parameters to perform projection processing on the three-dimensional feature points at the current depth to obtain the projected pixel coordinates of the three-dimensional feature points.

[0021] In one embodiment, based on the projected pixel coordinates and the two-dimensional pixel coordinates, obtaining the similarity degree includes:

[0022] Determine the Euclidean distance between the projected pixel coordinates and the two-dimensional pixel coordinates;

[0023] Based on the Euclidean distance, obtain the similarity degree.

[0024] In one embodiment, both the three-dimensional image and the two-dimensional image are divided into a central region and an outer ring region; the outer ring region of the three-dimensional image includes a plurality of three-dimensional image regions; the outer ring region of the two-dimensional image includes a plurality of two-dimensional image regions;

[0025] Based on the uneliminated three-dimensional feature points and the corresponding two-dimensional feature points, perform image fusion to obtain a fused image, including:

[0026] Determine the three-dimensional feature points and the two-dimensional feature points corresponding to the central region, and based on the three-dimensional feature points and the two-dimensional feature points corresponding to the central region, determine the external parameter calibration parameters of the central region;

[0027] Based on the uneliminated three-dimensional feature points in the outer ring region and the two-dimensional feature points corresponding to the uneliminated three-dimensional feature points, determine the external parameter calibration parameters of the outer ring region;

[0028] According to the external parameter calibration parameters of the central region and the external parameter calibration parameters of the outer ring region, image fusion is performed on the central region and the outer ring region respectively to obtain a fused image.

[0029] In one embodiment, the method is applied to a power robot; the three-dimensional image is a point cloud image;

[0030] Obtaining a three-dimensional image and a two-dimensional image to be subjected to image fusion includes:

[0031] Obtaining a point cloud image through a lidar installed on the power robot;

[0032] Obtaining a two-dimensional image through a camera device installed on the power robot.

[0033] In a second aspect, the present application provides a multi-modal image fusion device, which includes:

[0034] An image region determination module, configured to obtain a three-dimensional image and a two-dimensional image to be subjected to image fusion, and perform region division according to the same preset shape to obtain a three-dimensional image region corresponding to the three-dimensional image and a two-dimensional image region corresponding to the two-dimensional image;

[0035] A feature point determination module, configured to determine three-dimensional feature points located at the same depth distance in the three-dimensional image region; and determine two-dimensional feature points corresponding to the three-dimensional feature points in the two-dimensional image region;

[0036] A similarity degree determination module, configured to obtain the similarity degree between the three-dimensional feature points and the two-dimensional feature points;

[0037] A quantity determination module, configured to determine the quantity of three-dimensional feature points corresponding to a similarity degree less than a preset similarity degree threshold in the same depth distance of the three-dimensional image region; if the quantity is greater than a preset quantity threshold, then eliminate the three-dimensional feature points with a similarity degree less than the preset similarity degree threshold;

[0038] An image fusion module, configured to perform image fusion based on the uneliminated three-dimensional feature points and the corresponding two-dimensional feature points to obtain a fused image.

[0039] In a third aspect, the present application provides a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the steps of the above method are implemented.

[0040] In a fourth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above method are implemented.

[0041] It can be seen from the above technical solutions that the embodiments of the present application have the following advantages:

[0042] The multi-modal image fusion method, device, computer device, and storage medium provided by this application divide the three-dimensional image corresponding to the three-dimensional image region and the two-dimensional image corresponding to the two-dimensional image region by dividing the region according to the same preset shape; and by determining the similarity degree between the three-dimensional feature points and the two-dimensional feature points at the same depth distance, it is possible to judge the number of three-dimensional feature points with a similarity degree less than the preset similarity degree threshold at a certain depth distance. In this way, when the number is greater than the preset number threshold, the three-dimensional feature points corresponding to the feature with a similarity degree less than the preset similarity degree threshold can be removed. That is, when it is judged that there are many three-dimensional feature points with too large fusion errors in a certain three-dimensional image region, the three-dimensional feature points with too large fusion errors can be removed, so as to improve the fusion accuracy of the three-dimensional features and the two-dimensional feature points, and further improve the image fusion quality. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0044] Figure 1 Flow schematic of the multi-modal image fusion method provided by the embodiment of the present application Figure 1 ;

[0045] Figure 2 Flow schematic of the multi-modal image fusion method provided by the embodiment of the present application Figure 2 ;

[0046] Figure 3 Flow schematic of the multi-modal image fusion method provided by the embodiment of the present application Figure 3 ;

[0047] Figure 4 Flow schematic of the multi-modal image fusion method provided by the embodiment of the present application Figure 4 ;

[0048] Figure 5 Schematic diagram of an image region division provided by the embodiment of the present application;

[0049] Figure 6 Flow schematic of the multi-modal image fusion method provided by the embodiment of the present application Figure 5 ;

[0050] Figure 7 Schematic diagram of the structure of a multi-modal image fusion device provided by the embodiment of the present application;

[0051] Figure 8 The internal structure diagram of a computer device provided by an embodiment of the present application. Detailed implementation manners

[0052] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0053] With the rapid development of operation robots in various industries (for example, power operation robots), multi-sensor fusion technology provides a perception basis for operation robots. For example, the fusion of lidar and camera is widely used in power operation scenarios due to its wide perception range and high positioning accuracy.

[0054] Currently, for the laser point cloud and image fusion method, the PnP algorithm is mainly used to solve the external parameter matrix for two-dimensional and three-dimensional feature point pairs to achieve the direct fusion of the point cloud and the image.

[0055] During the implementation process, the applicant found that the fusion accuracy of the laser point cloud and image fusion method is highly dependent on the accuracy of feature point extraction and the spatial distribution relationship of the feature points, and the calibration data is mainly manually collected, with low efficiency and easy to generate errors.

[0056] Based on this, the present application provides a fusion method, device, computer device and storage medium for multi-modal images. In the present application, the three-dimensional image and the two-dimensional image are divided into regions according to the same preset shape to obtain a three-dimensional image region corresponding to the three-dimensional image and a two-dimensional image region corresponding to the two-dimensional image; further, by determining the similarity degree between the three-dimensional feature points and the two-dimensional feature points at the same depth distance, the number of three-dimensional feature points with a similarity degree less than a preset similarity degree threshold in this depth distance can be judged. Thus, when the above number in this depth distance is greater than a preset number threshold, the three-dimensional feature points corresponding to the feature with a similarity degree less than the preset similarity degree threshold can be removed. That is, in the present application, when it is judged that there are many three-dimensional feature points with too large fusion errors in a certain three-dimensional image region, the three-dimensional feature points with too large fusion errors can be removed, thereby improving the image fusion quality.

[0057] In an exemplary embodiment, Figure 1 The flow diagram of the fusion method for multi-modal images provided by the embodiment of the present application Figure 1 as Figure 1As shown, a method for fusing multi-modal images is provided. In this embodiment, an example is given where this method is applied to a server. It can be understood that this method can also be applied to a terminal, or to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. This method includes the following S101 to S105. Among them:

[0058] S101. Obtain a three-dimensional image and a two-dimensional image to be fused for image fusion, and perform regional division according to the same preset shape to obtain a three-dimensional image region corresponding to the three-dimensional image and a two-dimensional image region corresponding to the two-dimensional image.

[0059] Among them, the three-dimensional image and the two-dimensional image can be different-modal images obtained by different devices set on the same robot. The three-dimensional image can be a point cloud image. The two-dimensional image can be a camera image. As an example, this application illustrates the fusion method of a point cloud image and a camera image. The point cloud image can be data obtained by a lidar. The camera image can be data obtained by shooting through a camera device.

[0060] The preset shape can refer to a region division shape set in advance. For example, the preset shape can be a "square with a hole in the middle" shape, that is, the central region and the outer ring region of the image can be distinguished, and the outer ring region can be further divided into multiple sub-regions, and each of the multiple sub-regions can be used as a three-dimensional image region or a two-dimensional image region.

[0061] The three-dimensional image region can refer to the sub-region obtained after dividing the three-dimensional image. For example, the sub-region of the outer ring region obtained by dividing the three-dimensional image as described above. The two-dimensional image region can refer to the sub-region obtained after dividing the two-dimensional image. For example, the sub-region of the outer ring region obtained by dividing the two-dimensional image as described above.

[0062] Schematically, a working robot can obtain a three-dimensional image of the surrounding environment through a three-dimensional image device set on itself, and obtain a two-dimensional image of the surrounding environment through a two-dimensional image device; through the fusion of the three-dimensional image and the two-dimensional image, better visual perception of the working robot can be achieved. For example, a point cloud image can be collected by the lidar of a power robot, and a camera image can be collected by a camera.

[0063] Exemplarily, the server can obtain the point cloud image and the camera image collected by the power robot, and can perform regional division on the two images according to a preset shape set in advance to obtain a three-dimensional image region corresponding to the point cloud image and a two-dimensional image region corresponding to the camera image.

[0064] It can be understood that the three-dimensional image and the two-dimensional image need to be divided according to the same shape. After division, there are corresponding two-dimensional image regions in the three-dimensional image regions. For example, the three-dimensional image A and the two-dimensional image B are two images to be fused. After image division using the same shape, the three-dimensional image region A1 corresponds to the two-dimensional image region B1, and the three-dimensional image region A2 corresponds to the two-dimensional image region B2.

[0065] S102. Determine the three-dimensional feature points in the three-dimensional image region that are at the same depth distance; and determine the two-dimensional feature points corresponding to the three-dimensional feature points in the two-dimensional image region.

[0066] Exemplarily, the server can extract the three-dimensional feature points in the three-dimensional image, and for each three-dimensional image region, it can determine the three-dimensional feature points at the same depth in the three-dimensional image region, that is, the three-dimensional feature points in the same plane in the depth direction. The server can also determine the two-dimensional feature points corresponding to the three-dimensional feature points, and the two-dimensional feature points are the feature points in the two-dimensional image region. It can be understood that corresponding to the three-dimensional feature points and the two-dimensional feature points, the image regions where they are located are also corresponding, that is, the three-dimensional image region where the three-dimensional feature points are located corresponds to the two-dimensional image region where the two-dimensional feature points are located.

[0067] Optionally, the robot can use a partition calibration data acquisition device to collect the laser point cloud and the camera image and transmit them to the server. The server can perform feature point extraction to obtain 2D and 3D feature point pairs for external parameter calibration (for example, the corresponding two-dimensional feature points and three-dimensional feature points can be used as feature point pairs). Among them, the camera and the lidar in the partition calibration data acquisition device can be connected and fixed on the turntable through a tooling. The turntable can rotate around the horizontal and vertical directions. The 2D and 3D feature point pairs can be a set of image feature points (two-dimensional feature points) and point cloud feature points (three-dimensional feature points). The set of image feature points can be a set of 2D feature points in the pixel coordinate system, and the set of point cloud feature points can be a set of 3D feature points in the radar coordinate system.

[0068] S103. Obtain the similarity degree between the three-dimensional feature points and the two-dimensional feature points.

[0069] Among them, the similarity degree can refer to the matching degree in terms of feature attributes. For example, the similarity degree can be measured by the distance between the feature points. For example, the smaller the Euclidean distance, the greater the similarity degree.

[0070] Exemplarily, the server can calculate and determine the similarity degree between the three-dimensional feature points and the corresponding two-dimensional feature points. For example, the server can obtain the similarity by determining the feature vector. For example, the similarity degree is characterized by the Euclidean distance; the similarity degree can also be determined through geometric features. For example, the similarity degree can be judged by the relative position relationship between the feature points.

[0071] Optionally, the server may perform error calculation on the projected pixel coordinates of the three-dimensional feature point and the two-dimensional pixel coordinates of the two-dimensional feature point to obtain the Euclidean distance between the three-dimensional feature point and the two-dimensional feature point, or may characterize the similarity between the three-dimensional feature point and the two-dimensional feature point by the Euclidean distance.

[0072] S104. In the same depth distance of the three-dimensional image area, determine the number of three-dimensional feature points corresponding to a similarity less than a preset similarity threshold; if the number is greater than the preset number threshold, eliminate the three-dimensional feature points with a similarity less than the preset similarity threshold.

[0073] The preset similarity threshold may refer to a threshold preset for the similarity, and the preset quantity threshold may refer to a threshold preset for the statistical quantity.

[0074] Exemplarily, for the three-dimensional feature points at the same depth of the three-dimensional image area, the number of three-dimensional feature points corresponding to the preset similarity threshold value at the depth is counted. If the number of three-dimensional feature points with too low similarity counted at the depth is greater than the preset number threshold, it can be explained that the number of three-dimensional feature points at the depth with too low similarity to the corresponding two-dimensional feature points is too large, and the three-dimensional feature points corresponding to the similarity less than the preset similarity threshold value at the current depth can be eliminated, which is conducive to image fusion based on eliminating the three-dimensional feature points with too low similarity, thereby improving the quality of image fusion.

[0075] Optionally, if the similarity is represented by Euclidean distance, then the number of three-dimensional feature points whose statistical similarity is less than a preset similarity threshold may be the number of three-dimensional feature points whose statistical Euclidean distance is greater than a preset Euclidean distance threshold.

[0076] S105 , performing image fusion based on the three-dimensional feature points that are not eliminated and the corresponding two-dimensional feature points to obtain a fused image.

[0077] For example, the server may remove the three-dimensional feature points that do not meet the similarity requirement, and may fuse the three-dimensional image and the two-dimensional image based on the retained three-dimensional feature points and their corresponding two-dimensional feature points to obtain a fused multimodal data image. In this way, the fusion of multimodal images after removing the three-dimensional feature points with large errors is conducive to improving the quality of the fused image.

[0078] In this embodiment, the three-dimensional image region corresponding to the three-dimensional image and the two-dimensional image region corresponding to the two-dimensional image are obtained by dividing the region according to the same preset shape; and by determining the similarity between the three-dimensional feature points and the two-dimensional feature points at the same depth distance, it is possible to judge the number of three-dimensional feature points with a similarity less than the preset similarity threshold at a certain depth distance. In this way, when the number is greater than the preset number threshold, the three-dimensional feature points corresponding to the feature with a similarity less than the preset similarity threshold can be further removed. That is, when it is judged that there are many three-dimensional feature points with too large fusion errors in a certain three-dimensional image region, the three-dimensional feature points with too large fusion errors can be removed, so as to improve the fusion accuracy of the three-dimensional features and the two-dimensional feature points, and further improve the image fusion quality.

[0079] In an exemplary embodiment, Figure 2 is a schematic flow chart of the fusion method of multi-modal images provided by the embodiments of the present application Figure 2 , as Figure 2 shown, on the basis of Figure 1 , an exemplary description is made of the steps of the fusion method of multi-modal images. In the step of S103, obtaining the similarity between the three-dimensional feature points and the two-dimensional feature points includes S201 to S202, where:

[0080] S201, obtain the projected pixel coordinates corresponding to the three-dimensional feature points, and determine the two-dimensional pixel coordinates of the two-dimensional feature points.

[0081] S202, based on the projected pixel coordinates and the two-dimensional pixel coordinates, obtain the similarity.

[0082] Among them, the projected pixel coordinates may refer to the pixel coordinates obtained by projecting the three-dimensional feature point data into the camera coordinate system, and the projected pixel coordinates may be two-dimensional coordinates. The two-dimensional pixel coordinates may refer to the coordinates of the two-dimensional feature points in the camera coordinate system.

[0083] Exemplarily, the server can project and convert the coordinate data of the three-dimensional feature points by using the parameters of the camera device to obtain the projected pixel coordinates in the camera coordinate system. And the server can directly obtain the two-dimensional pixel coordinates of the two-dimensional feature points based on the two-dimensional image data. In this way, the server can use the projected pixel coordinates and the two-dimensional pixel coordinates to calculate the pixel error by using error calculation such as the Euclidean distance formula, and can use the pixel error to characterize the similarity between the three-dimensional feature points and the two-dimensional feature points.

[0084] In this embodiment, by determining the projected pixel coordinates corresponding to the three-dimensional feature points and the two-dimensional pixel coordinates of the two-dimensional feature points, it is possible to calculate the pixel error based on the projected pixel coordinates and the two-dimensional pixel coordinates to obtain the similarity. In this way, the similarity between the three-dimensional feature points and the two-dimensional feature points can be accurately obtained through the pixel error.

[0085] In an exemplary embodiment, Figure 3 is a schematic flow chart of the multi-modal image fusion method provided by the embodiment of the present application Figure 3 As shown in Figure 3 shown, based on Figure 1 and Figure 2 an exemplary description of the steps of the multi-modal image fusion method is given. In the step of S201, obtaining the projection pixel coordinates corresponding to the three-dimensional feature points includes S301 and S302, where:

[0086] S301. Based on the three-dimensional feature points at the current depth and the two-dimensional feature points corresponding to the three-dimensional feature points, obtain the external parameter calibration parameters corresponding to the current depth; where the current depth is any depth.

[0087] S302. Use the external parameter calibration parameters to project the three-dimensional feature points at the current depth to obtain the projection pixel coordinates of the three-dimensional feature points.

[0088] Among them, the external parameter calibration parameters can refer to the parameters of the relative position and attitude relationship between the coordinate system of a three-dimensional image device (such as a lidar) and the camera coordinate system. For example, it can be a matrix of the relative position and attitude relationship between the two. As an example, the external parameter calibration matrix usually can include a rotation matrix and a translation vector. The current depth can refer to any depth distance, that is, for the three-dimensional feature points at any depth, the method of this embodiment can be used for processing.

[0089] Exemplarily, for any depth, the server can calculate the external parameter calibration parameters corresponding to the current depth through multiple three-dimensional feature points at the current depth and the two-dimensional feature points corresponding to the three-dimensional feature points, and can use the external parameter calibration parameters at the current depth to project each three-dimensional feature point at the current depth to obtain the projection pixel coordinates of each three-dimensional feature point.

[0090] In this embodiment, through the three-dimensional feature points at the current depth and their corresponding two-dimensional feature points, the external parameter calibration parameters corresponding to the current depth can be determined, and by using the external parameter calibration parameters corresponding to the current depth, the projection can be effectively performed to obtain the projection pixel coordinates of the three-dimensional feature points, which is beneficial to determining the external parameter calibration parameters, thereby being able to accurately determine the pixel error and improve the image fusion quality.

[0091] Optionally, in the step of S302, using the external parameter calibration parameters to project the three-dimensional feature points at the current depth to obtain the projection pixel coordinates of the three-dimensional feature points may specifically include:

[0092] Obtain the internal parameters of the camera corresponding to the camera device of the two-dimensional image.

[0093] Using the external parameter calibration parameters and the camera internal parameters, project the three-dimensional feature points of the current depth to obtain the projected pixel coordinates of the three-dimensional feature points.

[0094] Among them, the camera internal parameters can be the camera parameters used to convert three-dimensional coordinates.

[0095] Exemplarily, the server can directly determine the camera internal parameters of the camera device set on the operation robot based on the data of the camera device. In this way, the server can further use the external parameter calibration parameters and the camera internal parameters of the current depth to convert the two-dimensional coordinates of the three-dimensional feature points of the current depth to obtain the projected pixel coordinates of the three-dimensional feature points.

[0096] In this embodiment, by using the camera internal parameters of the camera device set on the operation robot for coordinate projection, the projected pixel coordinates of the three-dimensional feature points can be effectively obtained, which is beneficial to the pixel error calculation between the three-dimensional features and the two-dimensional feature points, and further beneficial to improving the quality of image fusion.

[0097] In some specific embodiments, for the three-dimensional image region divide it according to different depth distances to obtain where and are the set of 2D feature points and the set of 3D feature points of the i-th region with a depth distance of d, respectively.

[0098] Using the set of feature points perform external parameter calibration on the three-dimensional image region with different depth distances d to obtain the calibration matrix for different distances where , is the number of three-dimensional image regions, and the calibration matrix for different distances represents the external parameter calibration matrix with a depth distance of d.

[0099] If the depth distance is d, the i-th region and the j-th 3D feature point , using for projection to obtain the pixel coordinates , that is:

[0100]

[0101] where is the projection function for converting 3D feature points to 2D coordinates; is the camera internal parameter.

[0102] In this embodiment, an exemplary expression is provided, which can achieve the projection coordinate conversion of three-dimensional feature points, thereby facilitating the calculation of pixel errors between three-dimensional features and two-dimensional feature points, and further facilitating the improvement of the quality of image fusion.

[0103] Optionally, in the step of S202, based on the projected pixel coordinates and the two-dimensional pixel coordinates, the similarity degree is obtained, which may specifically include:

[0104] Determine the Euclidean distance between the projected pixel coordinates and the two-dimensional pixel coordinates.

[0105] Based on the Euclidean distance, obtain the similarity degree.

[0106] Exemplarily, for and The Euclidean distance formula is used to calculate the error, and the pixel error is stored in the error set Furthermore, it can be further compared with the similarity degree threshold through the pixel errors in the error set . Among them, is the j-th 2D feature point; is the projected pixel coordinate of the above 3D feature point.

[0107] Optionally, by traversing all depth distances d, is obtained, where .

[0108] In this embodiment, through the Euclidean distance between the projected pixel coordinates and the two-dimensional pixel coordinates, the similarity degree between the three-dimensional feature and the two-dimensional feature point can be accurately and effectively obtained, thereby improving the quality of image fusion.

[0109] In an exemplary embodiment, Figure 4 is the flow schematic of the multi-modal image fusion method provided by the embodiment of the present application Figure 4 , as Figure 4 shown, on the basis of Figure 1 , the steps of the multi-modal image fusion method are illustrated exemplarily. Among them, both the three-dimensional image and the two-dimensional image are divided into a central region and an outer ring region; the outer ring region of the three-dimensional image includes multiple three-dimensional image regions; the outer ring region of the two-dimensional image includes multiple two-dimensional image regions; in the step of S105, based on the uneliminated three-dimensional feature points and the corresponding two-dimensional feature points, image fusion is performed to obtain a fused image, including S401 to S403, where:

[0110] S401. Determine the three-dimensional feature points and two-dimensional feature points corresponding to the central region, and based on the three-dimensional feature points and two-dimensional feature points corresponding to the central region, determine the external parameter calibration parameters of the central region;

[0111] S402. Determine the external parameter calibration parameters of the outer ring area based on the uneliminated three-dimensional feature points in the outer ring area and the two-dimensional feature points corresponding to the uneliminated three-dimensional feature points.

[0112] S403. Perform image fusion on the central area and the outer ring area respectively according to the external parameter calibration parameters of the central area and the external parameter calibration parameters of the outer ring area to obtain a fused image.

[0113] Exemplarily, Figure 5 FIG. is a schematic diagram of image area division provided by an embodiment of the present application. As Figure 5 shown, the image can be divided into areas in a "hui"-shaped area. Among them, the center of the "hui"-shaped area can be used as a direct calibration area, with a total of 1 central area, and the set of feature points within the area is , and the outer ring of the "hui"-shaped area is a partition calibration area, with a total of sub-areas. It can be understood that the image area has a total of sub-areas, and the sets of feature points within each sub-area in the outer ring area are respectively , obtaining the feature point set . The set of feature points within the area , is the set of 2D feature points within area , and is the set of 3D feature points corresponding to .

[0114] Schematically, the external parameter calibration parameters of the central area can be determined through the three-dimensional feature points of the three-dimensional image corresponding to the central area and the two-dimensional feature points of the two-dimensional image. And the external parameter calibration parameters of the outer ring area can be determined through the uneliminated three-dimensional feature points in the outer ring area and the two-dimensional feature points corresponding to the uneliminated three-dimensional feature points. It can be understood that the above embodiment of eliminating three-dimensional feature points can be implemented for the outer ring area.

[0115] Furthermore, the images of the central area and the outer ring area can be respectively subjected to image fusion according to the external parameter calibration parameters of the central area and the external parameter calibration parameters of the outer ring area to obtain a fused image.

[0116] In this embodiment, by using the external parameter calibration parameters of the central area and the external parameter calibration parameters of the outer ring area, image fusion is respectively performed on the central area and the outer ring area, and a fused image can be obtained. In this way, the three-dimensional feature points in the outer ring area can be eliminated. Since the three-dimensional feature points in the central area are usually valid feature points, while there may be invalid feature points in the three-dimensional feature points in the outer ring area, eliminating the invalid feature points in the outer ring area without processing the feature points in the central area. In this way, through differential processing of specific areas in this embodiment, the quality of image fusion can be improved while the efficiency of image fusion is improved.

[0117] Optionally, use and respectively to calibrate the direct calibration area (central area) and the partition calibration area (outer ring area), and obtain the external parameter calibration matrix of the "hui"-shaped central area and the external parameter calibration matrix of the "hui"-shaped outer ring area . Further, through and , image fusion can be performed on the central area and the outer ring area respectively to obtain a fused image.

[0118] In an exemplary embodiment, it is applied to a power robot; the three-dimensional image is a point cloud image;

[0119] Obtain the three-dimensional image and the two-dimensional image to be subjected to image fusion, including:

[0120] Obtain a point cloud image through a lidar installed on the power robot;

[0121] Obtain a two-dimensional image through a camera device installed on the power robot.

[0122] Exemplarily, the power robot can adopt a partition calibration data acquisition device to collect the laser point cloud and the camera image and transmit them to the server, and the server can obtain the point cloud image collected by the lidar and the two-dimensional image collected by the camera device.

[0123] Among them, the camera device and the lidar in the partition calibration data acquisition device can be connected and fixed to the turntable through a tooling, and the turntable can rotate around the horizontal and vertical directions.

[0124] In this embodiment, it can be applied to a power robot. A point cloud image is obtained through a lidar installed on the power robot; and a two-dimensional image is obtained through a camera device installed on the power robot. In this way, it is beneficial to improve the quality of the fusion of the laser point cloud and the camera image.

[0125] In some specific embodiments, Figure 6 is a schematic flow of the multi-modal image fusion method provided by the embodiment of the present application Figure 5 , as Figure 6 shown, on the basis of Figure 1 , an exemplary description of the multi-modal image fusion method is made. Among them, the multi-modal image fusion method may include:

[0126] S601. Collect and extract the feature points of the camera image and the laser point cloud;

[0127] Among them, a partition calibration data acquisition device is used to collect the laser point cloud and the camera image and transmit them to the server.

[0128] S602. Perform a "return" shaped partitioning on the feature points; where each region can have x points;

[0129] Optionally, the server can perform feature point extraction to obtain 2D and 3D feature point pairs for extrinsic parameter calibration (e.g., the corresponding 2D feature points and 3D feature points can be used as feature point pairs). Among them, the camera and lidar in the partitioning calibration data acquisition device can be connected and fixed to the pan-tilt through a tooling. The pan-tilt can rotate around the horizontal and vertical directions. The 2D and 3D feature point pairs can be a set of image feature points (2D feature points) and point cloud feature points (3D feature points). The set of image feature points can be a set of 2D feature points in the pixel coordinate system, and the set of point cloud feature points can be a set of 3D feature points in the radar coordinate system.

[0130] S603. Solve the extrinsic parameter calibration matrix T of the depth distance d d (d = 1, …, A);

[0131] Among them, using the set of feature points , perform extrinsic parameter calibration on the 3D image regions with different depth distances d to obtain the sub-distance calibration matrix , where , is the number of 3D image regions, and the sub-distance calibration matrix represents the extrinsic parameter calibration matrix with a depth distance of d.

[0132] S604. Use T d to back-project onto the image to calculate the pixel error ;

[0133] Exemplarily, if the depth distance is d, the i-th region and the j-th 3D feature point can be projected using to obtain the pixel coordinates , that is:

[0134]

[0135] Among them, is the projection function for converting 3D feature points into 2D coordinates; is the camera internal parameter.

[0136] For and , use the Euclidean distance formula to calculate the error, obtain the pixel error , and store it in the error set . Further, the pixel error in the error set can be compared with the similarity degree threshold. Among them, is the jth 2D feature point.

[0137] S605, Save To error set .

[0138] S606, determine whether d is greater than A; if not, execute S607; if so, execute S608.

[0139] S607, d=d+1.

[0140] S608, Calculation The number of points in the ;

[0141] Among them, comparison With error threshold , and record middle Greater than The number of ;

[0142] S609, if Greater than the minimum confidence number τ ( ), then remove ;

[0143] Among them, when Greater than the minimum confidence number ( ) when the point Eliminate;

[0144] S610, determine whether j is greater than x, if not, execute S611; if so, execute S612.

[0145] S611, d=1, j=j+1.

[0146] S612, determine whether i is greater than m×np×q, if not, execute S613, if yes, execute S614;

[0147] S613, d=1, j=1, i=i+1;

[0148] S614, calibrate the central area and the outer ring area of ​​the “回” shape respectively.

[0149] In this embodiment, a "hui"-shaped zoning strategy is adopted to separate the edge region (with large fusion error) from the central region data (with small fusion error), and a multi-region joint optimization algorithm is used to accurately eliminate invalid feature points, perform regional fusion on the lidar point cloud and the image, effectively solve the problem of large fusion error between the lidar point cloud and the image in the edge region, achieve accurate and stable fusion of the lidar point cloud and the image in a large scene and at a long distance, and enable high-precision and highly reliable fusion of the lidar point cloud and the image in the power operation scenario.

[0150] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are sequentially shown according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise clearly stated in this article, there is no strict order limit for the execution of these steps, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or steps in other steps.

[0151] The fusion device for multi-modal images provided in the embodiments of the present application will be described below. The fusion device for multi-modal images has the same inventive concept as the above-described fusion method for multi-modal images. The implementation solutions provided by this device to solve problems are similar to the implementation solutions described in the above method. Therefore, the specific limitations in one or more embodiments of the fusion device for multi-modal images provided below can refer to the limitations on the fusion method for multi-modal images in the above text. The fusion device for multi-modal images described below and the fusion method for multi-modal images described above can be correspondingly referred to each other, and will not be elaborated here.

[0152] In an exemplary embodiment, Figure 7 is a schematic structural diagram of a fusion device for multi-modal images provided in an embodiment of the present application. As Figure 7 shown, the fusion device 70 for multi-modal images includes: an image region determination module 710, a feature point determination module 720, a similarity degree determination module 730, a quantity determination module 740, and an image fusion module 750, where:

[0153] The image region determination module 710 is configured to obtain a three-dimensional image and a two-dimensional image to be subjected to image fusion, and perform region division according to the same preset shape to obtain a three-dimensional image region corresponding to the three-dimensional image and a two-dimensional image region corresponding to the two-dimensional image.

[0154] A feature point determination module 720, configured to determine three-dimensional feature points located at the same depth distance in a three-dimensional image region; and determine two-dimensional feature points corresponding to the three-dimensional feature points in a two-dimensional image region.

[0155] A similarity degree determination module 730, configured to obtain the similarity degree between the three-dimensional feature points and the two-dimensional feature points.

[0156] A quantity determination module 740, configured to determine the quantity of three-dimensional feature points corresponding to a similarity degree less than a preset similarity degree threshold in the same depth distance of the three-dimensional image region; if the quantity is greater than a preset quantity threshold, then eliminate the three-dimensional feature points with a similarity degree less than the preset similarity degree threshold.

[0157] An image fusion module 750, configured to perform image fusion based on the uneliminated three-dimensional feature points and the corresponding two-dimensional feature points to obtain a fused image.

[0158] In an exemplary embodiment, the similarity degree determination module is configured to obtain the projected pixel coordinates corresponding to the three-dimensional feature points, and determine the two-dimensional pixel coordinates of the two-dimensional feature points; based on the projected pixel coordinates and the two-dimensional pixel coordinates, obtain the similarity degree.

[0159] In an exemplary embodiment, the similarity degree determination module is configured to obtain the external parameter calibration parameters corresponding to the current depth based on the three-dimensional feature points at the current depth and the two-dimensional feature points corresponding to the three-dimensional feature points; where the current depth is any depth; use the external parameter calibration parameters to project the three-dimensional feature points at the current depth to obtain the projected pixel coordinates of the three-dimensional feature points.

[0160] In an exemplary embodiment, the similarity degree determination module is configured to obtain the internal parameters of the camera device corresponding to the two-dimensional image; use the external parameter calibration parameters and the internal parameters of the camera to perform projection processing on the three-dimensional feature points at the current depth to obtain the projected pixel coordinates of the three-dimensional feature points.

[0161] In an exemplary embodiment, the similarity degree determination module is configured to determine the Euclidean distance between the projected pixel coordinates and the two-dimensional pixel coordinates; based on the Euclidean distance, obtain the similarity degree.

[0162] In an exemplary embodiment, both the three-dimensional image and the two-dimensional image are divided into a central region and an outer ring region; the outer ring region of the three-dimensional image includes a plurality of three-dimensional image regions; the outer ring region of the two-dimensional image includes a plurality of two-dimensional image regions.

[0163] The image fusion module is used to determine the three-dimensional feature points and two-dimensional feature points corresponding to the central region, and based on the three-dimensional feature points and two-dimensional feature points corresponding to the central region, determine the external parameter calibration parameters of the central region; based on the three-dimensional feature points not excluded in the outer ring region and the two-dimensional feature points corresponding to the three-dimensional feature points not excluded, determine the external parameter calibration parameters of the outer ring region; according to the external parameter calibration parameters of the central region and the external parameter calibration parameters of the outer ring region, perform image fusion on the central region and the outer ring region respectively to obtain a fused image.

[0164] In an exemplary embodiment, the method is applied to a power robot; the three-dimensional image is a point cloud image. The image region determination module is used to obtain a point cloud image through a lidar installed on the power robot; and obtain a two-dimensional image through a camera device installed on the power robot.

[0165] In an exemplary embodiment, the present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by one or more processors, one or more processors are caused to execute the steps of the multi-modal image fusion method as described in any one of the above embodiments.

[0166] In an exemplary embodiment, the present application also provides a computer device, in which a computer program is stored. When the computer-readable instructions are executed by one or more processors, one or more processors are caused to execute the steps of the multi-modal image fusion method as described in any one of the above embodiments.

[0167] In an exemplary embodiment, the present application also provides a computer program product, including a computer program, which when executed by a processor implements the steps of the multi-modal image fusion method as described in any one of the above embodiments.

[0168] Schematically, as Figure 8 shown, Figure 8 is a schematic internal structure diagram of a computer device provided by an embodiment of the present application. The computer device 800 can be provided as a server. Referring to Figure 8 , the computer device 800 includes a processing component 802, which further includes one or more processors, and memory resources represented by a memory 801 for storing instructions executable by the processing component 802, such as application programs. The application programs stored in the memory 801 can include one or more modules each corresponding to a set of instructions. In addition, the processing component 802 is configured to execute instructions to perform the text recognition method of any of the above embodiments.

[0169] The computer device 800 may further include a power supply component 803 configured to perform power management of the computer device 800, a wired or wireless network interface 804 configured to connect the computer device 800 to a network, and an input / output (I / O) interface 805. The computer device 800 may operate based on an operating system stored in the memory 801, such as Windows Server TM, Mac OS XTM, Unix TM, Linux TM, Free BSDTM, or the like.

[0170] Those skilled in the art can understand that Figure 8 the structure shown in is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different component arrangement.

[0171] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.

[0172] Finally, it should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element.

[0173] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The various embodiments can be combined as needed, and the same or similar parts can be referred to each other.

[0174] The foregoing description of the disclosed embodiments enables those skilled in the art to practice or use the present application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Thus, the present application is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A multimodal image fusion method, characterized in that: The method comprises: Acquire a three-dimensional image and a two-dimensional image to be fused, and divide the regions according to the same preset shape to obtain a three-dimensional image region corresponding to the three-dimensional image and a two-dimensional image region corresponding to the two-dimensional image; Determine three-dimensional feature points located at the same depth distance in the three-dimensional image area; and determine that the three-dimensional feature points correspond to two-dimensional feature points in the two-dimensional image area; Obtaining a degree of similarity between the three-dimensional feature point and the two-dimensional feature point; In the same depth distance of the three-dimensional image area, determine the number of three-dimensional feature points corresponding to a similarity less than a preset similarity threshold; if the number is greater than the preset number threshold, remove the three-dimensional feature points with a similarity less than the preset similarity threshold; Based on the unremoved three-dimensional feature points and the corresponding two-dimensional feature points, image fusion is performed to obtain a fused image; Wherein, obtaining the similarity between the three-dimensional feature point and the two-dimensional feature point includes: Obtaining the projection pixel coordinates corresponding to the three-dimensional feature points, and determining the two-dimensional pixel coordinates of the two-dimensional feature points; Obtaining the degree of similarity based on the projected pixel coordinates and the two-dimensional pixel coordinates; The step of obtaining the projection pixel coordinates corresponding to the three-dimensional feature points includes: Based on the three-dimensional feature points of the current depth and the two-dimensional feature points corresponding to the three-dimensional feature points, obtain the external parameter calibration parameters corresponding to the current depth; wherein the current depth is any depth; Projecting the three-dimensional feature points at the current depth using the external parameter calibration parameters to obtain the projected pixel coordinates of the three-dimensional feature points; Wherein, the three-dimensional image and the two-dimensional image are both divided into a central area and an outer ring area; the outer ring area of ​​the three-dimensional image includes a plurality of the three-dimensional image areas; the outer ring area of ​​the two-dimensional image includes a plurality of the two-dimensional image areas; The step of fusing the images based on the unremoved three-dimensional feature points and the corresponding two-dimensional feature points to obtain a fused image includes: Determine the three-dimensional feature points and the two-dimensional feature points corresponding to the central area, and determine the extrinsic calibration parameters of the central area based on the three-dimensional feature points and the two-dimensional feature points corresponding to the central area; Determine the extrinsic calibration parameters of the outer ring area based on the three-dimensional feature points that are not eliminated in the outer ring area and the two-dimensional feature points corresponding to the three-dimensional feature points that are not eliminated; According to the extrinsic calibration parameters of the central area and the extrinsic calibration parameters of the outer ring area, the images of the central area and the outer ring area are fused respectively to obtain the fused image.

2. The method according to claim 1, characterized in that The method of projecting the three-dimensional feature point at the current depth by using the external parameter calibration parameter to obtain the projected pixel coordinates of the three-dimensional feature point includes: Obtaining camera internal parameters corresponding to the camera device of the two-dimensional image; The three-dimensional feature points at the current depth are projected using the external calibration parameters and the camera internal parameters to obtain the projected pixel coordinates of the three-dimensional feature points.

3. The method according to claim 1, characterized in that The obtaining the similarity based on the projection pixel coordinates and the two-dimensional pixel coordinates includes: Determining the Euclidean distance between the projected pixel coordinates and the two-dimensional pixel coordinates; Based on the Euclidean distance, the similarity degree is obtained.

4. The method according to any one of claims 1 to 3, characterized in that: The method is applied to electric robots; The three-dimensional image is a point cloud image; The step of acquiring the three-dimensional image and the two-dimensional image to be fused includes: Acquiring the point cloud image by using a laser radar installed on the electric robot; The two-dimensional image is acquired by a camera device installed on the electric robot.

5. A multimodal image fusion device, characterized in that: For executing the steps of the method according to any one of claims 1 to 4, the device comprises: An image region determination module is used to obtain a 3D image and a 2D image to be fused, and divide the regions according to the same preset shape to obtain a 3D image region corresponding to the 3D image and a 2D image region corresponding to the 2D image; A feature point determination module, used to determine three-dimensional feature points located at the same depth distance in the three-dimensional image area; and determine that the three-dimensional feature points correspond to two-dimensional feature points in the two-dimensional image area; A similarity determination module, used to obtain the similarity between the three-dimensional feature point and the two-dimensional feature point; a quantity determination module, configured to determine the number of three-dimensional feature points corresponding to a similarity less than a preset similarity threshold value at the same depth distance of the three-dimensional image area; if the number is greater than the preset number threshold value, then eliminate the three-dimensional feature points whose similarity is less than the preset similarity threshold value; The image fusion module is used to perform image fusion based on the unremoved three-dimensional feature points and the corresponding two-dimensional feature points to obtain a fused image.

6. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 4 are implemented.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 4 are implemented.

Citation Information

Patent Citations

  • Airborne sounding radar and multispectral satellite image registration method based on feature fusion

    CN112686935A

  • Danger alarm method, device and equipment based on radar induction and storage medium

    CN119091433A