Panoramic image and laser point cloud data fusion method and system based on laser SLAM

By combining feature matching and multimodal deep learning model methods, the efficient fusion problem of laser point cloud data and panoramic image data is solved, the registration accuracy and fusion quality are improved, and the dynamic environment is adapted to the dynamic environment, especially for autonomous driving and robot navigation.

CN120387934APending Publication Date: 2025-07-29CHINA ELECTRIC POWER RESEARCH INSTITUTE CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510278266.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The existing registration and fusion methods for laser point cloud data and panoramic image data have problems such as low accuracy, low efficiency or poor adaptability, making it difficult to achieve efficient environmental perception and modeling.

Method used

The traditional method based on feature matching is used to combine it with the multimodal deep learning model, and data fusion is carried out through feature registration, global optimization algorithm and multimodal deep learning model, including voxel filtering, bilateral filtering, scale-invariant feature transformation, random sampling consistency algorithm and least squares optimization graph model, and data processing is carried out by combining convolutional neural networks and deep generation adversarial networks.

Benefits of technology

It significantly improves the registration accuracy and fusion quality of panoramic images and laser point cloud data, improves the computing efficiency in large-scale environments, adapts to dynamically changing scenarios, and provides accurate environmental perception, especially for autonomous driving and robot navigation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120387934A_ABST
    Figure CN120387934A_ABST
Patent Text Reader

Abstract

The invention discloses a panoramic image and laser point cloud data fusion method and system based on laser SLAM, and the method comprises the steps: carrying out the feature registration based on obtained laser point cloud data and corresponding panoramic image data, and determining feature matching point pairs; constructing a graph model by using a global optimization algorithm based on the feature matching point pairs; and based on the graph model, performing data fusion by using a multi-modal deep learning model to obtain fused data. According to the method, a traditional method based on feature matching is organically combined with a frontier multi-mode deep learning model, the registration precision and fusion quality of the panoramic image and the laser point cloud data can be improved, the calculation efficiency in a large-scale environment is remarkably improved, the method adapts to dynamically changing scenes, and the method is suitable for large-scale application. And particularly, accurate environment perception is provided for applications such as automatic driving and robot navigation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data fusion, and more specifically, to a method and system for fusing panoramic images and laser point cloud data based on laser SLAM. Background Art

[0002] With the rapid development of autonomous driving, robotic navigation, and environmental modeling, laser technology (Simultaneous Localization and Mapping, SLAM) has been widely used as a high-precision sensor data processing method for building environmental maps and positioning. However, in practical applications, how to efficiently register and fuse laser point cloud data with panoramic image data remains a technical challenge that needs to be solved.

[0003] Laser SLAM technology captures laser point cloud data of the environment, reflecting the spatial geometry, but lacks rich texture and color information. Panoramic images, on the other hand, provide rich visual information but are limited in depth perception. Efficiently registering and fusing these two methods enables more comprehensive and accurate environmental perception and modeling. However, existing registration and fusion methods suffer from low accuracy, inefficiency, and poor adaptability. Summary of the Invention

[0004] The present invention proposes a method and system for fusing panoramic images and laser point cloud data based on laser SLAM to solve the problem of how to efficiently fuse laser point cloud data and panoramic images.

[0005] In order to solve the above problems, according to one aspect of the present invention, a method for fusing panoramic images and laser point cloud data based on laser SLAM is provided, the method comprising:

[0006] Perform feature registration based on the acquired laser point cloud data and the corresponding panoramic image data to determine feature matching point pairs;

[0007] Based on the feature matching point pairs, a graph model is constructed using a global optimization algorithm;

[0008] Based on the graph model, a multimodal deep learning model is used to perform data fusion to obtain fused data.

[0009] Preferably, before registering the acquired laser point cloud data and the corresponding panoramic image data, the method further comprises:

[0010] Performing filtering processing on the laser point cloud data using a voxel filtering algorithm;

[0011] The panoramic image data is subjected to denoising and enhancement processing using a bilateral filtering algorithm.

[0012] Preferably, the method further includes:

[0013] When performing feature registration, the scale-invariant feature transform algorithm is used to extract feature points in the panoramic image data.

[0014] Preferably, the method further includes:

[0015] The random sample consensus algorithm is used to screen the feature matching point pairs and eliminate the incorrect feature matching point pairs.

[0016] Preferably, constructing a graph model based on the feature matching point pairs by using a global optimization algorithm includes:

[0017] Using nodes to represent laser point cloud frames or image frames, using edges to represent the constraint relationships between frames, the weights of the edges are determined according to the registration error, and the graph model is optimized by using the least squares method;

[0018] Wherein, let the node set V in the graph be {v1, v2,..., v n}, the edge set E be {e ij |i, j = 1, 2,..., n}, the weight w ij of the edge e ij , for the pose T i corresponding to the node v i , the optimization objective function is expressed as:

[0019]

[0020] Wherein, n is the number of nodes; T j is the pose corresponding to the node v j .

[0021] Preferably, for the multi-modal deep learning model, a convolutional neural network model is adopted, and the network structure includes: a convolutional layer, a pooling layer, and a fully connected layer. Let the input panoramic image be I, and the feature map after passing through the convolutional layer be F l (I), and the convolution operation formula is:

[0022]

[0023] Wherein, ω l,k is the kth convolutional kernel of the lth layer, * represents the convolution operation, b l is the bias term, and σ is the activation function.

[0024] Preferably, the multimodal deep learning model adopts a deep generative adversarial network model, including: a generator G and a discriminator D. The generator G converts the laser point cloud data into a representation G(P) similar to image features. The discriminator D is used to distinguish real image features from the features generated by the generator. The loss function of the generator is expressed as: The loss function of the discriminator is: where P data represents the distribution of the laser point cloud data, and P image represents the distribution of the image data.

[0025] According to another aspect of the present invention, a panoramic image and laser point cloud data fusion system based on laser SLAM is provided. The system includes:

[0026] A feature matching unit for performing feature registration based on the acquired laser point cloud data and the corresponding panoramic image data to determine feature matching point pairs;

[0027] A graph model construction unit for constructing a graph model based on the feature matching point pairs by using a global optimization algorithm;

[0028] A data fusion unit for performing data fusion based on the graph model by using a multimodal deep learning model to obtain fused data.

[0029] Preferably, the system further includes:

[0030] A preprocessing unit for filtering the laser point cloud data by using a voxel filtering algorithm and denoising and enhancing the panoramic image data by using a bilateral filtering algorithm before registering the acquired laser point cloud data and the corresponding panoramic image data.

[0031] Preferably, the feature matching unit further includes:

[0032] When performing feature registration, using the scale-invariant feature transform algorithm to extract feature points in the panoramic image data.

[0033] Preferably, the system further includes:

[0034] A screening unit for screening the feature matching point pairs by using the random sample consensus algorithm to remove incorrect feature matching point pairs.

[0035] Preferably, the graph model construction unit constructs a graph model based on the feature matching point pairs by using a global optimization algorithm, including:

[0036] Using nodes to represent laser point cloud frames or image frames, using edges to represent the constraint relationships between frames, determining the weights of the edges according to the registration error, and optimizing the graph model by using the least squares method;

[0037] Among them, let the node set V in the figure be {v1, v2, …, v n}, the edge set E be {e ij | i, j = 1, 2, …, n}, the weight w ij of the edge e ij , for the pose T i corresponding to the node v i , the optimization objective function is expressed as:

[0038]

[0039] Among them, n is the number of nodes; T j is the pose corresponding to the node v j .

[0040] Preferably, in the data fusion unit, the multi-modal deep learning model adopts a convolutional neural network model, and the network structure includes: a convolutional layer, a pooling layer, and a fully connected layer. Let the input panoramic image be I, and the feature map after passing through the convolutional layer be expressed as F l (I), and the convolution operation formula is:

[0041]

[0042] Among them, ω l,k is the k-th convolutional kernel of the l-th layer, * represents the convolution operation, b l is the bias term, and σ is the activation function.

[0043] Preferably, in the data fusion unit, the multi-modal deep learning model adopts a deep generative adversarial network model, including: a generator G and a discriminator D. The generator G converts the laser point cloud data into a representation G(P) similar to image features, and the discriminator D is used to distinguish real image features from the features generated by the generator. The loss function of the generator is expressed as: The loss function of the discriminator is: Among them, P data represents the distribution of the laser point cloud data, and P image represents the distribution of the image data.

[0044] Based on another aspect of the present invention, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements any one of the above-mentioned methods for fusing panoramic images and laser point cloud data based on laser SLAM.

[0045] Based on another aspect of the present invention, the present invention provides an electronic device, including:

[0046] The above-mentioned computer-readable storage medium; and

[0047] One or more processors are configured to execute the program in the computer-readable storage medium.

[0048] The present invention provides a method and system for fusing panoramic images and laser point cloud data based on laser SLAM, comprising: performing feature registration based on the acquired laser point cloud data and the corresponding panoramic image data to determine feature matching point pairs; constructing a graph model based on the feature matching point pairs using a global optimization algorithm; and performing data fusion based on the graph model using a multimodal deep learning model to obtain fused data. By organically combining traditional feature matching-based methods with cutting-edge multimodal deep learning models, the method of the present invention can improve the registration accuracy and fusion quality of panoramic images and laser point cloud data, significantly enhance computational efficiency in large-scale environments, adapt to dynamically changing scenarios, and, in particular, provide accurate environmental perception for applications such as autonomous driving and robotic navigation. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] A more complete understanding of exemplary embodiments of the present invention may be obtained by referring to the following drawings:

[0050] Figure 1 Flowchart of a panoramic image and laser point cloud data fusion method 100 of the laser SLAM technology according to an embodiment of the present invention;

[0051] Figure 2 Schematic diagram of the structure of a panoramic image and laser point cloud data fusion system 200 using laser SLAM technology according to an embodiment of the present invention. DETAILED DESCRIPTION

[0052] Exemplary embodiments of the present invention will now be described with reference to the accompanying drawings. However, the present invention may be embodied in many different forms and is not limited to the embodiments described herein. These embodiments are provided to provide a thorough and complete disclosure of the present invention and to fully convey the scope of the present invention to those skilled in the art. The terminology used in the exemplary embodiments shown in the accompanying drawings is not intended to limit the present invention. In the accompanying drawings, identical elements are denoted by the same reference numerals.

[0053] Unless otherwise specified, the terms used herein (including technical terms) have the meanings commonly understood by those skilled in the art. In addition, it is understood that terms defined in commonly used dictionaries should be understood to have the same meanings as those in the context of the relevant fields, and should not be understood as idealized or overly formal meanings.

[0054] The object of the present invention is to address the problems existing in the background art and propose a method for fusing panoramic images and laser point cloud data based on laser SLAM, aiming to solve the problems of the huge data volume of laser point clouds and panoramic images, how to achieve fast processing and calculation, and avoid calculation problems in large-scale scenarios.

[0055] Figure 1 It is a flowchart of a method 100 for fusing panoramic images and laser point cloud data of laser SLAM technology according to an embodiment of the present invention. As Figure 1 shown, the method for fusing panoramic images and laser point cloud data based on laser SLAM provided by the embodiment of the present invention can improve the registration accuracy and fusion quality of panoramic images and laser point cloud data by organically combining the traditional method based on feature matching with the advanced multimodal deep learning model, significantly improve the calculation efficiency in large-scale environments, adapt to dynamically changing scenarios, and especially provide accurate environmental perception for applications such as autonomous driving and robot navigation. The method 100 for fusing panoramic images and laser point cloud data based on laser SLAM provided by the embodiment of the present invention starts from step 101. In step 101, feature registration is performed based on the acquired laser point cloud data and the corresponding panoramic image data to determine feature matching point pairs.

[0056] Preferably, before registering the acquired laser point cloud data and the corresponding panoramic image data, the method further includes:

[0057] Filtering the laser point cloud data by using a voxel filtering algorithm;

[0058] Denosing and enhancing the panoramic image data by using a bilateral filtering algorithm.

[0059] Preferably, the method further includes:

[0060] When performing feature registration, extracting feature points in the panoramic image data by using a scale-invariant feature transform algorithm.

[0061] Preferably, the method further includes:

[0062] Screening the feature matching point pairs by using a random sample consensus algorithm to eliminate incorrect feature matching point pairs.

[0063] In the present invention, in the registration method based on feature matching, when extracting the color and texture features of the panoramic image, feature points in the image are extracted by using a scale-invariant feature transform algorithm. Using the geometric features of the laser point cloud and the color and texture features of the panoramic image, an iterative closest point algorithm is used for preliminary registration. In the iterative closest point algorithm, let the laser point cloud set P = {P1, P2,..., P m}, corresponding to the image feature point cloud set Q = {Q1, Q2, ..., Q m}, the initial transformation matrix, the error function of the iterative process is expressed as:

[0064]

[0065] This function quantifies the overall deviation between the transformed laser point cloud and the image feature point cloud from a mathematical perspective. By continuously adjusting the transformation matrix to minimize the error function value, the iterative formula T k+1 =T k +ΔT, then the update method of each iteration is clarified, where

[0066]

[0067] This calculation gradually corrects the transformation direction and step size based on information such as the partial derivatives of the error function with respect to the transformation parameters, ensuring that the point cloud and the image feature point cloud gradually align. For example, in an indoor scene reconstruction project, the ICP algorithm uses an initial rough alignment of a room's laser scan point cloud and the feature point cloud of a simultaneously captured panoramic image. After multiple iterations, the point cloud and image features of objects such as furniture and walls are closely aligned.

[0068] In the present invention, before feature-based registration, a voxel filtering algorithm is used to filter the laser point cloud data, and a bilateral filtering algorithm is used to denoise and enhance the panoramic image.

[0069] Among them, for the pixel I(x,y) in the image, its filtered pixel value is expressed as:

[0070]

[0071] Among them, N(x,y) is the neighborhood of pixel (x,y), and are Gaussian functions in the spatial domain and the range, respectively.

[0072]

[0073] In the present invention, in the registration method based on feature matching, when extracting the geometric features of the laser point cloud, for a point P in the laser point cloud i =(x i ,y i ,z i ), by calculating its local geometric features, such as the normal vector Determine whether it is an edge or corner point. If the normal vector changes significantly, it is determined to be an edge or corner point.

[0074] In the registration method based on feature matching, the random sampling consensus algorithm is used to remove incorrect matching point pairs. Among them, the sampling strategy is: based on the preliminary registration results, the RANSAC algorithm is introduced to purify the set of matching point pairs. The algorithm cleverly randomly extracts multiple sample points from many matching point pairs. Generally speaking, for the rigid transformation registration scenario from two-dimensional plane to two-dimensional plane or three-dimensional space to three-dimensional space, it can meet the needs of building a basic transformation model. Using this small but representative sample point, a temporary transformation model is calculated through specific geometric calculation rules. The internal point screening and model selection process includes: for the remaining other point pairs (P i , Q i ), and rigorously calculate its error e i =∥P i -T s Q i ∥ 2 , and makes a detailed comparison with the preset threshold. If the error is less than the threshold, the point pair is determined to be an inlier, that is, the point pair is well matched under the current temporary transformation model. Due to the influence of many interference factors such as light changes, occlusion, and reflection during the actual data collection process, the initial matching is very likely to be mixed with erroneous matching point pairs. Through hundreds or even more random sampling, model construction and inlier statistical processes, RANSAC finally selects the transformation model with the largest number of inliers, establishes it as the most reliable preliminary registration model, effectively eliminates erroneous matches, and greatly improves the registration accuracy. Taking outdoor building surveying as an example, changes in sunlight angle may cause the building surface to present different features in the image and point cloud, resulting in erroneous matches. RANSAC can accurately identify and remove these anomalies to ensure the accuracy of registration.

[0075] In step 102, a graph model is constructed using a global optimization algorithm based on the feature matching point pairs.

[0076] Preferably, the step of constructing a graph model based on the feature matching point pairs using a global optimization algorithm includes:

[0077] Nodes are used to represent laser point cloud frames or image frames, edges are used to represent the constraint relationship between frames, the weight of the edge is determined according to the registration error, and the least squares method is used to optimize the graph model;

[0078] In which, let the node set V in the graph be {v1,v2,…,v n}, edge set E={e ij |i,j=1,2,…,n}, edge e ij The weight w ij , for node v i The corresponding pose T i , the optimization objective function is expressed as:

[0079]

[0080] where n is the number of nodes; T j is the pose corresponding to node v j .

[0081] In the present invention, to further deepen the registration accuracy, constructing a graph model becomes a key link. In this model, nodes are given profound meanings. They respectively represent laser point cloud frames or image frames, which are derived from continuous acquisitions during the movement of intelligent devices. And the edges vividly represent the constraint relationships between frames. This kind of constraint is closely related to the registration error, that is, the smaller the registration error, the greater the edge weight between adjacent frames, meaning the relative position relationship between these two frames is more reliable; conversely, when the registration error is large, the edge weight is smaller. For example, in the scenario where an autonomous driving vehicle continuously travels and collects data, the laser point cloud frames and the corresponding panoramic image frames collected at adjacent moments form tightly connected node pairs. Since the vehicle moves relatively smoothly in a short time, the registration error is tiny, and the edge weight is correspondingly large; while for frames with a relatively large interval, due to the cumulative error during long-term movement, the registration error increases, and the edge weight decreases accordingly.

[0082] Among them, the well-constructed graph model is optimized and solved by using the mature least squares method. Let the node set V = {v1, v2,..., v n} in the graph, the edge set E = {e ij | i, j = 1, 2,..., n}, the weight w ij of edge e ij , for the pose T i corresponding to node v i , the optimization objective function is expressed as:

[0083]

[0084] A full-scale calibration is performed on the entire data sequence, which can globally adjust the poses of each frame, so that from a macroscopic perspective, the registration of the entire data sequence reaches an almost optimal state, significantly reducing the cumulative error and providing guarantee for high-precision map construction or accurate environmental perception.

[0085] The registration method based on feature matching further includes constructing a graph model through a global optimization algorithm. Nodes represent laser point cloud frames or image frames, edges represent the constraint relationships between frames, the weights of edges are determined according to the registration error, and the graph model is optimized by using the least squares method. Let the node set V = {v1, v2,..., v n} in the graph, the edge set E = {e ij | i, j = 1, 2,..., n}, the weight w ij of edge e ij , for the pose T i corresponding to node v i, the optimization objective function is expressed as:

[0086]

[0087] In step 103, based on the graph model, a multi-modal deep learning model is used for data fusion to obtain fused data.

[0088] Preferably, the multi-modal deep learning model adopts a convolutional neural network model, and the network structure includes: a convolutional layer, a pooling layer, and a fully connected layer. Let the input panoramic image be I, and the feature map after passing through the convolutional layer is expressed as F l (I), and the convolution operation formula is:

[0089]

[0090] Among them, ω l,k is the k-th convolutional kernel of the l-th layer, * represents the convolution operation, b l is the bias term, and σ is the activation function.

[0091] Preferably, the multi-modal deep learning model adopts a deep generative adversarial network model, including: a generator G and a discriminator D. The generator G converts the laser point cloud data into a representation G(P) similar to image features, and the discriminator D is used to distinguish real image features from the features generated by the generator. The loss function of the generator is expressed as: The loss function of the discriminator is: Among them, P data represents the distribution of the laser point cloud data, and P image represents the distribution of the image data.

[0092] In the present invention, a convolutional neural network model CNN is used for data fusion. As one of the core components of the fusion, the network architecture of the CNN model is clear, covering a convolutional layer, a pooling layer, and a fully connected layer. Let the input panoramic image be I. When the image data flows into the convolutional layer, after the fine scanning and feature extraction by the convolutional kernel, the obtained feature map is expressed as F l (I), and the convolution operation formula is:

[0093]

[0094] Among them, ω l,k is the k-th convolutional kernel of the l-th layer, * represents the convolution operation, b l is the bias term, and σ is the activation function.

[0095] As for the activation function, it plays a vital role in the operation of the CNN model. This method uses the ReLU function σ(x) = max(0,x) as the activation function, which has significant advantages. On the one hand, the ReLU function can effectively accelerate the training speed of the model. Compared with the traditional saturated activation function, it avoids the gradient vanishing problem, so that the gradient can be smoothly transmitted during the back propagation process of the model, and the parameter update can be carried out efficiently; on the other hand, it has a certain sparsity induction ability, which can enable the model to automatically learn and focus on the key feature areas in the image, reduce unnecessary computing resource consumption, and improve the overall model performance. In the digital protection of cultural relics, high-resolution panoramic images of cultural relics are input. The CNN model uses the progressive convolution layer and the small-size convolution kernel to accurately capture key features such as the fine carved texture on the surface of the cultural relics. After multi-layer processing, the image features are converted into a compact vector form, building a solid bridge for the subsequent deep fusion with laser point cloud data.

[0096] In this paper, the deep generative adversarial network model GAN is introduced to further deepen the multimodal fusion effect. It is composed of two core components, the generator and the discriminator. The generator accurately distinguishes the real image features from the features generated by the generator, and the two continuously improve their respective capabilities during the adversarial training process. The loss function of the generator is designed to be Loss function of the discriminator Among them, P data represents the distribution of laser point cloud data, P image Represents the distribution of image data. Driven by this loss function, the generator continuously attempts to generate more realistic image features, while the discriminator continuously improves its identification capabilities, and the two repeatedly compete with each other. In the map construction phase of autonomous driving scenarios, laser point cloud data accurately reflects the geometric structure of the road. The generator cleverly converts it into image-like texture features, and the discriminator uses its keen discrimination to judge its authenticity. After multiple rounds of adversarial training, the features generated by the generator are increasingly close to real image features, successfully achieving deep fusion of laser point clouds and panoramic images at the feature level, providing rich and accurate environmental information for autonomous driving decision-making.

[0097] Compared with the prior art, the advantages of the present invention are:

[0098] 1. By organically combining traditional feature matching-based methods with cutting-edge multimodal deep learning models, and coordinating adaptive filtering and noise reduction processing, the registration accuracy and fusion quality of panoramic images and laser point cloud data are comprehensively improved.

[0099] 2. The method of the present invention can flexibly adjust parameters and adapt algorithms according to the requirements of different scenarios, make full use of laser SLAM technology and image data fusion, improve the registration accuracy of panoramic images and laser point clouds, and reduce errors.

[0100] 3. Through adaptive filtering and noise reduction processing methods, not only can the noise and redundancy in the original data be effectively eliminated, but also the storage, transmission and computing costs can be reduced, further improving the data quality, making subsequent algorithm processing more efficient and the overall system operation smoother.

[0101] Figure 2 Schematic diagram of the structure of a panoramic image and laser point cloud data fusion system 200 using laser SLAM technology according to an embodiment of the present invention.

[0102] like Figure 2 As shown, the panoramic image and laser point cloud data fusion system based on laser SLAM provided by the embodiment of the present invention includes: a feature matching unit 201, a graph model construction unit 202 and a data fusion unit 203.

[0103] Preferably, the feature matching unit 201 is configured to perform feature registration based on the acquired laser point cloud data and the corresponding panoramic image data to determine feature matching point pairs.

[0104] Preferably, the system further comprises:

[0105] The preprocessing unit is used to filter the laser point cloud data using a voxel filtering algorithm before aligning the acquired laser point cloud data with the corresponding panoramic image data; and to denoise and enhance the panoramic image data using a bilateral filtering algorithm.

[0106] Preferably, the feature matching unit 201 further includes:

[0107] When performing feature registration, a scale-invariant feature transformation algorithm is used to extract feature points in the panoramic image data.

[0108] Preferably, the system further comprises:

[0109] The screening unit is used to screen the feature matching point pairs using a random sampling consistency algorithm to eliminate erroneous feature matching point pairs.

[0110] Preferably, the graph model construction unit 202 is configured to construct a graph model based on the feature matching point pairs using a global optimization algorithm.

[0111] Preferably, the graph model building unit 202 builds a graph model based on the feature matching point pairs using a global optimization algorithm, including:

[0112] Nodes are used to represent laser point cloud frames or image frames, edges are used to represent the constraint relationship between frames, the weight of the edge is determined according to the registration error, and the least squares method is used to optimize the graph model;

[0113] Among them, let the node set V in the graph be V = {v1, v2, …, v n}, the edge set E = {e ij | i, j = 1, 2, …, n}, and the weight w ij of the edge e ij . For the pose T i corresponding to the node v i , the optimization objective function is expressed as:

[0114]

[0115] Among them, n is the number of nodes; T j is the pose corresponding to the node v j .

[0116] Preferably, the data fusion unit 203 is configured to perform data fusion based on the graph model by using a multimodal deep learning model to obtain fused data.

[0117] Preferably, in the data fusion unit 203, the multimodal deep learning model adopts a convolutional neural network model, and the network structure includes: a convolutional layer, a pooling layer, and a fully connected layer. Let the input panoramic image be I, and the feature map after passing through the convolutional layer be F l (I), and the convolution operation formula is:

[0118]

[0119] Among them, ω l,k is the kth convolutional kernel of the lth layer, * represents the convolution operation, b l is the bias term, and σ is the activation function.

[0120] Preferably, in the data fusion unit 203, the multimodal deep learning model adopts a deep generative adversarial network model, including: a generator G and a discriminator D. The generator G converts the laser point cloud data into a representation G(P) similar to image features, and the discriminator D is used to distinguish real image features from the features generated by the generator. The loss function of the generator is expressed as: The loss function of the discriminator is: Among them, P data represents the distribution of the laser point cloud data, and P image represents the distribution of the image data.

[0121] The panoramic image and laser point cloud data fusion system 200 of the laser SLAM technology in the embodiment of the present invention corresponds to the panoramic image and laser point cloud data fusion method 100 of the laser SLAM technology in another embodiment of the present invention, which will not be elaborated here.

[0122] Based on another aspect of the present invention, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements any one of the panoramic image and laser point cloud data fusion methods of a laser SLAM technology.

[0123] Based on another aspect of the present invention, the present invention provides an electronic device, including:

[0124] The above-mentioned computer-readable storage medium; and

[0125] One or more processors for executing the program in the computer-readable storage medium.

[0126] The present invention has been described by referring to a few embodiments. However, as is well known to those skilled in the art, other embodiments equivalent to those disclosed above of the present invention equally fall within the scope of the present invention.

[0127] Generally, all terms used in the present invention are interpreted according to their ordinary meanings in the technical field, unless otherwise explicitly defined therein. All references to "a / the [device, component, etc.]" are open to interpretation as at least one instance of the device, component, etc., unless otherwise explicitly stated. The steps of any method disclosed herein need not be run in the exact order disclosed, unless explicitly stated.

[0128] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0129] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for realizing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0130] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0131] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0132] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the field should understand that the specific implementation methods of the present invention can still be modified or replaced by equivalents. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered by the scope of protection of the present invention.

Claims

1. A method for fusing panoramic images and laser point cloud data based on laser SLAM, characterized in that, The method comprises: Perform feature registration based on the acquired laser point cloud data and the corresponding panoramic image data to determine feature matching point pairs; Based on the feature matching point pairs, a graph model is constructed using a global optimization algorithm; Based on the graph model, a multimodal deep learning model is used to perform data fusion to obtain fused data.

2. The method according to claim 1, characterized in that Before registering the acquired laser point cloud data and the corresponding panoramic image data, the method further includes: Performing filtering processing on the laser point cloud data using a voxel filtering algorithm; The panoramic image data is subjected to denoising and enhancement processing using a bilateral filtering algorithm.

3. The method according to claim 1, characterized in that, The method further comprises: When performing feature registration, a scale-invariant feature transformation algorithm is used to extract feature points in the panoramic image data.

4. The method according to claim 1, wherein The method further comprises: The feature matching point pairs are screened using a random sampling consistency algorithm to eliminate erroneous feature matching point pairs.

5. The method according to claim 1, wherein The step of constructing a graph model based on the feature matching point pairs using a global optimization algorithm includes: Nodes represent laser point cloud frames or image frames, edges represent the constraints between frames, edge weights are determined based on the registration error, and the least squares method is used to optimize the graph model. Among them, let the node set V in the graph be V = {v1, v2, …, v n}, the edge set E = {e ij | i, j = 1, 2, …, n}, the weight w ij of the edge e ij , for the pose T i corresponding to the node v i , the optimization objective function is expressed as: F(T) = ∑ eij∈E w ij ∥T i -T j -1 ∥ 2 , Where n is the number of nodes; T j For node v j The corresponding pose.

6. The method according to claim 1, characterized in that The multimodal deep learning model adopts a convolutional neural network model. The network structure includes: convolution layer, pooling layer and fully connected layer. Let the input panoramic image be I, and the feature map after the convolution layer is represented as F l (I), the convolution operation formula is: where ω l,k is the k-th convolution kernel of the l-th layer, * represents the convolution operation, b l is the bias term, and σ is the activation function.

7. The method according to claim 6, wherein The multimodal deep learning model adopts a deep generative adversarial network model, including a generator G and a discriminator D. The generator G converts the laser point cloud data into a representation G(P) similar to image features. The discriminator D is used to distinguish between real image features and features generated by the generator. The loss function of the generator is expressed as: The loss function of the discriminator is: Among them, P data represents the distribution of laser point cloud data, P image Represents the distribution of image data.

8. A panoramic image and laser point cloud data fusion system based on laser SLAM, characterized in that, The system comprises: A feature matching unit, configured to perform feature registration based on the acquired laser point cloud data and the corresponding panoramic image data, and determine feature matching point pairs; A graph model construction unit, configured to construct a graph model based on the feature matching point pairs using a global optimization algorithm; The data fusion unit is used to perform data fusion based on the graph model using a multimodal deep learning model to obtain fused data.

9. The system according to claim 8, characterized in that, The system further comprises: The preprocessing unit is used to filter the laser point cloud data using a voxel filtering algorithm before aligning the acquired laser point cloud data with the corresponding panoramic image data; and to denoise and enhance the panoramic image data using a bilateral filtering algorithm.

10. The system according to claim 8, characterized in that, The feature matching unit further includes: When performing feature registration, a scale-invariant feature transformation algorithm is used to extract feature points in the panoramic image data.

11. The system according to claim 8, wherein The system further comprises: The screening unit is used to screen the feature matching point pairs using a random sampling consistency algorithm to eliminate erroneous feature matching point pairs.

12. The system according to claim 8, wherein The graph model construction unit constructs a graph model based on the feature matching point pairs using a global optimization algorithm, including: Nodes are used to represent laser point cloud frames or image frames, edges are used to represent the constraint relationship between frames, the weight of the edges is determined according to the registration error, and the least squares method is used to optimize the graph model; In which, let the node set V in the graph be {v1,v2,…,v n }, edge set E={e ij |i,j=1,2,…,n}, edge e ij The weight w ij , for node v i The corresponding pose T i , the optimization objective function is expressed as: F(T)=∑ eij∈E w ij ∥T i -T j -1 ∥ 2 , where n is the number of nodes; T j is the pose corresponding to node v j ​ 13. The system according to claim 8, wherein: In the data fusion unit, the multi-modal deep learning model adopts a convolutional neural network model, and the network structure includes: a convolutional layer, a pooling layer, and a fully connected layer. Let the input panoramic image be I, and the feature map after passing through the convolutional layer is represented as F l (I), and the convolution operation formula is: Among them, ω l,k is the kth convolution kernel of the lth layer, * represents the convolution operation, b l is the bias term, and σ is the activation function.

14. The system according to claim 13, characterized in that In the data fusion unit, the multimodal deep learning model adopts a deep generative adversarial network model, including: a generator G and a discriminator D. The generator G converts the laser point cloud data into a representation G(P) similar to image features. The discriminator D is used to distinguish between real image features and features generated by the generator. The loss function of the generator is expressed as: The loss function of the discriminator is: Among them, P data represents the distribution of laser point cloud data, P image Represents the distribution of image data.

15. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

16. An electronic device, characterized in that, include: The computer-readable storage medium of claim 15; as well as One or more processors are configured to execute the program in the computer-readable storage medium.