A method, apparatus, device, and storage medium for fusing scenario data

By acquiring and processing original image data, point cloud and millimeter radar data in preset scenarios, combining homography matrix mapping and weighted averaging method for fusion, and optimizing weight values ​​using the global optimization algorithm, the three-party direct fusion of image, point cloud and radar is achieved, solving the problem of low error accumulation and accuracy in the existing technology, and improving the fusion accuracy and effect.

CN119671869BActive Publication Date: 2025-06-24POWER CHINA KUNMING ENG CORP LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510197097.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-06-24
Estimated Expiration
2045-02-21

AI Technical Summary

Technical Problem

There are many fusion methods in the prior art, which leads to different errors generated by different fusion methods, and requires registration of multiple stages, especially when fusion of three parties, there are many processing steps, which leads to continuous accumulation and amplification of the errors in each stage, resulting in a low accuracy of the final fusion effect.

Method used

By obtaining the original image data of the preset scene, using the target detection algorithm to mark and pinpoint fixed points, combining point cloud and millimeter radar data, iterating the weight value through homography algorithm to improve clarity, and realizing the three-party direct fusion of image, point cloud and radar.

Benefits of technology

The error introduction of intermediate processes is reduced, the fusion accuracy is improved, the error caused by incorrect threshold setting is avoided, the optimal fusion effect is achieved, and the image distortion and pixel loss caused by the direct superposition of the pair-to-two fusion technology is avoided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119671869B_ABST
    Figure CN119671869B_ABST
Patent Text Reader

Abstract

The present application discloses a method, device, equipment and storage medium for fusing scenario data, which relates to the field of remote sensing data processing. The method realizes the direct fusion of images, point clouds and radars. Compared with other pairwise fusion technologies, the direct superposition has fewer calculation steps, reduces the error introduction in the intermediate process, improves the fusion accuracy, and makes the final fusion effect better. Moreover, the algorithms are all adaptive algorithms, without setting an absolute static threshold, avoiding large errors in the final fusion result due to incorrect threshold setting. At the same time, the present application also globally optimizes and iterates the weight values of each pixel point in the final fusion to make the final fused image maintain the optimal fusion effect, avoiding problems such as image distortion, pixel point loss, and tearing caused by the direct superposition of pairwise fusion technologies.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of electronic digital data processing, and particularly relates to a method, device, equipment and storage medium for fusing scenario data. Background Art

[0002] With the development of the electrification and intelligence of construction machinery, the ways of obtaining scenario data are increasing day by day. The mainstream acquisition means such as cameras, radars, and 3D mapping have also become important means for perceiving the surrounding environment and have become the recognized mainstream choices in the industry.

[0003] Among them, cameras have low cost and low power consumption, and can perform high-level target detection and image processing functions, but are greatly affected by factors such as light, weather, and image quality; radars have strong environmental adaptability, can work normally under bad weather and light conditions, and at the same time have stable performance and low power consumption, but have relatively low spatial resolution; 3D mapping has the advantages of high flexibility, variable data density, rich additional information, and retention of original data and is widely used, but at the same time brings the disadvantages of high data processing complexity, large data volume, and non-structured characteristics, making the acquisition of scenario data for the same scenario have their own advantages and disadvantages.

[0004] Fusing the scenario data obtained by the above means is an effective means to improve the data accuracy. However, the existing fusing means are usually camera-radar fusion and point cloud-radar fusion, and there are many fusing means for each. The errors generated by different fusing means are different and multiple stages of registration are required. If three-party fusion is performed, there are more processing steps, resulting in the continuous accumulation and amplification of errors in each stage, and the accuracy of the final fusion effect is relatively low. Summary of the Invention

[0005] The main purpose of the present application is to provide a method, device, equipment and storage medium for fusing scenario data, so as to solve the problem in the prior art that there are many fusing means for each, the errors generated by different fusing means are different and multiple stages of registration are required. If three-party fusion is performed, there are more processing steps, resulting in the continuous accumulation and amplification of errors in each stage, and the accuracy of the final fusion effect is relatively low.

[0006] To achieve the above purpose, the present application provides the following technical solutions:

[0007] A method for fusing scenario data, the scenario data is from the same preset scenario, and there are a preset number of calibration points in the preset scenario. The fusing method includes:

[0008] Step S1, obtaining the original image data of all calibration points in the preset scenario;

[0009] Step S2, mark all calibration points in the original image data with detection frames through a target detection algorithm to form first-order image data;

[0010] Step S3, obtain a point cloud data set of the preset scene with all calibration points through a first preset mapping strategy;

[0011] Step S4, project all points in the point cloud data set onto the first-order image data to form second-order image data;

[0012] Step S5, obtain millimeter radar data of the preset scene with all calibration points through a second preset mapping strategy;

[0013] Step S6, map the millimeter radar data to the second-order image data through a homography matrix to form third-order image data;

[0014] Step S7, fuse each pixel point of the third-order image data through the weighted average method respectively, and obtain the fused image data after traversing all pixel points;

[0015] Step S8, iterate the weight value of the weighted average method through a global optimization algorithm to make the clarity evaluation index of the fused image data reach the maximum value;

[0016] Step S9, obtain the fused image data corresponding to the maximum value and define it as the final image data.

[0017] As a further improvement of the present application, Step S2, mark all calibration points in the original image data with detection frames through a target detection algorithm to form first-order image data, including:

[0018] Step S21, evenly divide the original image data into several square grids;

[0019] Step S22, predict several bounding boxes for all calibration points based on all square grids according to the target detection algorithm;

[0020] Step S23, obtain the confidence of each bounding box respectively, and obtain the bounding box with the highest confidence and mark it as the first-order bounding box;

[0021] Step S24, calculate the intersection over union of each other bounding box and the first-order bounding box respectively;

[0022] Step S25, select the bounding boxes with an intersection over union greater than or equal to a preset threshold as second-order bounding boxes;

[0023] Step S26, obtain the second-order bounding box with the highest confidence and define it as the detection frame of all calibration points;

[0024] Step S27: Output all the detection frames to the original image data to form the first-order image data.

[0025] As a further improvement of this application, in step S4, project all the points in the point cloud dataset onto the first-order image data to form the second-order image data, including:

[0026] Step S41: Obtain the coordinate value of the th point in the point cloud dataset ;

[0027] Step S42: Define the plane equation of the first-order image data as ;

[0028] Step S43: Solve the projection coordinate value of the th point according to formula (1) :

[0029] (1);

[0030] Step S44: Solve the projection coordinate values of all the points in the point cloud dataset according to formula (1);

[0031] Step S45: Output all the projection coordinate values to the first-order image data to form the second-order image data.

[0032] As a further improvement of this application, in step S6, map the millimeter radar data to the second-order image data through the homography matrix to form the third-order image data, including:

[0033] Step S61: Define the mapping relationship between the millimeter radar data and the first-order image data through formula (2):

[0034] (2);

[0035] Where is the radar position coordinate of the current calibration point, is the pixel coordinate value of the current calibration point based on the original image data, is the homography matrix;

[0036] Step S62: Solve the homography matrix according to formula (2) to obtain the mapping relationship;

[0037] Step S63: Substitute all the radar position coordinates in the millimeter radar data into formula (2) to solve for all the radar mapping coordinates;

[0038] Step S64: Output all the radar mapping coordinates to the second-order image data to form the third-order image data.

[0039] As a further improvement of the present application, in step S7, each pixel point of the third-order image data is fused by the weighted average method, and after traversing all pixel points, the fused image data is obtained, including:

[0040] Step S71, each pixel point of the third-order image data is fused by the weighted average method according to formula (3):

[0041] (3);

[0042] Wherein, is the fused image data, is the pixel coordinate of each order of image data, is the weight value based on the third-order image data, is the weight value based on the second-order image data, is the weight value based on the first-order image data, is the original image data;

[0043] Step S72, all fused pixel points are obtained after traversing all pixel points of the third-order image data according to formula (3);

[0044] Step S73, all fused pixel points are integrated to obtain the fused image data.

[0045] As a further improvement of the present application, in step S8, the weight value of the weighted average method is iterated by the global optimization algorithm to make the clarity evaluation index of the fused image data reach the maximum value, including:

[0046] Step S81, a number of random solutions are defined for the weight value of the weighted average method according to formula (4), and the optimization result of all random solutions is defined as the maximum value of the clarity evaluation index;

[0047] (4);

[0048] Wherein, is the set of all random solutions, are each random solution respectively, is the label of the random solution, is the number of all random solutions; is the set of the speeds of all random solutions, are the speeds of each random solution respectively;

[0049] Step S82: Initialize the positions of each random solution, and update the current position and current velocity respectively based on the same random solution according to Equation (5):

[0050] (5);

[0051] Wherein, is the velocity of the th random solution at the th step, is the velocity inertia of the th random solution at the th step, is the inertia coefficient, is the self - awareness representation of the th random solution, is the social - awareness representation of the th random solution; and are both learning factors, is a random number of , is the individual optimal solution obtained by the th random solution, is the global optimal solution obtained by the th random solution, is the th random solution at the th step, is the th random solution at the th step;

[0052] Step S83: Iterate each random solution according to the said Equation (5) to update each and each ;

[0053] Step S84: Respectively judge whether the difference of each compared with the previous iteration is less than or equal to the first preset fitness threshold. If the differences of each compared with the previous iteration are all less than or equal to the first preset fitness threshold, then execute Step S85;

[0054] Step S85: Respectively judge whether the difference of each compared with the previous iteration is less than or equal to the second preset fitness threshold. If the differences of each compared with the previous iteration are all less than or equal to the second preset fitness threshold, then execute Step S86;

[0055] Step S86: Determine that the optimal solution of the said weight value has been obtained.

[0056] As a further improvement of the present application, in step S83, each random solution is iterated according to the formula (5) to update each and each , including:

[0057] In step S831, the inertia coefficient is linearly decreased once according to the formula (6) based on each iteration:

[0058] (6);

[0059] wherein, is the inertia coefficient after optimization of the th random solution at the th step of optimization, is the initial inertia coefficient, is the current iteration step number, is the maximum iteration step number.

[0060] To achieve the above object, the present application also provides the following technical solutions:

[0061] A fusion device for scene data, the fusion device for scene data is applied to the fusion method as described above, and the fusion device includes:

[0062] An original image data acquisition module, configured to acquire original image data of all calibration points in the preset scene;

[0063] A first-order image data acquisition module, configured to mark all calibration points in the original image data through a detection frame by using an object detection algorithm to form first-order image data;

[0064] A calibration point cloud data set acquisition module, configured to acquire a point cloud data set of all calibration points in the preset scene through a first preset surveying and mapping strategy;

[0065] A second-order image data acquisition module, configured to project all points in the point cloud data set onto the first-order image data to form second-order image data;

[0066] A millimeter radar data acquisition module, configured to acquire millimeter radar data of all calibration points in the preset scene through a second preset surveying and mapping strategy;

[0067] A third-order image data acquisition module, configured to map the millimeter radar data to the second-order image data through a homography matrix to form third-order image data;

[0068] A fused image data acquisition module, configured to perform fusion on each pixel point of the third-order image data through a weighted average method, and obtain fused image data after traversing all pixel points;

[0069] The weighted average method weight value acquisition module is used to iterate the weight value of the weighted average method through a global optimization algorithm so that the clarity evaluation index of the fused image data reaches the maximum value;

[0070] The final image data acquisition module is used to acquire the fused image data corresponding to the maximum value and define it as the final image data.

[0071] To achieve the above object, the present application also provides the following technical solutions:

[0072] An electronic device includes a processor and a memory coupled to the processor. The memory stores program instructions executable by the processor. When the processor executes the program instructions stored in the memory, the above-described method for fusing scene data is implemented.

[0073] To achieve the above object, the present application also provides the following technical solutions:

[0074] A storage medium stores program instructions. When the program instructions are executed by a processor, the method for fusing scene data as described above can be implemented.

[0075] The present application obtains the original image data of all calibration points in a preset scene; marks all calibration points in the original image data through a target detection algorithm to form first-order image data; obtains a point cloud data set of all calibration points in the preset scene through a first preset mapping strategy; projects all points in the point cloud data set onto the first-order image data to form second-order image data; obtains millimeter radar data of all calibration points in the preset scene through a second preset mapping strategy; maps the millimeter radar data to the second-order image data through a homography matrix to form third-order image data; fuses each pixel point of the third-order image data through the weighted average method respectively, and obtains the fused image data after traversing all pixel points; iterates the weight value of the weighted average method through a global optimization algorithm so that the clarity evaluation index of the fused image data reaches the maximum value; obtains the fused image data corresponding to the maximum value and defines it as the final image data. The present application realizes the direct fusion of images, point clouds, and radars. Compared with other pairwise fusion technologies, the direct superposition has fewer calculation steps, reduces the error introduction in the intermediate process, improves the fusion accuracy, makes the final fusion effect better, and the algorithms are all adaptive algorithms, without setting an absolute static threshold, avoiding large errors in the final fusion result due to incorrect threshold setting. At the same time, the present application also iterates the weight values of the last fused pixel points through global optimization to keep the final fused image in the optimal fusion effect, avoiding problems such as image distortion, pixel point loss, and tearing caused by the direct superposition of pairwise fusion technologies. BRIEF DESCRIPTION OF THE DRAWINGS

[0076] Figure 1 Schematic diagram of the process steps of an embodiment of the method for fusing scenario data of the present application;

[0077] Figure 2 Schematic diagram of the functional modules of an embodiment of the device for fusing scenario data of the present application;

[0078] Figure 3 Schematic diagram of the structure of an embodiment of the electronic device of the present application;

[0079] Figure 4 Schematic diagram of the structure of an embodiment of the storage medium of the present application. Detailed implementation manners

[0080] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0081] The terms "first", "second", and "third" in the present application are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first", "second", and "third" may explicitly or implicitly include at least one of such features. In the description of the present application, the meaning of "a plurality" is at least two, such as two, three, etc., unless otherwise specifically and clearly defined. All directional indications (such as up, down, left, right, front, back...) in the embodiments of the present application are only used to explain the relative positional relationship and movement conditions between components in a specific posture (as shown in the drawings). If the specific posture changes, the directional indications will also change accordingly. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally further include steps or units not listed, or may optionally further include other steps or units inherent to these processes, methods, products, or devices.

[0082] Referring to "embodiment" herein means that a specific feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the present application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0083] As Figure 1 shown, this embodiment provides an embodiment of the method for fusing scene data. In this embodiment, the scene data is derived from the same preset scene, such as natural landforms, building areas, municipal related facilities, etc. There are a preset number of calibration points in the preset scene, and the number of calibration points is generally two to four.

[0084] Specifically, the fusion method includes the following steps:

[0085] Step S1, obtain the original image data of all calibration points in the preset scene.

[0086] Preferably, the original image data can be directly obtained by shooting with a camera. In order to ensure the subsequent fusion accuracy, several original image data can be shot at the same shooting position.

[0087] Step S2, mark all calibration points in the original image data with detection frames through the target detection algorithm to form first-order image data.

[0088] Preferably, the target detection algorithm can be implemented by VJ, HOG, DPMDetector; deep learning Two-stage RCNN, SPPNet, FastRCNN, FasterRCNN; Trick algorithm FPN, CascadeRCNN; deep learning one-stage Yolo, X, SSD, RetinaNet; deep learning Anchor-free CornerNet, CenterNet, FCOS; Transformer DETR, etc.

[0089] Preferably, this embodiment preferably uses the Yolo algorithm.

[0090] Step S3, obtain the point cloud data set of all calibration points in the preset scene through the first preset mapping strategy.

[0091] Preferably, the first preset mapping strategy for obtaining point cloud data is one of the laser ranging method, the 3D scanner method, and the structured light method.

[0092] Among them, the laser ranging method emits laser to the surface of an object through a laser emitter. After the laser irradiates the surface of the object, a photoelectric converter is used to convert the reflected laser signal into an electrical signal, and the three-dimensional coordinates of the object surface are calculated through an algorithm to obtain point cloud data; the three-dimensional scanner method is a device that scans the surface of an object by laser or optical means to obtain point cloud data. Its working principle is to control the scanning angle and distance of the scanner to convert the information on the object surface into point cloud data; the structured light method is a method of three-dimensional scanning by projecting a grating pattern. It usually uses a projector to project a specific grating pattern, and then a camera captures the reflected image to obtain the three-dimensional coordinates of the object surface and get point cloud data.

[0093] Preferably, the processing flow of point cloud data is generally data acquisition, point cloud preprocessing, feature extraction and matching. Among them, data acquisition: Obtain the point cloud data of the object surface. Common data acquisition methods include laser scanning, structured light scanning and image acquisition. Point cloud preprocessing: Process the collected point cloud data, such as noise filtering, outlier removal, data downsampling and data registration. Feature extraction and matching: Extract features from the point cloud data and use a matching algorithm to align the point cloud.

[0094] Step S4, project all points in the point cloud dataset onto the first-order image data to form second-order image data.

[0095] Step S5, obtain millimeter radar data with all calibration points in the preset scene through the second preset mapping strategy.

[0096] Preferably, the second preset mapping strategy for obtaining millimeter radar data is usually to directly obtain it on an online platform or measure it by oneself. Online platforms such as Radar Journal, National Comprehensive Earth Observation Data Sharing Platform, Chinese Science Data, IEEE Dataport, Zenodo, Kaggle, Papers With Code, Mendeley Data, Google Dataset Search, Github; Measuring by oneself can be directly obtained through a millimeter-level radar without difficulty in acquisition.

[0097] Step S6, map the millimeter radar data to the second-order image data through a homography matrix to form third-order image data.

[0098] Step S7, fuse each pixel point of the third-order image data through the weighted average method. After traversing all pixel points, the fused image data is obtained.

[0099] Step S8, iterate the weight value of the weighted average method through a global optimization algorithm to make the clarity evaluation index of the fused image data reach the maximum value.

[0100] Preferably, the clarity evaluation metrics are mainly divided into subjective evaluation metrics and objective evaluation metrics.

[0101] Among them, the subjective evaluation metrics are evaluated by manually observing the image, and the results are greatly affected by the subjective factors of the observer and are difficult to quantify. Several common subjective evaluation metrics include:

[0102] Clarity level evaluation: The observer divides the image into different levels according to its clarity, such as very clear, clear, average, blurred, very blurred, etc.

[0103] Blur evaluation: The observer scores the image according to its blur degree, such as 0 - 10 points, and the higher the score, the clearer the image.

[0104] Detail visibility evaluation: The observer evaluates the image according to the visibility of details in the image, such as details are clearly visible, details are blurred but visible, details are not visible, etc.

[0105] Among them, the objective evaluation metrics are evaluated based on the characteristics of the image itself, are not affected by subjective factors, and can be quantitatively analyzed. Several common objective evaluation metrics include:

[0106] Gray variance: It reflects the distribution of gray values in the image. The larger the gray variance, the clearer the image.

[0107] Edge gradient: It reflects the sharpness of edges in the image. The larger the edge gradient, the clearer the image.

[0108] Information entropy: It reflects the richness of information in the image. The larger the information entropy, the clearer the image.

[0109] Spectral entropy: It reflects the richness of frequency information in the image. The larger the spectral entropy, the clearer the image.

[0110] Structural similarity: It reflects the similarity of structural information in the image. The higher the structural similarity, the clearer the image.

[0111] In practical applications, combining multiple metrics can provide a more accurate evaluation. For example, MTF (Modulation Transfer Function) and SFR (Spatial Frequency Response) are important metrics for measuring the resolution of cameras, and through these metrics, the spatial resolution levels of different cameras can be compared and verified. In addition, Brenner gradient method, Tenegrad gradient method, Laplace gradient method, variance method, energy gradient method, etc. are also common image clarity evaluation methods.

[0112] Preferably, one or a combination of multiple objective evaluation indicators can be selected in this embodiment to evaluate the clarity. If multiple combinations are selected, the average score of the multiple combinations can be used for evaluation.

[0113] Step S9: Obtain the fused image data corresponding to the maximum value and define it as the final image data.

[0114] Further, in step S2, all calibration points are marked by detection frames in the original image data through the target detection algorithm to form the first-order image data, including:

[0115] Step S21: Divide the original image data into several square grids on average.

[0116] Specifically, the size of the original picture of the image data can be adjusted to 448×448, and then the adjusted picture is evenly divided into S×S (for example, 7×7) grids, and the size of each grid is 64×64.

[0117] Step S22: Based on all the square grids, predict several bounding boxes for all calibration points according to the target detection algorithm.

[0118] Step S23: Obtain the confidence level of each bounding box respectively, and obtain the bounding box with the highest confidence level and mark it as the first-order bounding box.

[0119] Preferably, the confidence level can be understood as whether there is a target in the current grid and the accuracy of the detection frame.

[0120] Step S24: Calculate the intersection over union of each of the other bounding boxes and the first-order bounding box respectively.

[0121] Step S25: Select the bounding boxes with the intersection over union greater than or equal to the preset threshold as the second-order bounding boxes.

[0122] Preferably, the preset threshold can be set to 80%.

[0123] Step S26: Obtain the second-order bounding box with the highest confidence level and define it as the detection frame of all calibration points.

[0124] Step S27: Output all the detection frames to the original image data to form the first-order image data.

[0125] Preferably, each grid is used to predict the abscissa, ordinate, width, and height of a detection frame, as well as the confidence level of each detection frame, that is, each grid needs to predict values.

[0126] Each grid needs to predict ; among them, is the offset of the center of the detection box relative to the grid, is the ratio of the detection box relative to the above-mentioned resized image, is the confidence of the grid, with a value of 1 or 0.

[0127] Preferably, the confidence can be understood as whether there is a target in the current grid and the accuracy of the detection box.

[0128] For example: Suppose there is a target in a resized image, and the width and height of the resized image are Then:

[0129] The image is evenly divided into 7×7 (S×S) grids. If there is a grid located at the center of the target, the coordinates of this grid are , suppose the coordinates of the center of the target are , then the above-mentioned offset can be calculated according to the following formula : .

[0130] Preferably, in actual detection, if the predicted detection box and the actual bounding box perfectly overlap, the value of the intersection over union is 1. In actual application, the value of the first preset threshold is generally set to 0.5 to determine whether the predicted bounding box is correct, and the accuracy of the bounding box is positively correlated with the intersection over union.

[0131] Preferably, the YOLO algorithm also needs to train the detection box to improve the accuracy of object detection.

[0132] Next, train the above training model through a preset pedestrian and vehicle training set, and iteratively adjust the weights and biases of the training model a first preset number of times through the backpropagation algorithm to reduce the value of the loss function of the training model.

[0133] Preferably, the loss function is shown as the following formula:

[0134] .

[0135] Among them, is the indicator function of whether the th grid's th detection box is responsible for the target, with a value of 1 or 0; , , , , respectively correspond to the th predicted value.

[0136] It can be understood that the loss function includes the deviation of the coordinate values of the detection boxes, the deviation of the confidence, and the deviation of the predicted probability (or class deviation).

[0137] Among them, is the midpoint loss of the detection box in the coordinate value deviation, is the width and height loss of the detection box in the coordinate value deviation, is the deviation of the confidence, is the deviation of the predicted probability (or class deviation).

[0138] Among them, is the positioning error penalty. Generally, ; is the above-mentioned S×S grids; is the number of bounding boxes; and is the estimated values of the abscissa and ordinate of the midpoint of the th bounding box; and are the estimated values of the width and height of the th bounding box; is the confidence of the th bounding box; is the estimated value of the confidence of the th bounding box; is the confidence prediction loss. Generally, is the th class probability of the bounding box; is the th estimated value of the class probability of the bounding box; and in correspond to .

[0139] It should be noted that since each grid does not necessarily contain a target, if there is no target in the grid, it will cause to take a value of 0, resulting in too large a gradient span in the subsequent backpropagation algorithm. Therefore, is introduced to control the loss of the predicted position of the detection box, and is introduced to control the loss of no target in a single grid.

[0140] It should be noted that the symbolic meanings explained in the principle of the above object detection algorithm do not communicate with the meanings of other symbols in the context.

[0141] Furthermore, in step S4, all points in the point cloud dataset are projected onto the first-order image data to form second-order image data, including:

[0142] Step S41, obtaining the first The coordinates of the points .

[0143] Step S42, define the plane equation of the first-order image data as .

[0144] Preferably, if the first-order image data is coplanar with the xy plane, then Can be optimized to , in order to reduce the computing burden.

[0145] Step S43, solve the first The projected coordinates of the points :

[0146] (1).

[0147] Step S44, solving the projection coordinate values ​​of all points in the point cloud data set according to equation (1).

[0148] In step S45, all projection coordinate values ​​are output to the first-order image data to form the second-order image data.

[0149] Furthermore, step S6, mapping the millimeter radar data to the second-order image data through the homography matrix to form the third-order image data, includes:

[0150] Step S61, defining the mapping relationship between the millimeter radar data and the first-order image data by equation (2):

[0151] (2).

[0152] in, is the radar position coordinate of the current calibration point, is the pixel coordinate value of the current calibration point based on the original image data, is the homography matrix.

[0153] Step S62, solving the homography matrix according to equation (2): Get the mapping relationship.

[0154] Step S63: Substitute all radar position coordinates in the millimeter radar data into equation (2) to obtain all radar mapping coordinates.

[0155] In step S64, all radar mapping coordinates are output to second-order image data to form third-order image data.

[0156] For example: If the coordinates of the four positions of the corner reflector detected by the millimeter wave radar They are respectively (12, 4.9, 0), (16.2, -0.9, 0), (16.4, 3.1, 0), (18.2, -3.7, 0), and the corresponding pixel coordinates are (373.68, 575.71, 0), (453.22, 544.49, 0), (264.38, 546.6, 0), (177, 547.74, 0).

[0157] Then, the homography matrix is calculated through Equation (2). .

[0158] Furthermore, in step S7, each pixel point of the third-order image data is fused by the weighted average method. After traversing all pixel points, the fused image data is obtained, including:

[0159] Step S71, each pixel point of the third-order image data is fused by the weighted average method according to Equation (3):

[0160] (3).

[0161] Wherein, is the fused image data, are the pixel coordinates of each-order image data, is the weight value based on the third-order image data, is the weight value based on the second-order image data, is the weight value based on the first-order image data, is the original image data.

[0162] Step S72, all fused pixel points are obtained after traversing all pixel points of the third-order image data according to Equation (3).

[0163] Step S73, all fused pixel points are integrated to obtain the fused image data.

[0164] Furthermore, in step S8, the weight value of the weighted average method is iterated through the global optimization algorithm to make the clarity evaluation index of the fused image data reach the maximum value, including:

[0165] Step S81, a number of random solutions are defined for the weight value of the weighted average method according to Equation (4), and the optimization result of all random solutions is defined as the maximum value of the clarity evaluation index.

[0166] (4).

[0167] Wherein, is the set of all random solutions, For each random solution, is the label of the random solution, is the number of all random solutions; is the set of velocities of all random solutions, are the velocities of each random solution respectively.

[0168] Step S82, initialize the position of each random solution, and update the current position and current velocity respectively based on the same random solution according to Equation (5):

[0169] (5).

[0170] Among them, is the velocity of the th random solution at the th step, is the velocity inertia of the th random solution at the th step, is the inertia coefficient, is the self - cognitive representation of the th random solution, is the social - cognitive representation of the th random solution; and are both learning factors, is a random number, is the individual optimal solution obtained by the th random solution, is the global optimal solution obtained by the th random solution, is the th step of the th random solution, is the th step of the th random solution.

[0171] Preferably, the value range of is ; the value range of is .

[0172] Step S83, iterate each random solution according to Equation (5) to update each and each .

[0173] Step S84, judge each Whether the difference compared to the previous iteration is less than or equal to the first preset adaptation threshold. If each The difference compared to the previous iteration is less than or equal to the first preset adaptation threshold for all, then step S85 is executed.

[0174] In step S85, it is determined for each Whether the difference compared to the previous iteration is less than or equal to the second preset adaptation threshold. If each The difference compared to the previous iteration is less than or equal to the second preset adaptation threshold for all, then step S86 is executed.

[0175] In step S86, it is determined that the optimal solution of the weight value has been obtained.

[0176] Preferably, the values of the first preset adaptation threshold and the second preset adaptation threshold need to be adjusted according to the specific problem, generally according to the calculation results. If the adaptation threshold is set too small, it may cause the algorithm to stop prematurely and the optimal solution cannot be obtained; if the adaptation threshold is set too large, it may cause the algorithm to over-iterate and waste computing resources.

[0177] Preferably, the adaptation threshold can also be evaluated by one of the Griewank function, Rastrigin function, Schaffer function, Ackley function, and Rosenbrock function.

[0178] Furthermore, in step S83, each random solution is iterated according to formula (5) to update each and each , including:

[0179] In step S831, the inertia coefficient is linearly decreased once according to formula (6) based on each iteration:

[0180] (6).

[0181] Wherein, is the inertia coefficient after optimization of the th random solution at the th step of optimization, is the initial inertia coefficient, is the current iteration step number, is the maximum iteration step number.

[0182] Preferably, the initial inertia coefficient is generally set to 0.5, and the maximum iteration step number is generally set according to actual needs. In this embodiment, it can be set to 1000 times.

[0183] In this embodiment, the original image data of all calibration points in a preset scenario is obtained; all calibration points are marked in the original image data through a target detection algorithm by detection frames to form first-order image data; a point cloud data set of all calibration points in the preset scenario is obtained through a first preset mapping strategy; all points in the point cloud data set are projected onto the first-order image data to form second-order image data; millimeter radar data of all calibration points in the preset scenario is obtained through a second preset mapping strategy; the millimeter radar data is mapped onto the second-order image data through a homography matrix to form third-order image data; each pixel point of the third-order image data is fused through a weighted average method, and after traversing all pixel points, the fused image data is obtained; the weight value of the weighted average method is iteratively optimized through a global optimization algorithm so that the clarity evaluation index of the fused image data reaches the maximum value; the fused image data corresponding to the maximum value is obtained and defined as the final image data. This embodiment realizes the direct fusion of images, point clouds, and radars. Compared with the direct superposition of other pairwise fusion technologies, the calculation steps are fewer, the error introduction in the intermediate process is reduced, the fusion accuracy is improved, and the final fusion effect is better. Moreover, the algorithms are all adaptive algorithms and do not require setting an absolute stillness threshold, avoiding large errors in the final fusion result due to incorrect threshold setting. At the same time, this embodiment also iteratively optimizes the weight values of the last fused pixel points through global optimization to keep the final fused image in the optimal fusion effect, avoiding problems such as image distortion, pixel point loss, and tearing caused by the direct superposition of pairwise fusion technologies.

[0184] As Figure 2 shown, this embodiment provides an embodiment of a fusion device for scenario data. In this embodiment, the fusion device is applied to the fusion method in the above embodiment.

[0185] Specifically, the fusion device includes an original image data acquisition module 1, a first-order image data acquisition module 2, a calibration point cloud data set acquisition module 3, a second-order image data acquisition module 4, a millimeter radar data acquisition module 5, a third-order image data acquisition module 6, a fused image data acquisition module 7, a weighted average method weight value acquisition module 8, and a final image data acquisition module 9 that are electrically connected in sequence.

[0186] Among them, the original image data acquisition module 1 is used to acquire the original image data of a preset scene with all calibration points; the first-order image data acquisition module 2 is used to mark all calibration points in the original image data through a detection box by means of a target detection algorithm to form first-order image data; the calibration point cloud data set acquisition module 3 is used to acquire the point cloud data set of the preset scene with all calibration points through a first preset mapping strategy; the second-order image data acquisition module 4 is used to project all points in the point cloud data set onto the first-order image data to form second-order image data; the millimeter radar data acquisition module 5 is used to acquire the millimeter radar data of the preset scene with all calibration points through a second preset mapping strategy; the third-order image data acquisition module 6 is used to map the millimeter radar data to the second-order image data through a homography matrix to form third-order image data; the fused image data acquisition module 7 is used to perform fusion on each pixel point of the third-order image data through a weighted average method respectively, and obtain the fused image data after traversing all pixel points; the weighted average method weight value acquisition module 8 is used to iterate the weight value of the weighted average method through a global optimization algorithm so that the clarity evaluation index of the fused image data reaches the maximum value; the final image data acquisition module 9 is used to acquire the fused image data corresponding to the maximum value and define it as the final image data.

[0187] Further, the first-order image data acquisition module 2 specifically includes a first first-order image data acquisition sub-module, a second first-order image data acquisition sub-module, a third first-order image data acquisition sub-module, a fourth first-order image data acquisition sub-module, a fifth first-order image data acquisition sub-module, a sixth first-order image data acquisition sub-module, and a seventh first-order image data acquisition sub-module that are electrically connected in sequence; the first first-order image data acquisition sub-module is electrically connected to the original image data acquisition module 1, and the seventh first-order image data acquisition sub-module is electrically connected to the calibration point cloud data set acquisition module 3.

[0188] Among them, the first first-order image data acquisition sub-module is used to evenly divide the original image data into a plurality of square grids; the second first-order image data acquisition sub-module is used to predict a plurality of bounding boxes for all calibration points based on all square grids according to the target detection algorithm; the third first-order image data acquisition sub-module is used to respectively obtain the confidence level of each bounding box, and obtain the bounding box with the maximum confidence level and mark it as the first-order bounding box; the fourth first-order image data acquisition sub-module is used to calculate the intersection over union of each other bounding box and the first-order bounding box respectively; the fifth first-order image data acquisition sub-module is used to select the bounding boxes with the intersection over union greater than or equal to a preset threshold as the second-order bounding boxes; the sixth first-order image data acquisition sub-module is used to obtain the second-order bounding box with the highest confidence level and define it as the detection box of all calibration points; the seventh first-order image data acquisition sub-module is used to output all detection boxes to the original image data to form first-order image data.

[0189] Further, the second-order image data acquisition module 4 specifically includes a first second-order image data acquisition sub-module, a second second-order image data acquisition sub-module, a third second-order image data acquisition sub-module, a fourth second-order image data acquisition sub-module, and a fifth second-order image data acquisition sub-module that are electrically connected in sequence; the first second-order image data acquisition sub-module is electrically connected to the calibration point cloud data set acquisition module 3, and the fifth second-order image data acquisition sub-module is electrically connected to the millimeter-wave radar data acquisition module 5.

[0190] Among them, the first second-order image data acquisition sub-module is used to obtain the coordinate value of the th point in the point cloud data set .

[0191] The second second-order image data acquisition sub-module is used to define the plane equation of the first-order image data as .

[0192] The third second-order image data acquisition sub-module is used to solve the projected coordinate value of the th point according to Equation (1) :

[0193] (1).

[0194] The fourth second-order image data acquisition sub-module is used to solve the projected coordinate values of all points in the point cloud data set according to Equation (1).

[0195] The fifth second-order image data acquisition sub-module is used to output all the projected coordinate values to the first-order image data to form the second-order image data.

[0196] Further, the third-order image data acquisition module 6 specifically includes a first third-order image data acquisition sub-module, a second third-order image data acquisition sub-module, a third third-order image data acquisition sub-module, and a fourth third-order image data acquisition sub-module that are electrically connected in sequence; the first third-order image data acquisition sub-module is electrically connected to the millimeter-wave radar data acquisition module 5, and the fourth third-order image data acquisition sub-module is electrically connected to the fused image data acquisition module 7.

[0197] Among them, the first third-order image data acquisition sub-module is used to define the mapping relationship between the millimeter-wave radar data and the first-order image data through Equation (2):

[0198] (2).

[0199] Among them, is the radar position coordinate of the current calibration point, is the pixel coordinate value of the current calibration point based on the original image data, is the homography matrix.

[0200] The second third-order image data acquisition sub-module is used to solve the homography matrix according to Equation (2). The mapping relationship is obtained.

[0201] The third third-order image data acquisition sub-module is used to substitute all the radar position coordinates in the millimeter radar data into Equation (2) to solve for all the radar mapping coordinates.

[0202] The fourth third-order image data acquisition sub-module is used to output all the radar mapping coordinates to the second-order image data to form the third-order image data.

[0203] Furthermore, the fused image data acquisition module 7 specifically includes a first fused image data acquisition sub-module, a second fused image data acquisition sub-module, and a third fused image data acquisition sub-module that are electrically connected in sequence; the first fused image data acquisition sub-module is electrically connected to the fourth third-order image data acquisition sub-module, and the third fused image data acquisition sub-module is electrically connected to the weighted average method weight value acquisition module 8.

[0204] Among them, the first fused image data acquisition sub-module is used to fuse each pixel point of the third-order image data by the weighted average method according to Equation (3):

[0205] (3).

[0206] Among them, is the fused image data, are the pixel coordinates of each order of image data, is the weight value based on the third-order image data, is the weight value based on the second-order image data, is the weight value based on the first-order image data, is the original image data.

[0207] The second fused image data acquisition sub-module is used to obtain all the fused pixel points after traversing all the pixel points of the third-order image data according to Equation (3).

[0208] The third fused image data acquisition sub-module is used to integrate all the fused pixel points to obtain the fused image data.

[0209] Further, the weighted average method weight value acquisition module 8 specifically includes a first weighted average method weight value acquisition sub-module, a second weighted average method weight value acquisition sub-module, a third weighted average method weight value acquisition sub-module, a fourth weighted average method weight value acquisition sub-module, a fifth weighted average method weight value acquisition sub-module, and a sixth weighted average method weight value acquisition sub-module that are electrically connected in sequence; the first weighted average method weight value acquisition sub-module is electrically connected to the third fused image data acquisition sub-module, and the sixth weighted average method weight value acquisition sub-module is electrically connected to the final image data acquisition module 9.

[0210] Among them, the first weighted average method weight value acquisition sub-module is used to define a number of random solutions for the weight value of the weighted average method according to Equation (4), and define the optimization result of all random solutions as the maximum value of the clarity evaluation index.

[0211] (4).

[0212] Among them, is the set of all random solutions, are each random solution, is the label of the random solution, is the number of all random solutions; is the set of the speeds of all random solutions, are the speeds of each random solution respectively.

[0213] The second weighted average method weight value acquisition sub-module is used to initialize the position of each random solution, and update the current position and the current speed respectively based on the same random solution according to Equation (5):

[0214] (5).

[0215] Among them, is the speed of the th random solution at the th step, is the speed inertia of the th random solution at the th step, is the inertia coefficient, is the self-cognitive representation of the th random solution, is the social-cognitive representation of the th random solution; and are both learning factors, is a random number of is the individual optimal solution obtained by the th random solution, is the The globally optimal solution obtained by a random solution, At the step, the th random solution, At the step, the th random solution.

[0216] The third weighted average method weight value acquisition sub-module is used to iteratively update each random solution according to Equation (5) to update each and each .

[0217] The fourth weighted average method weight value acquisition sub-module is used to respectively determine whether the difference of each compared with the previous iteration is less than or equal to the first preset adaptation threshold.

[0218] The fifth weighted average method weight value acquisition sub-module is used to respectively determine whether the difference of each compared with the previous iteration is less than or equal to the first preset adaptation threshold. If so, then respectively determine whether the difference of each compared with the previous iteration is less than or equal to the second preset adaptation threshold.

[0219] The sixth weighted average method weight value acquisition sub-module is used to determine that the optimal solution of the weight value has been obtained if the difference of each compared with the previous iteration is less than or equal to the second preset adaptation threshold.

[0220] Furthermore, the third weighted average method weight value acquisition sub-module is specifically used to linearly decrease the inertia coefficient once according to Equation (6) based on each iteration:

[0221] (6).

[0222] Wherein, is the inertia coefficient after optimization of the th random solution at the th step, is the initial inertia coefficient, is the current iteration step number, is the maximum iteration step number.

[0223] It should be noted that this embodiment is a functional module item embodiment based on the above method embodiment. For additional content such as the preference, expansion, and illustration of this embodiment, please refer to the above method embodiment, and this embodiment will not be elaborated here.

[0224] In this embodiment, the original image data of all calibration points in the preset scenario is obtained; all calibration points are marked by detection frames in the original image data through a target detection algorithm to form first-order image data; a point cloud data set of all calibration points in the preset scenario is obtained through a first preset mapping strategy; all points in the point cloud data set are projected onto the first-order image data to form second-order image data; millimeter radar data of all calibration points in the preset scenario is obtained through a second preset mapping strategy; the millimeter radar data is mapped onto the second-order image data through a homography matrix to form third-order image data; each pixel point of the third-order image data is fused through a weighted average method, and after traversing all pixel points, the fused image data is obtained; the weight value of the weighted average method is iteratively optimized through a global optimization algorithm to make the clarity evaluation index of the fused image data reach the maximum value; the fused image data corresponding to the maximum value is obtained and defined as the final image data. This embodiment realizes the direct fusion of images, point clouds, and radars. Compared with other pairwise fusion technologies, the direct superposition has fewer calculation steps, reduces the error introduction in the intermediate process, improves the fusion accuracy, and makes the final fusion effect better. Moreover, the algorithms are all adaptive algorithms and do not require setting an absolute static threshold, avoiding large errors in the final fusion result due to incorrect threshold setting. At the same time, this embodiment also iteratively optimizes the weight values of the last fused pixel points through global optimization to keep the final fused image in the optimal fusion effect, avoiding problems such as image distortion, pixel point loss, and tearing caused by the direct superposition of pairwise fusion technologies.

[0225] Figure 3 An embodiment of the electronic device of the present application is shown. Refer to Figure 3 , the electronic device 10 includes a processor 101 and a memory 102 coupled to the processor 101.

[0226] The memory 102 stores program instructions for implementing the fusion method of the scenario data in any of the above embodiments.

[0227] The processor 101 is configured to execute the program instructions stored in the memory 102 to perform the fusion of scenario data.

[0228] Among them, the processor 101 can also be called a CPU (Central Processing Unit, central processing unit). The processor 101 may be an integrated circuit chip with signal processing capabilities. The processor 101 may also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0229] Further, Figure 4 is a schematic structural diagram of a storage medium according to an embodiment of the present application. Refer to Figure 4 . The storage medium 11 of the embodiment of the present application stores program instructions 111 that can implement all the above methods. Among them, the program instructions 111 can be stored in the above storage medium in the form of a software product, including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: various media that can store program codes such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, or a terminal device such as a computer, a server, a mobile phone, or a tablet.

[0230] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms.

[0231] In addition, each functional unit in various embodiments of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. The above is only the embodiment of the present application, and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made by using the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.

[0232] The specific implementation manners of the application have been described in detail above, but it is only an example, and the present application is not limited to the specific implementation manners described above. For those skilled in the art, any equivalent modification or substitution to the application is also within the scope of the present application. Therefore, equivalent transformations, modifications, improvements, etc. made without departing from the spirit and principles of the present application should all be covered within the scope of the present application.

Claims

1. A method for fusing scene data, wherein the scene data is derived from the same preset scene, and the preset scene has a preset number of calibration points, characterized in that: The fusion method comprises: Step S1, obtaining the original image data of all calibration points of the preset scene; Step S2, marking all calibration points in the original image data with detection frames using a target detection algorithm to form first-order image data; Step S3, obtaining a point cloud data set having all calibration points of the preset scene through a first preset surveying and mapping strategy; Step S4, projecting all points in the point cloud data set to the first-order image data to form second-order image data; Step S5, obtaining the millimeter radar data of all calibration points of the preset scene through a second preset surveying and mapping strategy; Step S6, mapping the millimeter radar data to the second-order image data through a homography matrix to form third-order image data; Step S7, fusing each pixel of the third-order image data by weighted averaging method, and obtaining fused image data after traversing all the pixels; Step S8, iterating the weight value of the weighted average method through a global optimization algorithm so that the clarity evaluation index of the fused image data reaches a maximum value; Step S9, obtaining fused image data corresponding to the maximum value and defining it as final image data; The first preset mapping strategy is one of a laser ranging method, a three-dimensional scanner method, and a structured light method; The second preset surveying and mapping strategy is one of online platform acquisition and self-measurement; Step S7, fusing each pixel of the third-order image data by weighted average method, and obtaining fused image data after traversing all the pixels, including: Step S71, each pixel of the third-order image data is fused by weighted average method: Step S72, traversing all pixel points of the third-order image data to obtain all fused pixel points; Step S73, integrating all fused pixel points to obtain the fused image data; Step S8, iterating the weight value of the weighted average method through a global optimization algorithm so that the clarity evaluation index of the fused image data reaches a maximum value, including: Step S81, defining a plurality of random solutions for the weight value of the weighted average method according to formula (4), and defining the optimization result of all random solutions as the clarity evaluation index reaching the maximum value; (4); in, is the set of all random solutions, For each random solution, is the label of the random solution, is the number of all random solutions; is the set of velocities of all random solutions, are the speeds of each random solution respectively; Step S82, initialize the position of each random solution, and update the current position and current speed respectively based on the same random solution according to formula (5): (5); in, For the The random solution is The speed of the step, For the The random solution is The speed inertia of the step, is the inertia coefficient, For the The self-perception representation of a random solution, For the social cognitive representation of a random solution; and are learning factors, for A random number, For the The individual optimal solution obtained by random solutions is For the The global optimal solution obtained by random solutions is For the Step 1 A random solution, For the Step 1 A random solution; Step S83, iterate each random solution according to equation (5) to update each and each ; Step S84, determine each Compared with the previous iteration, whether the difference is less than or equal to the first preset adaptation threshold, if each If the difference between the values ​​in the previous iteration and the values ​​in the previous iteration are all less than or equal to the first preset adaptation threshold, step S85 is executed; Step S85, determine each Compared with the difference of the previous iteration, whether it is less than or equal to the second preset adaptation threshold, if each If the difference between the values ​​in the previous iteration and the values ​​in the previous iteration are all less than or equal to the second preset adaptation threshold, step S86 is executed; Step S86, determining whether the optimal solution of the weight value has been obtained; The clarity evaluation index includes one of grayscale variance, edge gradient, information entropy, spectrum entropy, and structural similarity, wherein grayscale variance reflects the distribution of grayscale values ​​in the image, and the larger the grayscale variance, the clearer the image; edge gradient reflects the sharpness of the edge in the image, and the larger the edge gradient, the clearer the image; information entropy reflects the richness of information in the image, and the larger the information entropy, the clearer the image; spectrum entropy reflects the richness of frequency information in the image, and the larger the spectrum entropy, the clearer the image; and structural similarity reflects the similarity of structural information in the image, and the higher the structural similarity, the clearer the image.

2. The fusion method according to claim 1, characterized in that: Step S2, marking all calibration points in the original image data with a detection frame by using a target detection algorithm to form first-order image data, including: Step S21, dividing the original image data into a plurality of square grids on average; Step S22, predicting a number of bounding boxes for all calibrated points based on all square grids according to the target detection algorithm; Step S23, respectively obtaining the confidence of each bounding box, obtaining the bounding box with the largest confidence and marking it as a first-order bounding box; Step S24, calculating the intersection-over-union ratio of each other bounding box with the first-order bounding box; Step S25, selecting a bounding box whose intersection-over-union ratio is greater than or equal to a preset threshold as a second-order bounding box; Step S26, obtaining the second-order bounding box with the highest confidence and defining it as the detection box of all calibration points; Step S27 , outputting all detection frames to the original image data to form the first-order image data.

3. The fusion method according to claim 1, characterized in that: Step S4, projecting all points in the point cloud data set to the first-order image data to form second-order image data, including: Step S41, obtaining the first The coordinates of the points ; Step S42, defining the plane equation of the first-order image data as ; Step S43, solve the first The projected coordinates of the points : (1); Step S44, solving the projection coordinate values ​​of all points in the point cloud data set according to formula (1); Step S45 , outputting all projection coordinate values ​​to the first-order image data to form the second-order image data.

4. The fusion method according to claim 1, characterized in that: Step S6, mapping the millimeter radar data to the second-order image data through a homography matrix to form third-order image data, including: Step S61, defining the mapping relationship between the millimeter radar data and the first-order image data by equation (2): (2); in, is the radar position coordinate of the current calibration point, is the pixel coordinate value of the current calibration point based on the original image data, is the homography matrix; Step S62, solving the homography matrix according to equation (2): Obtaining the mapping relationship; Step S63, substituting all radar position coordinates in the millimeter radar data into equation (2) to obtain all radar mapping coordinates; Step S64: output all radar mapping coordinates to the second-order image data to form the third-order image data.

5. The fusion method according to claim 1, characterized in that: Step S83, iterate each random solution according to equation (5) to update each and each ,include: Step S831, based on each iteration, the inertia coefficient is linearly reduced once according to formula (6): (6); in, For the The random solution is The inertia coefficient after step optimization, is the initial inertia coefficient, is the current iteration number, is the maximum number of iteration steps.

6. A scene data fusion device, the scene data fusion device is applied to the fusion method according to any one of claims 1 to 5, characterized in that: The fusion device comprises: An original image data acquisition module, used to acquire original image data of all calibration points of the preset scene; A first-order image data acquisition module, used to mark all calibration points in the original image data with a detection frame by using a target detection algorithm to form first-order image data; A calibration point cloud data set acquisition module, used to acquire a point cloud data set having all calibration points of the preset scene through a first preset surveying and mapping strategy; A second-order image data acquisition module, used for projecting all points in the point cloud data set to the first-order image data to form second-order image data; A millimeter radar data acquisition module, used to acquire the millimeter radar data of all calibration points of the preset scene through a second preset surveying and mapping strategy; A third-order image data acquisition module, used for mapping the millimeter radar data to the second-order image data through a homography matrix to form third-order image data; A fused image data acquisition module is used to fuse each pixel of the three-order image data by weighted averaging method, and obtain fused image data after traversing all pixels; A weighted average method weight value acquisition module, used for iterating the weight value of the weighted average method through a global optimization algorithm so that the clarity evaluation index of the fused image data reaches a maximum value; A final image data acquisition module, used for acquiring fused image data corresponding to the maximum value and defining it as final image data; The first preset mapping strategy is one of a laser ranging method, a three-dimensional scanner method, and a structured light method; The second preset surveying and mapping strategy is one of online platform acquisition and self-measurement; The fused image data acquisition module includes a first fused image data acquisition submodule, a second fused image data acquisition submodule, and a third fused image data acquisition submodule which are electrically connected in sequence; the first fused image data acquisition submodule is electrically connected to the fourth third-order image data acquisition submodule, and the third fused image data acquisition submodule is electrically connected to the weighted average method weight value acquisition module 8; The first fused image data acquisition submodule is used to fuse each pixel of the third-order image data by weighted average method; The second fused image data acquisition submodule is used to traverse all pixel points of the third-order image data to obtain all fused pixel points; The third fused image data acquisition submodule is used to integrate all fused pixel points to obtain fused image data; The weighted average method weight value acquisition module specifically includes a first weighted average method weight value acquisition submodule, a second weighted average method weight value acquisition submodule, a third weighted average method weight value acquisition submodule, a fourth weighted average method weight value acquisition submodule, a fifth weighted average method weight value acquisition submodule, and a sixth weighted average method weight value acquisition submodule, which are electrically connected in sequence; the first weighted average method weight value acquisition submodule is electrically connected to the third fused image data acquisition submodule, and the sixth weighted average method weight value acquisition submodule is electrically connected to the final image data acquisition module; The first weighted average method weight value acquisition submodule is used to define a number of random solutions for the weight value of the weighted average method according to formula (4), and define the optimization result of all random solutions as the clarity evaluation index reaching the maximum value; (4); in, is the set of all random solutions, For each random solution, is the label of the random solution, is the number of all random solutions; is the set of velocities of all random solutions, are the speeds of each random solution respectively; The second weighted average weight value acquisition submodule is used to initialize the position of each random solution, and based on the same random solution, update the current position and current speed according to formula (5): (5); in, For the The random solution is The speed of the step, For the The random solution is The speed inertia of the step, is the inertia coefficient, For the The self-perception representation of a random solution, For the social cognitive representation of a random solution; and are learning factors, for A random number, For the The individual optimal solution obtained by random solutions is For the The global optimal solution obtained by random solutions is For the Step 1 A random solution, For the Step 1 A random solution; The third weighted average weight value acquisition submodule is used to iterate each random solution according to formula (5) to update each and each ; The fourth weighted average method weight value acquisition submodule is used to determine each Whether the difference compared to the previous iteration is less than or equal to a first preset adaptation threshold; The fifth weighted average weight value acquisition submodule is used for each Compared with the difference of the previous iteration, if the difference is less than or equal to the first preset adaptation threshold, then each Whether the difference compared to the previous iteration is less than or equal to a second preset adaptation threshold; The sixth weighted average weight value acquisition submodule is used for each If the difference between the values ​​in the previous iteration and the values ​​in the previous iteration are all less than or equal to the second preset adaptation threshold, it is determined that the optimal solution of the weight value has been obtained; The clarity evaluation index includes one of grayscale variance, edge gradient, information entropy, spectrum entropy, and structural similarity, wherein grayscale variance reflects the distribution of grayscale values ​​in the image, and the larger the grayscale variance, the clearer the image; edge gradient reflects the sharpness of the edge in the image, and the larger the edge gradient, the clearer the image; information entropy reflects the richness of information in the image, and the larger the information entropy, the clearer the image; spectrum entropy reflects the richness of frequency information in the image, and the larger the spectrum entropy, the clearer the image; and structural similarity reflects the similarity of structural information in the image, and the higher the structural similarity, the clearer the image.

7. An electronic device, characterized in that: It comprises a processor and a memory coupled to the processor, wherein the memory stores program instructions executable by the processor; when the processor executes the program instructions stored in the memory, the fusion method as described in any one of claims 1 to 5 is implemented.

8. A storage medium, characterized in that: The storage medium stores program instructions, and when the program instructions are executed by the processor, the fusion method according to any one of claims 1 to 5 can be implemented.

Citation Information

Patent Citations

  • Target fusion method and device based on radar and image, equipment and storage medium

    CN116630216A

  • Method for comprehensively measuring quality of Leiyu fusion system

    CN118918707A