Road apparent disease real-time identification method based on visible light and infrared image multi-modal fusion

By adopting a multimodal fusion method in road disease detection, combining feature decomposition and neural network model, the problem of difficulty in detecting road disease and low image depth fusion efficiency in complex scenarios is solved, and higher detection accuracy and real-time performance are achieved.

CN120147709AActive Publication Date: 2025-06-13HARBIN INST OF TECH +2

Patent Information

Application Number
CN202510211328.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-06-13
Estimated Expiration
2045-02-25

AI Technical Summary

Technical Problem

The problem of difficulty in detecting road diseases and low image depth fusion efficiency in complex scenarios.

Method used

Real-time identification method for road apparent diseases based on multimodal fusion of visible light and infrared images is adopted, and real-time identification of road apparent diseases is achieved through the combination of feature decomposition and extraction, image registration and neural network models.

Benefits of technology

The accuracy and real-time detection of road apparent disease in complex scenarios are improved, and the problems of poor multimodal image registration effect and low image depth fusion efficiency in traditional methods are overcome.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147709A_ABST
    Figure CN120147709A_ABST
Patent Text Reader

Abstract

The invention discloses a road apparent disease real-time identification method based on visible light and infrared image multi-modal fusion, and belongs to the technical field of road apparent disease real-time detection based on multi-modal image fusion. In order to solve the problems of difficult road disease detection and low image depth fusion efficiency in a complex scene, the method comprises the following steps: carrying out feature decomposition extraction on a visible light image to obtain a corresponding base layer and a detail layer, and carrying out feature decomposition extraction on an infrared image to obtain a corresponding base layer and a detail layer; obtaining fusion information FB based on the base layer of the visible light image and the base layer of the infrared image, and obtaining fusion information FD based on the detail layer of the visible light image and the detail layer of the infrared image; and a fusion image F is obtained based on the FB and the FD, and real-time identification of road apparent diseases is realized by adopting a neural network model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a real-time detection method for road surface diseases by multi-modal image fusion. Background Art

[0002] Road surface diseases directly affect driving safety and the service life of roads. Therefore, accurate detection of road surface diseases can provide reliable and effective technical support and data support for road maintenance management decisions, and is also an important part of intelligent road management and maintenance work. In recent years, automated road data collection equipment has developed rapidly, and road surface disease detection technology based on image processing has gradually become the mainstream research direction, and road surface disease detection has also transitioned from traditional manual inspections to semi-automated and non-destructive full-automation. However, due to the complexity of road surface textures and the external environment, traditional image analysis methods are difficult to ensure stable and good detection results in various complex road surface scenarios.

[0003] Although the image information obtained based on digital camera technology has the advantages of intuitiveness and rich features, this image acquisition and analysis technology is sensitive to shadows, oil stains, uneven lighting, night-time, and other complex interference factors, and there are also a large number of misjudgments and missed detections, making it difficult to meet the high-precision, multi-level, and all-weather work requirements of road detection in the new era. For this reason, some researchers have considered introducing infrared imaging technology into the field of road engineering. It can work without external light sources and is sensitive to temperature differences, which gives it unique advantages in detecting road damage (especially under conditions of extremely low light and different temperature difference ranges).

[0004] Although infrared imaging technology has unique advantages in extreme scenarios, it also has certain limitations. Infrared sensors image by capturing the thermal radiation information of objects and have a certain adaptability to the interference of external environmental factors. However, the detection algorithms based on infrared images also face the dilemmas of insufficient texture details and difficulty in identifying distant targets. Visible light sensors image by collecting the reflected light on the surface of objects, and the images contain rich texture detail information, but they also cannot well handle the interference of shadows, smoke, etc. in complex ground environments. To overcome these limitations, a system based on multi-modal fusion of infrared and visible light images has been proposed, which integrates the complementary information in the acquired source images to generate a high-contrast fusion image that can not only highlight significant targets but also contain rich texture details.

[0005] However, traditional methods cannot achieve synchronous acquisition and feature fusion of visible light and infrared images, and complex algorithms are required to ensure precise registration and rapid fusion between multi-modal images. Moreover, most existing disease detection models have problems such as inconsistent sample image features, lack of image data, low accuracy, and poor stability, and cannot well meet the needs of modern road detection and maintenance. Summary of the Invention

[0006] The object of the present invention is to solve the problems of difficult detection of road diseases in complex scenarios and low efficiency of image depth fusion.

[0007] A real-time road surface disease recognition method based on multimodal fusion of visible light and infrared images, comprising the following steps:

[0008] For the visible light image I vi , perform feature decomposition extraction to obtain the corresponding base layer B vi and detail layer D vi ; for the infrared image I iR , perform feature decomposition extraction to obtain the corresponding base layer B iR and detail layer D iR ;

[0009] The formula for feature decomposition extraction is as follows:

[0010] B x =WLSR(I x ), D x =I x -B x

[0011] wherein, WLSR(·) is edge filtering processing; B x is the base layer obtained by feature decomposition extraction, D x is the detail layer obtained by feature decomposition extraction, x = vi or iR, vi and iR respectively represent the visible light image and the infrared image, and I x represents the visible light image I vi or the infrared image I iR ;

[0012] Based on the base layer B vi of the visible light image and the base layer B iR of the infrared image, obtain the fusion information F B , and at the same time, based on the detail layer D vi of the visible light image and the detail layer D iR of the infrared image, obtain the fusion information F D ; furthermore, based on F B and F D , obtain the fusion image F = F B +F D ;

[0013] Based on the fusion image F, use a neural network model to realize real-time recognition of road surface diseases.

[0014] Furthermore, before recognition, it is necessary to perform image registration on the visible light image and the infrared image, including the following steps:

[0015] First, taking the visible light image as the reference image, transform the infrared image into the reference field of the visible light image; then register the image coordinates of the infrared image transformed into the reference field of the visible light image with the dynamic pixels of the visible light image.

[0016] Furthermore, taking the visible light image as the reference image, the process of transforming the infrared image into the reference field of the visible light image includes the following steps:

[0017] Taking the visible light image as the reference image, obtain the image coordinates of the infrared image transformed into the reference field of the visible light image through the homography mapping transformation formula described as follows

[0018]

[0019] where x iR , y iR represent the source pixel point coordinates from the infrared image, are the image coordinates of the infrared image transformed into the reference field of the visible light image; K vi is the intrinsic matrix obtained with the visible light image as the reference image; K iR is the intrinsic matrix obtained with the infrared image as the reference image; vi R iR is the rotation matrix for the mapping transformation from the infrared image coordinate system to the visible light image coordinate system.

[0020] Furthermore, the mapping transformation for registering the image coordinates of the infrared image transformed into the reference field of the visible light image with the dynamic pixels of the visible light image is as follows:

[0021]

[0022] where t x is the horizontal displacement, t y is the vertical displacement, is the rotation angle, s is the scaling factor; x vi , y vi represent the source pixel point coordinates of the visible light image.

[0023] Furthermore, the edge filtering processing formula is as follows:

[0024] Y = WLSR(X) = (1 + λL X ) -1 X, where

[0025] where X represents the input image before filtering, Y represents the edge filtering result; λ is the balance coefficient; A x , A y is the diagonal weight matrix, and the diagonal elements of the matrix are μ x,p , μy,p , μ x,p , μ y,p are the weight coefficients of the gradients in the x and y directions corresponding to the pixel of the image X at the spatial position p; M x , M y are the discrete difference operator matrices corresponding to the difference in the x direction and the difference in the y direction.

[0026] Furthermore, the diagonal weight matrix A x , A y is specifically as follows:

[0027] A x = diag(μ x,1 , μ x,2 , μ x,3 , μ x,4 , …, μ x,n )

[0028] A y = diag(μ y,1 , μ y,2 , μ y,3 , μ y,4 , …, μ y,n )

[0029] where n represents the total number of pixels in the image.

[0030] Furthermore, the discrete difference operator matrix corresponding to the difference in the x direction the discrete difference operator matrix corresponding to the difference in the y direction

[0031] Furthermore, the process of obtaining the fusion information F vi from the base layer B iR of the visible light image and the base layer B B of the infrared image includes:

[0032] For B vi and B iR , let I k represent the intensity value of the k-th pixel, and use the significance intensity difference distribution formula to obtain the significance intensity values Λ vi , Λ iR of the visible light and infrared images, where N is the number of pixels in B x , and B x represents B vi or B iR ;

[0033] Then calculate the adaptive weight of the fusion and then obtain the fusion information F B of the base layer = ωB vi+(1 - ω)B iR 。

[0034] Further, based on the detail layer D of the visible light image vi and the detail layer D of the infrared image iR to obtain the fusion information F D The process includes:

[0035] First, the detail layer D of the visible light image vi is subjected to SVF multi-scale decomposition to obtain the enhanced detail layer D E_vi :

[0036]

[0037] Among them, represents the SVF filtering results at different scales, represents the detail layer corresponding to , α, β are detail layer weight coefficients, and D E_vi represents the enhanced detail layer of the visible light image;

[0038] The detail layer D of the infrared image iR is processed in the same way to obtain the enhanced detail layer D of the infrared image E_iR ;

[0039] Furthermore, the fusion information F of the detail layer is obtained D = D E_vi + D E_iR 。

[0040] Further, in the process of using a neural network model to achieve real-time identification of road surface diseases, the neural network model used is the YOLOv8 network model.

[0041] Beneficial effects:

[0042] Aiming at the problem of poor multi-modal image registration effect in the existing methods, the present invention proposes an image registration scheme for visible light images and infrared images, which can effectively improve the multi-modal image registration effect and thus improve the subsequent recognition effect. Aiming at the problems of difficult road disease detection and low image depth fusion efficiency in complex scenarios, the method of the present invention breaks through the limitations of traditional single-sensor disease detection, establishes a multi-modal fusion system of visible light and infrared images, introduces new algorithms to ensure accurate registration and rapid fusion between multi-modal images, and uses a deep learning recognition network for training, and then realizes real-time identification of road surface diseases. The present invention can improve the accuracy and real-time performance of the road surface disease detection model in complex scenarios. Description of the Drawings

[0043] Figure 1It is a flowchart of a real-time recognition method for road surface diseases based on multimodal fusion of visible light and infrared images;

[0044] Figure 2 It is a schematic diagram of the image registration structure of the present invention;

[0045] Figure 3 It is a schematic diagram of the image fusion network structure;

[0046] Figure 4 It is a schematic diagram of the intelligent and rapid recognition network structure of the fused image;

[0047] Figure 5 It is the visualization effect of object detection of the visible light and infrared fusion image of the present invention. Detailed implementation manners

[0048] To make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be described below through specific embodiments shown in the drawings. However, it should be understood that these descriptions are only exemplary and are not intended to limit the scope of the present invention. In addition, in the following description, the descriptions of well-known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present invention.

[0049] Here, it should also be noted that in order to avoid obscuring the present invention due to unnecessary details, only the structures and / or processing steps closely related to the solution of the present invention are shown in the drawings, and other details less related to the present invention are omitted.

[0050] Detailed implementation manner 1: In combination with Figures 1-4 To illustrate this implementation manner,

[0051] A real-time recognition method for road surface diseases based on multimodal fusion of visible light and infrared images described in this implementation manner includes the following steps:

[0052] Step 1: Preprocess the collected visible light and infrared image data, including noise removal and contrast enhancement processing;

[0053] The infrared and visible light images of road surface diseases are collected by vehicle-mounted and handheld methods. The collection of visible light and infrared image data can be realized by using a camera and an infrared thermal imager device respectively, or an infrared thermal imager including an infrared lens and a visible light lens can be used. The FOTRIC 340X + cloud thermal imager series used in this implementation manner includes an infrared lens and a visible light lens. Considering that the all-day temperature change has a significant impact on the image temperature difference effect, the data at different time periods of a day are collected. Then, the collected data is subjected to noise removal and contrast enhancement to further improve the image quality and visual perception, and to improve the feature discrimination and extraction for subsequent graphic registration.

[0054] Step 2: Perform image registration on the visible light image and the infrared image processed in Step 1, including operations such as feature detection, feature matching, transformation model estimation, optimization iteration, similarity measurement, and image registration between images;

[0055] The process of performing visible light image and infrared image registration on the preprocessed data in Step 1 is as Figure 2 shown. Before this, it is necessary to calculate the intrinsic matrix K of the camera, the radial distortion coefficients (k 1 , k 2 , k 3 ), and the tangential distortion coefficients (p 1 , p 2 ) through calibration. The intrinsic matrix K is composed of the focal lengths f x , f y and the principal points c x , c y usually close to the center of the image, that is:

[0056]

[0057] At the same time, it is also necessary to determine the external parameters of the camera. In previous studies, most used the infrared image as the reference, which is beneficial for night detection. Considering the weak edge information of road surface diseases, the present invention takes the visible light image as the reference image to obtain the external parameters vi R iR , vi T iR for the mapping transformation from the infrared image coordinate system to the visible light image coordinate system. vi R iR , vi T iR are the rotation matrix and the translation vector respectively; and this mapping transformation process can be represented by the matrix vi G iR . Since the topological structure between the conversions of different coordinate systems is extremely complex and difficult to represent with a simple plane model, a simplification process of the homography mapping transformation is considered, that is:

[0058]

[0059] where x iR , y iR represent the source pixel coordinates from the infrared image, is the image coordinate in the infrared image transformed to the visible light image reference field; K vi is the intrinsic matrix obtained with the visible light image as the reference image; K iR is the intrinsic matrix obtained with the infrared image as the reference image.

[0060] The image coordinates of the infrared image transformed into the visible light image reference field are obtained through the foregoing process The above is only a rough correction registration that fixes only pixel points, and it is difficult to be well applied to the correction of dynamic pixels. Therefore, on the basis of the foregoing processing process, the present invention considers introducing the RIFT algorithm for fine dynamic pixel registration, that is, registering the image coordinates of the infrared image transformed into the visible light image reference field with the dynamic pixels of the visible light image. The two graphic transformations of this algorithm can be represented by a 3×3 matrix H RIFT to represent, this matrix combines the rigid transformation with the scaling factor, and to a certain extent, it can well explain the mapping transformation process of dynamic pixel points.

[0061]

[0062] Among them, t x is the horizontal displacement, t y is the vertical displacement, is the rotation angle, s is the scaling factor; x vi 、y vi represent the source pixel coordinates of the visible light image.

[0063] The present invention aims at the problems of difficult image registration and correction and low conversion efficiency, and realizes the rapid registration and robust alignment of infrared images and visible light images based on the preprocessing of rough correction registration and feature-based matching.

[0064] Step 3: Perform fast fusion processing on the registered infrared image and visible light image in Step 2, including feature decomposition extraction, feature fusion, and image reconstruction processes on the multimodal image to obtain a fused image;

[0065] The process of fusing the images registered in Step 2 is as Figure 3 shown. When performing feature decomposition extraction on the multimodal image, the present invention improves on the basis of the least squares filter, pays more attention to the edge gradient problem of cracks, and proposes an edge-preserving filter to decompose the original image into a base layer and a detail layer. Specifically, for the input image X, the goal is to make the filtering result Y as close as possible to the source image X, keep smooth in the area with small gradients, and be as close as possible to the source image in the edge area with strong gradients. This method innovatively converts the filtering problem into the problem of minimizing the loss function f(X), which can be expressed as:

[0066]

[0067] Among them, p represents the spatial position of the pixel; X p 、Y p are the pixels of the image X and the filtering result Y at the spatial position p; λ is the balance coefficient; μ x ,μ yThe weight coefficients representing the gradients in the x and y directions, μ x,p , μ y,p is the μ corresponding to the pixel of image X at spatial position p x , μ y ; l is the luminance channel of the input image X, α represents the sensitivity to gradient changes, and κ is a constant coefficient.

[0068] After performing a series of matrix transformations on the above equation and introducing the diagonal weight matrix A x , A y and the discrete difference operator matrix M x , M y , the matrix form of f′(Y) can be obtained, i.e.:

[0069] f′(Y) = 2Y - 2X + 2λ(M x T A x M x + M y T A y M y )Y = 0

[0070] f′(Y) is the derivative of f(Y);

[0071] The diagonal weight matrix A x , A y : The diagonal elements of the matrix correspond to the weight coefficients μ x,p , μ y,p , and specifically can be expressed as:

[0072] A x = diag(μ x,1 , μ x,2 , μ x,3 , μ x,4 , …, μ x,n )

[0073] A y = diag(μ y,1 , μ y,2 , μ y,3 , μ y,4 , …, μ y,n )

[0074] where n represents the total number of pixels in the image;

[0075] The discrete difference operator matrix M x , M y : Usually constructed as a diagonal matrix to represent the differences between adjacent pixels; M x represents the difference in the x direction, M yDenote the difference in the y direction, which can be specifically expressed as:

[0076]

[0077] It can be derived that:

[0078] Y = (1 + λL X ) -1 X = WLSR(X)

[0079]

[0080] where WLSR(·) is edge filtering processing;

[0081] For the visible light image denoted as I vi and the infrared image after dynamic pixel registration denoted as I iR , calculate the basic layer B vi and the detail layer D iR of I x and I x respectively:

[0082] B x = WLSR(I x ), D x = I x - B x

[0083] where x = vi or iR, vi and iR represent the visible light image and the infrared image respectively, that is, I x represents the visible light image I vi or the infrared image I iR ;

[0084] After the above image decomposition, it is necessary to fuse the decomposed basic layer and detail layer, and this part needs to be carried out in two parts:

[0085] (1) During the basic layer fusion process, consider introducing the vector space model and the idea of adaptive weight assignment. For the basic layer B x corresponding images B vi , B iR , use I k to represent the intensity value of the kth pixel, then the significant intensity difference distribution can be expressed as Λ k , that is:

[0086]

[0087] where Λ k can be used to calculate the intensity value of each pixel point, and the range is [0, 1]; where N is the number of pixels in B x .

[0088] The saliency intensity values Λ of visible light and infrared images can be calculated respectively using the above formula vi , Λ iR , and at this time, the adaptive weight ω for fusing the base layer can be calculated, which can be expressed as:

[0089]

[0090] At this time, the fused information F of the base layer can be calculated B :

[0091] F B = ωB vi + (1 - ω)B iR

[0092] where B vi represents the base layer of the visible light image; B iR represents the base layer of the infrared image;

[0093] (2) For the detail layer fusion, first, a sub-window variance filter is used to obtain the local edge information of the image, and a spatial statistical model is used to further improve the edge perception ability of the filter. This filter needs to consider the edge statistical model and the properties of the sub-window variance, and the output result can be expressed as a linear combination of the source image and the image after smoothing filtering.

[0094] For the detail layer D x (i.e., D vi , D iR ), perform SVF (sub-window variance filtering) multi-scale decomposition;

[0095] The SVF sub-window variance filtering process is as follows:

[0096]

[0097] where I p represents the local pixel block centered on pixel p in the image corresponding to the detail layer D x , I′ p is the local pixel block after sub-window variance filtering; φ p is the contribution parameter, used to control the contribution of the local pixel block I p in I′ p ; F(·) is the smoothing filter; ω p represents the support domain of the filter centered on pixel p, and k i represents the intensity value of the i-th pixel in the support domain.

[0098] Regarding the value of φ p , it depends on the global variance value and the sub-window variance value to be filtered. Assume is the global variance value to be filtered, represents the variance values corresponding to dividing the area to be filtered into four sub - windows. At this time, it can be set that:

[0099]

[0100] where, ε is the regularization parameter.

[0101] It can be seen that the above - obtained base layer B x and detail layer D x The filtering method also follows the Laplacian pyramid principle. For this reason, the present invention performs multi - scale decomposition based on SVF (sub - window variance filtering) on D vi and can obtain:

[0102]

[0103] where, respectively represent the fine - scale detail layer and the small - scale detail layer, and α, β represent the weight coefficients of the two detail layers. D E_vi represents the enhanced visible - light image detail layer.

[0104] Perform multi - scale decomposition of SVF on D iR to obtain D E_iR , D E_iR represents the enhanced infrared - image detail layer.

[0105] The decomposition scale of the multi - scale decomposition of SVF can be determined according to the actual situation.

[0106] Finally, the fusion information is obtained: F D = D E_vi + D E_iR .

[0107] It is necessary to perform inverse - transformation image fusion on the fused base layer and detail layer:

[0108] F = F B + F D

[0109] Step 4: Label the fused image F and divide it into a training set, a validation set, and a test set;

[0110] During the labeling process, use the LabelImg software to label the image to form a corresponding txt file, which contains the type, size, and position information of the target of interest to ensure the accuracy and reliability of the data set.

[0111] Step 5: Input the reconstructed image F corresponding to the dataset in Step 4 into the road apparent disease rapid detection network model, perform iterative optimization training on it to obtain a pre-trained weight model, and import this model into the mobile device to achieve real-time recognition of road apparent diseases.

[0112] During the process of real-time recognition, obtain the visible light and infrared images corresponding to the road to be recognized, obtain the reconstructed image F based on the visible light and infrared images, and use the road apparent disease rapid detection network model to achieve real-time recognition of road apparent diseases.

[0113] Use the training set obtained in Step 4 to train and optimize the YOLOv8 model. Its training structure diagram is as Figure 4 shown. First, the input image sequentially passes through multiple convolutional layers, C2f, and SPPF modules for deep feature extraction to achieve efficient fusion of local and global features. Then, the extracted feature maps are passed to the Neck network. By constructing a multi-scale feature pyramid and achieving beneficial transfer of cross-layer information, the network can better perceive the features of targets at different scales. The model uses anchor boxes to predict bounding boxes on the feature maps at multiple scales, predicting the position offset, confidence, and class probability of each anchor box. Finally, it enters the detection network. This area will stack the optimized feature maps to obtain a feature map with region proposals, and use fully connected operations for target localization, thereby performing bounding box regression and classification regression of image diseases. The prediction results undergo post-processing, including threshold filtering and non-maximum suppression, to remove low-confidence predictions and merge overlapping boxes, and finally output the class, bounding box coordinates, and confidence of the target, and finally obtain the accurate information of the detected underground target space body.

[0114] During the model training process, the present invention uses Precision, Recall, and mean Average Precision (mAP) as model evaluation indicators. In order to give full play to the advantages of the YOLOv8 architecture, the present invention continuously iterates the model structure and parameter configuration to achieve the best balance among these three indicators. The calculation formulas for precision P, recall R, and mean average precision mAP are as follows:

[0115] Precision P:

[0116]

[0117] Recall R:

[0118]

[0119] Mean average precision and mean average precision value mAP:

[0120]

[0121] Wherein, TP is the number of positive examples of the target and predicted as positive examples in the object detection task; FP is the number of negative examples of the target and predicted as positive examples in the object detection task; P(r) is the curve corresponding to both precision and recall; N is the number of categories in the multi-class object detection task.

[0122] Example:

[0123] Use an infrared thermal imager to detect about 10 km of urban roads and campus roads. Preprocess the collected infrared data according to the above steps, then perform model design. At the same time, data augmentation and image annotation are required to make a dataset, and finally, model experiments and comparative analysis are carried out.

[0124] The results of the model experiments are shown in Table 1 below:

[0125] Table 1 Comparison of iterative evaluation indicators of each model

[0126]

[0127] It can be seen that Yolo has played its unique advantages in the fields of object recognition and real-time detection. Although it has now developed to Yolov11, it can be seen from the results that the latest detection results are not necessarily the best, and it needs to be trained specifically according to the specific detection scenario, the target of concern, and the region of interest. The results show that Yolov8 has excellent performance in the recognition and detection of infrared and visible light fusion images, exceeding other models in various indicators, and can be well applied to the detection of road damage in extreme scenarios.

[0128] The recognition effect for visible light and infrared fusion images is as Figure 5 shown.

[0129] Based on rich practical experience and professional knowledge in this type of product, the present invention designs a real-time recognition method for road surface diseases based on multi-modal fusion of visible light and infrared images, which can solve problems such as poor detection effect of road damage under extremely low light conditions, difficult data interpretation, large workload of target recognition and classification, and poor real-time performance to a certain extent. Compared with the prior art, the model of the present invention performs cross-scale registration and fusion on infrared and visible light multi-modal image data to solve the problem of detecting road surface diseases in complex scenarios, and then uses an intelligent recognition algorithm to train and evaluate the fused image, breaking through the disadvantages of missed detection, misdetection, and low efficiency of conventional manual inspections, laying a foundation for building a disease detection model in extreme scenarios, enabling it to be more widely applied to various portable road information collection devices and online disease detection platforms, and promoting the construction of a smart road state online holographic precise sensing, diagnosis, and evaluation system with higher detection accuracy and lower computing cost.

[0130] The foregoing has shown and described the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments, and what is described in the above embodiments and the specification is only to illustrate the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will also have various changes and improvements, and these changes and improvements fall within the scope of the present invention claimed. The scope of the present invention claimed is defined by the appended claims and their equivalents.

Claims

1. A method for real-time recognition of road surface defects based on multimodal fusion of visible light and infrared images, characterized in that: The following steps are involved: For visible light image I vi , perform feature decomposition and extraction to obtain the corresponding base layer B vi and detail layer D vi ; For infrared image I iR , perform feature decomposition and extraction to obtain the corresponding base layer B iR and detail layer D iR ; The formula for feature decomposition extraction is as follows: B x =WLSR(I x ),D x =I x -B x Among them, WLSR(·) is edge filtering processing; B x is the base layer extracted by feature decomposition, D x is the detail layer extracted by feature decomposition, x = vi or iR, vi and iR represent visible light image and infrared image respectively, I x Represents the visible light image I vi or infrared image I iR ; Base layer B based on visible light image vi and the base layer B of the infrared image iR Get the fusion information F B , while the detail layer D based on the visible light image vi and the detail layer D of the infrared image iR Get the fusion information F D ; Based on F B and F D Get the fused image F = F B +F D ; Based on the fused image F, a neural network model is used to realize real-time recognition of road surface defects.

2. The method for real-time recognition of road surface defects based on multimodal fusion of visible light and infrared images according to claim 1 is characterized by: Before recognition, the visible light image and the infrared image need to be registered, including the following steps: Firstly, the visible light image is used as the reference image, and the infrared image is transformed into the visible light image reference field; then, the image coordinates of the infrared image transformed into the visible light image reference field are dynamically aligned with the visible light image.

3. The method for real-time recognition of road surface defects based on multimodal fusion of visible light and infrared images according to claim 2 is characterized by: Taking the visible light image as the reference image, the process of transforming the infrared image into the visible light image reference field includes the following steps: Taking the visible light image as the reference image, the image coordinates of the infrared image transformed to the visible light image reference field are obtained by the homography mapping transformation formula described below: Among them, x iR ,y iR represents the source pixel coordinates from the infrared image, K is the image coordinates transformed from the infrared image to the visible light image reference field; vi is the intrinsic matrix obtained by taking the visible light image as the reference image; K iR is the intrinsic matrix obtained by taking the infrared image as the reference image; vi R iR The rotation matrix for mapping transformation from infrared image coordinate system to visible light image coordinate system.

4. The method for real-time recognition of road surface defects based on multimodal fusion of visible light and infrared images according to claim 3 is characterized by: The mapping transformation of the image coordinates of the infrared image into the visible light image reference field and the dynamic pixel registration of the visible light image is as follows: Among them, t x is the horizontal displacement, t y is the vertical displacement, is the rotation angle, s is the scaling factor; x vi ,y vi Represents the source pixel coordinates of the visible light image.

5. A method for real-time recognition of road surface defects based on multimodal fusion of visible light and infrared images according to any one of claims 1 to 4, characterized in that: The edge filtering processing formula is as follows: Y=WLSR(X)=(1+λL X ) -1 X, where Among them, X represents the input image before filtering, Y represents the edge filtering result; λ is the balance coefficient; A x ,A y is a diagonal weight matrix, and the diagonal elements of the matrix are μ x,p , μ y,p , μ x,p , μ y,p is the weight coefficient of the x,y direction gradient corresponding to the pixel of image X at spatial position p; M x ,M y is the discrete difference operator matrix corresponding to the x-direction difference and the y-direction difference.

6. The method for real-time recognition of road surface defects based on multimodal fusion of visible light and infrared images according to claim 5 is characterized by: Diagonal weight matrix A x ,A y The details are as follows: A x =diag(μ x,1 ,m x,2 ,m x,3 ,m x,4 ,…,m x,n ) A y =diag(μ y,1 ,m y,2 ,m y,3 ,m y,4 ,…,m y,n ) Here, n represents the total number of pixels in the image.

7. The method for real-time recognition of road surface defects based on multimodal fusion of visible light and infrared images according to claim 5 is characterized by: The discrete difference operator matrix corresponding to the x-direction difference Discrete difference operator matrix corresponding to the difference in the y direction 8. The method for real-time recognition of road surface defects based on multimodal fusion of visible light and infrared images according to claim 5 is characterized by: Base layer B based on visible light image vi and the base layer B of the infrared image iR Get the fusion information F B The process includes: For B vi and B iR , use I k Represents the intensity value of the kth pixel, using the significant intensity difference distribution formula Obtain the saliency intensity value Λ of visible light and infrared images vi ,Λ iR , where N is B x The number of pixels in B x Indicates B vi or B iR ; Then calculate the fusion adaptive weight Then we get the fusion information F of the base layer B =ωB vi +(1-ω)B iR .

9. The method for real-time recognition of road surface defects based on multimodal fusion of visible light and infrared images according to claim 5, characterized in that: Detail layer D based on visible light image vi and the detail layer D of the infrared image iR Get the fusion information F D The process includes: First, the detail layer D of the visible light image vi Perform SVF multi-scale decomposition to obtain the enhanced detail layer D E_vi : in, Represents the SVF filtering results of different scales, Indicates that it corresponds to The detail layer, α, β are the detail layer weight coefficients, D E_vi represents an enhanced visible light image detail layer; Detail layer D of infrared image iR The enhanced infrared image detail layer D is obtained in the same way. E_iR ; Then we get the fusion information F of the detail layer D =D E_vi +D E_iR .

10. The method for real-time recognition of road surface defects based on multimodal fusion of visible light and infrared images according to claim 5, characterized in that: The neural network model used in the process of realizing real-time identification of road surface defects using a neural network model is the YOLOv8 network model.

Citation Information

Patent Citations

  • Multi-band image fusion method and system combined with saliency region detection

    CN115578304A

  • Infrared and visible light image fusion enhanced vehicle detection method and system

    CN116152778A

  • Multi-mode image fusion method and system for distribution network components

    CN116739955A

  • Infrared and visible light image fusion method, device and system and storage medium

    CN117152037A

  • Visible light and infrared light image fusion method and device for road crack detection

    CN119067867A

Cited By

  • Boiler scale detection method and device based on multi-modal image

    CN120385686A

  • A boiler scale detection method and device based on multimodal images

    CN120385686B

  • Structure surface disease diagnosis method and system based on multi-modal edge calculation

    CN121068609A

  • Foreign matter detection method, system and device in railway scene and storage medium

    CN122157210A

  • A foreign matter detection method, system, device and storage medium in a railway scene

    CN122157210B