Method and device for predicting trajectory of dangerous rock mass based on fusion of visible light and infrared thermal imaging

Through infrared thermal imaging and visible light imaging fusion technology, combined with video fusion model and dangerous rock mass collapse recognition model, the monitoring problem of dangerous rock mass collapse is solved at night and inclement weather, and accurate rockfall recognition and trajectory prediction are achieved, which improves monitoring efficiency and accuracy.

CN119832028BActive Publication Date: 2025-07-08CHANGAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510300603.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-07-08
Estimated Expiration
2045-03-14

AI Technical Summary

Technical Problem

Existing visible light monitoring equipment is difficult to accurately distinguish rockfall from background at night or in severe weather, resulting in inaccurate identification and trajectory prediction, and it is impossible to effectively monitor the collapse of dangerous rocks on the road slopes.

Method used

Using a method based on the fusion of visible light and infrared thermal imaging, the fusion processing of infrared thermal imaging video and visible light video is combined with the trained video fusion model and the dangerous rock mass collapse recognition model to achieve accurate prediction of dangerous rock mass trajectory.

Benefits of technology

It has achieved all-weather and full-time monitoring of dangerous rock mass collapse and slippage, improved the accuracy of rockfall location and trajectory prediction, reduced manpower and material consumption, and reduced maintenance costs and operation uncertainty of mountain traffic lines.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119832028B_ABST
    Figure CN119832028B_ABST
Patent Text Reader

Abstract

The present invention provides a method and device for predicting the trajectory of dangerous rock masses based on the fusion of visible light and infrared thermal imaging, belonging to the field of computer technology, including: obtaining the infrared thermal imaging video and visible light video of the dangerous rock mass area in the first time period, and obtaining a fused video based on the infrared thermal imaging video and the visible light video; inputting the fused video into a trained dangerous rock mass landslide recognition model, determining the falling rock position data corresponding to the current image frame based on the current image frame and the previous image frame in the fused video, obtaining a first feature vector based on the current image frame, performing an embedding operation on the first feature vector to obtain an embedding vector, inputting the embedding vector into a downsampling module to obtain an image feature vector corresponding to the current image frame, and inputting the image feature vector corresponding to the current image frame and the falling rock position data into an upsampling module to predict the falling rock position data corresponding to the next image frame. This method can accurately identify falling rocks and improve the prediction accuracy of the falling rock trajectory.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular, to a method and device for predicting the trajectory of dangerous rock masses based on the fusion of visible light and infrared thermal imaging. Background Art

[0002] The occurrence of highway slope dangerous rock collapses often has the characteristics of strong suddenness and great destructiveness, which can block traffic in a short time, causing huge economic losses and potential safety hazards to personnel. In addition, geological disasters also increase the difficulty and cost of maintaining mountainous area traffic lines. Frequent repair work not only consumes a large amount of manpower and material resources, but also increases the uncertainty of traffic operation. Therefore, frequent geological disasters have become one of the key factors restricting the safe and efficient operation of mountainous area traffic. The dangerous rocks on the slopes of transportation corridors are characterized by numerous, scattered, and small individual scales. If traditional methods are used for monitoring, it will consume a large amount of manpower, material resources, and financial resources, and there are also many problems in actual operation. For dangerous rock masses with potential hazards, monitoring them in a simple, effective, and intelligent way will be the key to effectively preventing and avoiding dangerous rock mass landslide disasters.

[0003] In the prior art, most use visible light monitoring devices to monitor the video of dangerous rock areas. The visible light sensor forms an image through light reflection and provides high-resolution scene information. However, it is very difficult for the visible light sensor to distinguish obvious targets from the background environment at night or in bad weather. For example, when using visible light monitoring devices to monitor the falling rocks on the rock slopes of highways, it is impossible to accurately distinguish the falling rocks from the background throughout the day, and the sensitivity to the recognition of the movement trajectory of the falling rocks is relatively low, resulting in technical problems such as failure to recognize the falling rocks and inaccurate predicted trajectories of the falling rocks. Therefore, how to accurately identify the falling rocks and improve the prediction accuracy of the falling rock trajectories is a technical problem to be solved urgently. Summary of the Invention

[0004] In view of this, the embodiments of the present invention provide a method and device for predicting the trajectory of dangerous rock masses based on the fusion of visible light and infrared thermal imaging to eliminate or improve one or more defects existing in the prior art.

[0005] One aspect of the present invention provides a method for predicting the trajectory of dangerous rock masses based on the fusion of visible light and infrared thermal imaging, the method comprising:

[0006] Obtain the infrared thermal imaging video and visible light video of the dangerous rock area in the first time period, and obtain the fused video based on the infrared thermal imaging video and visible light video;

[0007] Input the fused video into the trained dangerous rock mass landslide recognition model. The dangerous rock mass landslide recognition model determines the falling rock position data corresponding to the current image frame based on the current image frame and the previous image frame in the fused video, obtains a first feature vector based on the current image frame, performs an embedding operation on the first feature vector to obtain an embedding vector, inputs the embedding vector into the downsampling module of the dangerous rock mass landslide recognition model to obtain the image feature vector corresponding to the current image frame, and inputs the image feature vector corresponding to the current image frame and the falling rock position data into the upsampling module to predict the falling rock position data corresponding to the next image frame. The falling rock position data includes the vertex coordinates of the detection box corresponding to the falling rock area.

[0008] In some embodiments of the present invention,

[0009] Obtaining a fused video based on the infrared thermal imaging video and the visible light video includes:

[0010] Input the infrared thermal imaging video and the visible light video into the trained first video fusion model. The first video fusion model splices the image frames of the infrared thermal imaging video and the image frames of the visible light video to obtain a spliced image, and sequentially passes through a convolutional layer, a downsampling layer, a first dilated convolutional layer, an upsampling layer, a transposed convolutional layer, and a second dilated convolutional layer to obtain the fused image corresponding to the spliced image;

[0011] Splice the fused images corresponding to the respective spliced images to obtain a fused video.

[0012] In some embodiments of the present invention,

[0013] Obtaining a fused video based on the infrared thermal imaging video and the visible light video includes:

[0014] Determine the brightness values of the image frames in the visible light video, calculate a first weight value of the infrared thermal imaging video and a second weight value of the visible light video based on the respective brightness values, calculate a normalized weight coefficient based on the first weight value and the second weight value, and calculate a fused video based on the normalized weight coefficient, the infrared thermal imaging video, and the visible light video.

[0015] In some embodiments of the present invention, the calculation formula for the brightness value is:

[0016] ;

[0017] where l(t) represents the brightness value of the t-th image frame in the visible light video, r(t) represents the R-channel data of the t-th image frame in the visible light video, g(t) represents the G-channel data of the t-th image frame in the visible light video, and b(t) represents the B-channel data of the t-th image frame in the visible light video;

[0018] The calculation formula for the first weight value is: , where l min represents the minimum brightness value, and l max represents the maximum brightness value;

[0019] The calculation formula for the second weight value is: , where is the second weight value, is the first weight value;

[0020] The calculation formula for the normalized weight coefficient is: , , , where α represents the normalized weight coefficient, represents the calculation result of the gradient amplitude corresponding to the image frame in the visible light video, represents the normalization result of the image matrix corresponding to the image frame in the infrared thermal imaging video.

[0021] In some embodiments of the present invention, the fused video is calculated based on the normalized weight coefficient, the infrared thermal imaging video, and the visible light video, including:

[0022] Based on calculate the R-channel data of the fused image corresponding to the image frame;

[0023] Based on calculate the G-channel data of the fused image corresponding to the image frame;

[0024] Based on calculate the B-channel data of the fused image corresponding to the image frame;

[0025] where A r , A g and A b respectively represent the R-channel data, G-channel data, and B-channel data of the image frame in the visible light video, B r , B g and B b respectively represent the R-channel data, G-channel data, and B-channel data of the image frame in the infrared thermal imaging video, and α represents the normalized weight coefficient.

[0026] In some embodiments of the present invention, the method further includes:

[0027] Construct a first sample data set, a first loss function, and an initial first video fusion model. Based on the first sample data set and the first loss function, pre-train the initial first video fusion model to obtain a trained first video fusion model. The sample data in the first sample data set includes infrared thermal imaging video sample data, visible light video sample data, and fused video sample data;

[0028] The first loss function is: , where N represents the total number of sample data in the first sample data set, L i represents the predicted value of the fused image, represents the label value of the fused image.

[0029] In some embodiments of the present invention, obtaining a fused video based on the infrared thermal imaging video and the visible light video includes:

[0030] Performing image registration on the image frames in the infrared thermal imaging video and the image frames in the visible light video; and / or,

[0031] Obtaining an infrared thermal imaging video and a visible light video of a dangerous rock area in a first time period includes:

[0032] Based on an infrared thermal imaging device and a visible light device having a preset distance from the dangerous rock area, respectively capture the infrared thermal imaging video and the visible light video of the dangerous rock area. The infrared thermal imaging device and the visible light device are both fixed on a device support rod.

[0033] In some embodiments of the present invention, the method includes:

[0034] Construct a second sample data set, a second loss function, and an initial dangerous rock mass landslide identification model. Based on the second sample data set and the second loss function, pre-train the initial dangerous rock mass landslide identification model to obtain a trained dangerous rock mass landslide identification model. The sample data in the second sample data set includes fused video sample data and sample data of the falling rock position data corresponding to the next image frame;

[0035] The second loss function is: , where, , ρ(p, g) represents the distance between the predicted value of the falling rock center point position data and the label value of the falling rock center point position data, Q represents the predicted value of the falling rock position data, R represents the label value of the falling rock position data, p represents the predicted value of the falling rock center point position data, g represents the label value of the falling rock center point position data, represents the intersection over union loss, and Loss represents the second loss function.

[0036] According to another aspect of the present invention, there is also disclosed a device for predicting the trajectory of dangerous rock masses based on the fusion of visible light and infrared thermal imaging. The device includes a processor, a memory, and a computer program stored on the memory. The processor is configured to execute the computer program, and when the computer program is executed, the device implements the steps of the method described in any of the above embodiments.

[0037] According to still another aspect of the present invention, there is also disclosed a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, it implements the steps of the method described in any of the above embodiments.

[0038] The method for predicting the trajectory of dangerous rock masses disclosed in the above embodiments of the present invention combines infrared thermal imaging and visible light imaging technologies, breaking through the limitations of single visible light monitoring under night or bad weather conditions, and enabling all-weather and full-time monitoring. Moreover, the fusion video of the present application further realizes a more comprehensive and detailed prediction of the position of falling rocks through a dangerous rock mass landslide and collapse recognition model, and can efficiently monitor the initiation of dangerous rock mass landslides and the movement trajectories of falling rocks. Therefore, the method for predicting the trajectory of dangerous rock masses can not only accurately identify falling rocks, but also improve the prediction accuracy of the trajectories of falling rocks.

[0039] In addition, the method for predicting the trajectory of dangerous rock masses in the present application adopts intelligent monitoring means, which not only improves the recognition ability of the model, but also greatly reduces the consumption of manpower, material resources, and financial resources, while reducing the maintenance cost and operation uncertainty of mountainous traffic lines.

[0040] The additional advantages, objectives, and features of the present invention will be partially described below, and will become partially apparent to those of ordinary skill in the art after studying the following text, or may be learned from the practice of the present invention. The objectives and other advantages of the present invention can be achieved and obtained by the structures specifically pointed out in the specification and the drawings.

[0041] Those skilled in the art will understand that the objectives and advantages that can be achieved by the present invention are not limited to the above specifically described, and the above and other objectives that the present invention can achieve will be more clearly understood according to the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] The drawings described herein are used to provide a further understanding of the present invention, form a part of this application, and do not limit the present invention. The components in the drawings are not drawn to scale, but are only for showing the principles of the present invention. For the convenience of showing and describing some parts of the present invention, the corresponding parts in the drawings may be enlarged, that is, they may become larger relative to other components in the exemplary device actually manufactured according to the present invention. In the drawings:

[0043] Figure 1Schematic flowchart of a method for predicting the trajectory of dangerous rock masses based on the fusion of visible light and infrared thermal imaging according to an embodiment of the present application.

[0044] Figure 2 Schematic flowchart of a method for predicting the trajectory of dangerous rock masses based on the fusion of visible light and infrared thermal imaging according to another embodiment of the present application.

[0045] Figure 3 Schematic layout diagram of dangerous rock mass landslide monitoring equipment according to an embodiment of the present application.

[0046] Figure 4 Schematic architecture diagram of a first video fusion model according to an embodiment of the present application.

[0047] Figure 5 Schematic architecture diagram of a dangerous rock mass landslide identification model according to an embodiment of the present application.

[0048] Figure 6 Schematic composition diagram of an electronic device according to an embodiment of the present application.

[0049] Reference numerals:

[0050] Acousto-optic alarm 1, Solar energy 2, Infrared thermal imaging device 3, Visible light device 4, Storage battery 5, Video preliminary processing device 6, Electronic device 300, Processor 310, Memory 320. Detailed implementation manners

[0051] To make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with the embodiments and the drawings. Herein, the illustrative embodiments and descriptions thereof of the present invention are used to explain the present invention, but not to limit the present invention.

[0052] Herein, it should also be noted that in order to avoid obscuring the present invention due to unnecessary details, only the structures and / or processing steps closely related to the solution of the present invention are shown in the drawings, while other details less related to the present invention are omitted.

[0053] It should be emphasized that the term "including / comprising" when used herein refers to the presence of features, elements, steps or components, but does not exclude the presence or addition of one or more other features, elements, steps or components.

[0054] Herein, it should also be noted that if not otherwise specified, the term "connection" in this document can not only refer to a direct connection, but also represent an indirect connection with an intermediate, and can not only represent a wired connection, but also a wireless connection, which can be specifically changed based on the actual application scenario.

[0055] For a better understanding of the present invention, the following are the explanations of the terms related to the technical solution:

[0056] The infrared thermal imaging monitoring device passive receives the infrared thermal radiation of the target itself and is independent of climatic conditions, and can work normally day and night. The core of the infrared and visible light image fusion is to integrate the complementary information of the two source images respectively to generate a fused image with better visual effects; in the field of intelligent recognition of rockfall trajectories, the visible light and infrared thermal imaging fusion technology provides more comprehensive information for the machine learning model, improves the accuracy of rockfall recognition, and shows good adaptability.

[0057] In the following, embodiments of the present invention will be described with reference to the accompanying drawings. In the drawings, the same reference numerals represent the same or similar components, or the same or similar steps.

[0058] Figure 1 The flowchart of the dangerous rock mass trajectory prediction method based on the fusion of visible light and infrared thermal imaging according to an embodiment of the present application is shown as Figure 1 shown, and the dangerous rock mass trajectory prediction method at least includes steps S10 to S20.

[0059] Step S10: Obtain the infrared thermal imaging video and visible light video of the dangerous rock mass area in the first time period, and obtain a fused video based on the infrared thermal imaging video and visible light video.

[0060] In this step, the infrared thermal imaging video and visible light video of the dangerous rock mass area can be respectively captured by an infrared thermal imaging device and a visible light device having a preset distance from the dangerous rock mass area. Both the infrared thermal imaging device 3 and the visible light device 4 are fixed on the device support rod. Specifically, first, install the dangerous rock mass landslide monitoring device (as Figure 3 shown), ensure that the entire movement path of the dangerous rock mass landslide is within the field of view of the infrared thermal imaging device and the visible light device; exemplarily, the support rod can be fixed at a certain distance from the slope. The support rod can specifically be an iron vertical rod, and both the infrared thermal imaging device and the visible light device are fixed in the upper middle part of the iron vertical rod; in addition, the support rod can specifically be located in the area 30m - 50m away from the slope. In order to ensure the reliable monitoring of the infrared thermal imaging device and the visible light device, the infrared thermal imaging device and the visible light device can be powered by solar energy 2 (360w / 60A). Further, a waterproof box can be installed in the middle of the support rod, and a storage battery is placed in the waterproof box. The storage battery stores the excess electric energy and provides electric energy for the entire system when the solar energy 2 does not generate electricity.

[0061] In one embodiment, a fused video is obtained based on the infrared thermal imaging video and the visible light video, including: inputting the infrared thermal imaging video and the visible light video into a trained first video fusion model, and the first video fusion model splices the image frames of the infrared thermal imaging video and the image frames of the visible light video to obtain a spliced image, and sequentially passes through a convolutional layer, a downsampling layer, a first dilated convolutional layer, an upsampling layer, a transposed convolutional layer, and a second dilated convolutional layer to obtain a fused image corresponding to the spliced image; splicing the fused images corresponding to the respective spliced images to obtain a fused video. In this embodiment, the fusion of the infrared thermal imaging video and the visible light video is completed based on the trained first video fusion model. In order to obtain the trained first video fusion model, in some embodiments, the dangerous rock mass trajectory prediction method further includes the following steps: constructing a first sample data set, a first loss function, and an initial first video fusion model, and pre-training the initial first video fusion model based on the first sample data set and the first loss function to obtain the trained first video fusion model, and the sample data in the first sample data set includes infrared thermal imaging video sample data, visible light video sample data, and fused video sample data; the first loss function is: , where N represents the total number of sample data in the first sample data set, L i represents the predicted value of the fused image, represents the label value of the fused image.

[0062] In order to obtain the sample data in the first sample data set, dangerous rock videos can be simulated from slopes with landslide risks to obtain dangerous rock video sample data. Exemplarily, dangerous rock mass landslide phenomena can be simulated on the slope, and an infrared thermal imaging device and a visible light device are used to capture videos in real time. The camera can be an auto-focus camera; further, the captured video data is transmitted to a video recorder and uploaded to a remote computer through a 4G router and a switch, and a dangerous rock mass landslide video database is established in the computer. When using the infrared thermal imaging device and the visible light device to capture dangerous rock videos, the infrared thermal imaging device and the visible light device are kept on the same horizontal plane, and the distance between the infrared thermal imaging device and the visible light device can be set to 20 cm. Further, the distance between the detection device and the object to be measured (such as 20 m, 40 m, 60 m, etc.) is adjusted to obtain infrared thermal imaging video sample data and visible light video sample data.

[0063] In addition, when obtaining the fused video based on the infrared thermal imaging video and the visible light video, the image frames in the infrared thermal imaging video and the image frames in the visible light video can also be image-registered. Specifically, images are extracted frame by frame from the obtained visible light video and infrared thermal imaging video, and the thermal infrared image and the visible light image are initially registered. The registration algorithm can be a feature-based method (such as SIFT, RIFT, etc.) or a region-based method (such as CFOG, HOPC, etc.). After the initial registration, precise registration is performed. After registration, the shooting areas of the image frames in the infrared thermal imaging video and the image frames in the visible light video are made consistent. In addition, the image frames are fine-tuned to reduce image distortion, and thermal infrared image frames and visible light image frames with the same region and resolution are obtained. The thermal infrared image frames provide temperature distribution, and the visible light image frames provide texture and details. Then, the registered thermal infrared image frames and visible light image frames are fused to obtain the fused image.

[0064] In another embodiment, the fused video can also be obtained based on the following steps: Determine the brightness values of the image frames in the visible light video, calculate the first weight value of the infrared thermal imaging video and the second weight value of the visible light video based on the brightness values, calculate the normalized weight coefficient based on the first weight value and the second weight value, and calculate the fused video based on the normalized weight coefficient, the infrared thermal imaging video, and the visible light video. In one embodiment, the above steps can be used to determine the fused video sample data based on the infrared thermal imaging video sample data and the visible light video sample data.

[0065] Further, the calculation formulas for the brightness value, the first weight value, the second weight value, and the normalized weight coefficient are specifically as follows:

[0066] ;

[0067] ;

[0068] ;

[0069] ;

[0070] ;

[0071] ;

[0072] Among them, l(t) represents the brightness value of the t-th image frame in the visible light video, r(t) represents the R-channel data of the t-th image frame in the visible light video, g(t) represents the G-channel data of the t-th image frame in the visible light video, b(t) represents the B-channel data of the t-th image frame in the visible light video; l min represents the minimum brightness value (when there are multiple sample data, lmin represents the average value of the minimum visible light intensity of multiple sample data in a day), l max represents the maximum brightness value (when there are multiple sample data, l max represents the average value of the maximum visible light intensity of multiple sample data in a day); is the second weight value (the weight value corresponding to the infrared thermal imaging video), is the first weight value (the weight value corresponding to the visible light video); represents the feature importance corresponding to the visible light video, represents the feature importance corresponding to the infrared thermal imaging video; α represents the normalized weight coefficient, represents the calculation result of the gradient amplitude corresponding to the image frame in the visible light video (that is, the gradient result in [0, 1] obtained by using edge algorithms such as the Canny operator and Sobel operator for the image frame in the visible light video), represents the normalization result of the image matrix corresponding to the image frame in the infrared thermal imaging video (that is, the normalization result obtained by normalizing the temperature matrix exported from the image frame in the infrared thermal imaging video). Additionally, when determining the feature importance corresponding to the infrared thermal imaging video, the position detection frame of the rockfall can be marked in the visible light image as the target area, and the temperature difference matrix between the local temperature of the target area and the surrounding background temperature is calculated, and the temperature difference matrix is normalized.

[0073] In some embodiments, a fused video is calculated based on the normalized weight coefficient, the infrared thermal imaging video, and the visible light video. Specifically, it may include: Based on calculate the R-channel data of the fused image corresponding to the image frame; Based on calculate the G-channel data of the fused image corresponding to the image frame; Based on calculate the B-channel data of the fused image corresponding to the image frame. Among them, A r 、A g and A b respectively represent the R-channel data, G-channel data, and B-channel data of the image frame in the visible light video, and B r 、B g and B b respectively represent the R-channel data, G-channel data, and B-channel data of the image frame in the infrared thermal imaging video. In the above embodiments, determining the weight value corresponding to the image frame of the visible light video based on the brightness value of the image frame of the visible light video can make the weight of the image frame of the visible light video higher during the day, while the weight of the image frame of the infrared thermal imaging video is higher at night.

[0074] Step S20: Input the fused video into the trained dangerous rock mass landslide recognition model. The dangerous rock mass landslide recognition model determines the falling rock position data corresponding to the current image frame based on the current image frame and the previous image frame in the fused video, obtains a first feature vector based on the current image frame, performs an embedding operation on the first feature vector to obtain an embedding vector, inputs the embedding vector into the downsampling module of the dangerous rock mass landslide recognition model to obtain the image feature vector corresponding to the current image frame, and inputs the image feature vector corresponding to the current image frame and the falling rock position data into the upsampling module to predict the falling rock position data corresponding to the next image frame. The falling rock position data includes the vertex coordinates of the detection box corresponding to the falling rock area.

[0075] In this step, the falling rock position data corresponding to the next image frame is predicted based on the trained dangerous rock mass landslide recognition model. The input information of the dangerous rock mass landslide recognition model is the fused video obtained in step S10, and the output information is the falling rock position data corresponding to the next image frame. When specifically predicting, the input of the dangerous rock mass landslide recognition model is the video at the current moment and before the current moment, and the output is the falling rock position data at the next moment.

[0076] Exemplarily, in order to obtain the prediction of the trained dangerous rock mass landslide recognition model, the dangerous rock mass trajectory prediction method may further include the following steps: constructing a second sample data set, a second loss function, and an initial dangerous rock mass landslide recognition model, and pre-training the initial dangerous rock mass landslide recognition model based on the second sample data set and the second loss function to obtain the trained dangerous rock mass landslide recognition model. The sample data in the second sample data set includes fused video sample data and falling rock position data sample data corresponding to the next image frame.

[0077] Furthermore, the fused video sample data may be the fused video sample data in the first sample data set. When constructing the second sample data set, falling rock detection boxes are labeled for each image frame in each fused video sample data, that is, each image frame is exported from the fused video, and the falling rock position detection box is labeled in each image frame as the target area. The four vertex coordinates of the detection box are used as the label, and this label is the falling rock position data in the corresponding image frame. The dangerous rock mass landslide recognition data set data is determined based on the falling rock position data in each image frame. , , denotes the fused video corresponding dangerous rock mass landslide recognition data set, denotes the i-th image frame exported from the fused video simulating the dangerous rock mass landslide The four vertex coordinates of the falling rock detection box are combined to form a 4×2 matrix , and the arrangement order of the four vertex coordinates in the matrix is bottom left, top left, top right, bottom right. , where \(n\) is the total number of image frames; export a dataset from all the fused videos , to form a dataset data for identifying the collapse and landslide of dangerous rock masses, , where \(m\) is the total number of fused videos. In this embodiment, the prediction of the initial dangerous rock mass collapse and landslide identification model is pre-trained based on the dataset data for identifying the collapse and landslide of dangerous rock masses.

[0078] Figure 5 is a schematic diagram of the architecture of the dangerous rock mass collapse and landslide identification model according to an embodiment of the present application. This model is improved by a transformer and combines the encoding and decoding (encode and decode) forms. The model starts reasoning through the start token and the landslide start \(x, y, w, h\) detected by the optical flow algorithm. The generation of each token will refer to the previously generated tokens, and finally \([x, y, w, h, end]\) is generated to inform the user that the prediction is completed. As Figure 5 shown, when pre-training this model, the fused video can be first analyzed for images, that is, existing dynamic object detection algorithms can be used for dynamic object detection; for example, the start of the collapse and landslide of dangerous rock masses can be predicted by analyzing the optical flow field, and the rough range of the collapse and landslide of dangerous rock masses can be circled with an initial detection frame. The four vertex coordinates of the initial detection frame are combined to form a \(4\times2\) matrix , . Further, the corresponding of the fused video is extracted from the dataset data as a sequence set, and the image frame where the initial detection frame is located is , and the next image frame of is . In addition, , are used as model inputs, , are imported into the linear projection layer through image slicing to output feature vectors and feature vector , and , are respectively combined with the position embedding vectors and , and then the image features are extracted through the downsampling module. The output image features are further imported into the upsampling module. At the same time, plus the start entry (such as start) is imported into the upsampling module through the text encoder. The vector output by the upsampling module is then output through the text decoder as the predicted position of the falling rock .

[0079] During the training process, the loss function Loss is used to calculate the predicted bounding boxes and the labels for the differences

[0080] ;

[0081] ;

[0082] Among them, ρ(p, g) represents the distance between the predicted value of the rockfall center point position data and the label value of the rockfall center point position data, Q represents the predicted value of the rockfall position data, R represents the label value of the rockfall position data, p represents the predicted value of the rockfall center point position data, and g represents the label value of the rockfall center point position data represents the intersection over union loss, and Loss represents the second loss function

[0083] For the dangerous rock mass collapse and landslide recognition model, a machine learning optimizer of the prior art, such as the adam optimizer, can be selected, and the dangerous rock mass collapse and landslide recognition model is trained through the loss function and the optimizer until it converges

[0084] In addition Figure 4 The following is a schematic diagram of the architecture of the first video fusion model according to an embodiment of the present application. For example, based on this first video fusion model, a visible light video and an infrared thermal imaging video are used to generate a fused video, that is, a visible light image D1 and a thermal infrared image D2 are used to generate a fused image, and finally multiple fused images are stitched together to obtain a fused video; in this embodiment, based on an improved UNet network, the infrared thermal imaging video and the visible light video are fused to obtain a more comprehensive and detailed fused video. When generating the fused image, the visible light image D1 and the thermal infrared image D2 are stitched into 6-channel data as the model input, and after a convolution operation with a convolution kernel of 7×7, a stride of 2, and a number of channels of 64, the output is a tensor of 256×256×64; then the tensor of 256×256×64 is further used to extract image features through a downsampling module. The downsampling module is four pooling layers and a resnet network, and the output tensor size is 16×16×512; the output of the downsampling module is further passed through four dilated convolutions, and the outputs of the dilated convolutions in five stages are added together, and the output is a tensor of 16×16×512; the output of the dilated convolution is further upsampled four times to obtain a tensor of 256×256×64. At the same time, the output of each downsampling feature is added to the input of each upsampling through a self-attention module; finally, the output of the model is obtained through a 4×4 transposed convolution operation and a convolution operation with a dilation rate of 1, that is, a three-channel fused image

[0085] In addition, when training the first video fusion model, the adam optimizer can be selected, and the deep network is trained through the loss function and the optimizer until convergence. In the first video fusion model, each frame of the infrared thermal imaging video and the visible light video is used as input information, and each frame of the fused image is output and re - spliced into a video, finally generating a more comprehensive and detailed fused video.

[0086] Figure 3 The layout schematic diagram of the dangerous rock collapse and landslide monitoring device according to an embodiment of the present application is as follows Figure 3 As shown, a waterproof box can be further arranged in the middle and lower part of the equipment support rod of the dangerous rock collapse and landslide monitoring device, and a video recorder, a hard disk, a 4G router, a switch, etc. are located in the waterproof box. The video recorder, the hard disk, the 4G router, the switch, etc. are combined into a video preliminary processing device 6 to provide local area network and cloud support for the whole system, facilitating users to view online. In addition, an audible and visual alarm 1 can be further arranged at the top of the support rod to remind pedestrians and vehicles of falling rocks when there is a risk of collapse of the slope dangerous rock. That is, in some embodiments, based on the dangerous rock trajectory prediction method, the slope dangerous rock collapse and landslide video is monitored, and alarm measures are taken when the start of the dangerous rock collapse and landslide is detected. The alarm measures include audible and visual alarms at points 1000m before and after the highway line, audible and visual alarms on the monitoring device, and text message notifications to users, etc. In addition, when identifying the movement trajectory of the dangerous rock in the fused video, if the start of the dangerous rock collapse and landslide is detected, the detected frames of the dangerous rock collapse detection boxes identified in each image frame are exported, and the mid - points of the detection boxes are connected in sequence. The connection line of the mid - points is the trajectory of the dangerous rock collapse and landslide.

[0087] Correspondingly, the present invention also provides a dangerous rock trajectory prediction device based on the fusion of visible light and infrared thermal imaging. The device includes a processor, a memory, and a computer program stored on the memory. The processor is used to execute the computer program, and when the computer program is executed, the device realizes the steps of the method described in any of the above embodiments.

[0088] Figure 2 The flow schematic diagram of the dangerous rock trajectory prediction method based on the fusion of visible light and infrared thermal imaging according to another embodiment of the present application is as follows Figure 2 As shown, the dangerous rock trajectory prediction method specifically includes the following steps:

[0089] Step S110, set up the dangerous rock collapse and landslide monitoring device. Ensure that the entire movement path of the dangerous rock collapse and landslide is within the field of view of the infrared thermal imaging device and the visible light device.

[0090] Step S120, simulate the dangerous rock collapse and landslide, and shoot a video. Simulate rockfall on the field slope and shoot a video of the slope with potential collapse and landslide danger in real time.

[0091] Step S130: Fuse the thermal imaging and visible light videos using machine learning methods. The infrared thermal imaging video and the visible light video are fused to obtain a more comprehensive and detailed fused video.

[0092] Step S140: Construct a video dataset for dangerous rock mass collapse and landslide. Mark the rockfall detection frames to construct a video dataset for dangerous rock mass collapse and landslide.

[0093] Step S150: Train a recognition model for dangerous rock mass collapse and landslide. Specifically, use the labeled video dataset of dangerous rock mass collapse and landslide to train the recognition model for dangerous rock mass collapse and landslide.

[0094] Step S160: Conduct recognition detection and alarm through the recognition model for dangerous rock mass collapse and landslide. Monitor the video of dangerous rock mass collapse and landslide on the slope. When the phenomenon of dangerous rock mass collapse and landslide is detected, take alarm measures and identify the collapse path.

[0095] According to the above embodiments, it can be found that the method for predicting the trajectory of dangerous rock masses based on the fusion of visible light and infrared thermal imaging in this application simulates rockfalls on slopes, captures visible light and infrared thermal imaging videos, and can achieve seamless monitoring during the day and night after fusion, overcoming the limitations of single imaging technologies. Mark the positions of dangerous rock mass collapses in each image frame of the simulated video, and use machine learning to learn these marks to obtain a trajectory recognition model for dangerous rock mass collapses, realizing the automatic recognition of the trajectories of dangerous rock mass collapses on field slopes, which is fast and efficient. This method comprehensively applies infrared thermal imaging, visible light imaging, and machine learning technologies, laying a foundation for the study of the movement characteristics of dangerous rock mass collapses on slopes and providing reliable technical support for the protection of transportation structures and geological disaster early warning.

[0096] An embodiment of the present invention also provides an electronic device 300. Figure 6 It is a schematic diagram of the composition of the electronic device according to an embodiment of the present application. As Figure 6 shown, the electronic device includes a processor 310 and a memory 320.

[0097] An embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it realizes the steps of the method described in any of the above embodiments. The computer-readable storage medium may be a tangible storage medium, such as a random access memory (RAM), memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, register, floppy disk, hard disk, removable storage disk, CD-ROM, or any other form of storage medium well-known in the technical field.

[0098] Those of ordinary skill in the art should understand that the various exemplary components, systems, and methods described in connection with the embodiments disclosed herein can be implemented in hardware, software, or a combination of both. Specifically, whether to implement in hardware or software depends on the specific application and design constraints of the technical solution. A person skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention. When implemented in hardware, it can be, for example, an electronic circuit, an application-specific integrated circuit (ASIC), appropriate firmware, a plug-in, a functional card, and so on. When implemented in software, the elements of the present invention are programs or code segments used to perform the required tasks. The program or code segment can be stored in a machine-readable medium or transmitted through a data signal carried in a carrier wave on a transmission medium or a communication link.

[0099] It should be clear that the present invention is not limited to the specific configurations and processes described above and illustrated in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and illustrated as examples. However, the method process of the present invention is not limited to the specific steps described and illustrated. Those skilled in the art can make various changes, modifications, and additions, or change the order between steps after understanding the spirit of the present invention.

[0100] In the present invention, the features described and / or exemplified for one embodiment can be used in the same or a similar manner in one or more other embodiments, and / or combined with the features of other embodiments or replace the features of other embodiments.

[0101] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, various changes and modifications can be made to the embodiments of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for predicting the trajectory of dangerous rock masses based on the fusion of visible light and infrared thermal imaging, characterized in that, The method includes: Obtain the infrared thermal imaging video and visible light video of the dangerous rock area in the first time period, and obtain a fused video based on the infrared thermal imaging video and visible light video; Input the fused video into a trained dangerous rock landslide recognition model. The dangerous rock landslide recognition model determines the rockfall position data corresponding to the current image frame based on the current image frame and the previous image frame in the fused video, and obtains a first feature vector based on the current image frame. Perform an embedding operation on the first feature vector to obtain an embedded vector, input the embedded vector into the downsampling module of the dangerous rock landslide recognition model to obtain the image feature vector corresponding to the current image frame, and input the image feature vector corresponding to the current image frame and the rockfall position data into the upsampling module to predict the rockfall position data corresponding to the next image frame. The rockfall position data includes the vertex coordinates of the detection box corresponding to the rockfall area; Obtaining a fused video based on the infrared thermal imaging video and visible light video includes: Determine the brightness values of each image frame in the visible light video, calculate the first weight value of the infrared thermal imaging video and the second weight value of the visible light video based on the brightness values, calculate the normalized weight coefficient based on the first weight value and the second weight value, and calculate the fused video based on the normalized weight coefficient, the infrared thermal imaging video and the visible light video; The calculation formula for the brightness value is: ; where l(t) represents the brightness value of the t-th image frame in the visible light video, r(t) represents the R-channel data of the t-th image frame in the visible light video, g(t) represents the G-channel data of the t-th image frame in the visible light video, and b(t) represents the B-channel data of the t-th image frame in the visible light video; The calculation formula for the first weight value is: , where l min represents the minimum luminance value, and l max represents the maximum luminance value; The calculation formula for the second weight value is as follows: , where is the second weight value, is the first weight value; The calculation formula for the normalized weight coefficient is as follows: , , where α represents the normalized weight coefficient, represents the calculation result of the gradient amplitude corresponding to the image frame in the visible light video, represents the normalization result of the image matrix corresponding to the image frame in the infrared thermal imaging video; Calculating the fused video based on the normalized weight coefficient, the infrared thermal imaging video and the visible light video includes: Based on calculate the R-channel data of the fused image corresponding to the image frame; Based on calculate the G-channel data of the fused image corresponding to the image frame; Based on calculate the B-channel data of the fused image corresponding to the image frame; Among them, A r , A g and A b respectively represent the R-channel data, G-channel data, and B-channel data of the image frames in the visible light video. B r , B g and B b respectively represent the R-channel data, G-channel data, and B-channel data of the image frames in the infrared thermal imaging video. α represents the normalized weight coefficient.

2. The method for predicting the trajectory of dangerous rock masses based on the fusion of visible light and infrared thermal imaging according to claim 1, wherein Obtaining a fused video based on the infrared thermal imaging video and visible light video includes: Input the infrared thermal imaging video and the visible light video into a trained first video fusion model. The first video fusion model splices the image frames of the infrared thermal imaging video and the image frames of the visible light video to obtain a spliced image, and successively passes through a convolutional layer, a downsampling layer, a first dilated convolutional layer, an upsampling layer, a transposed convolutional layer, and a second dilated convolutional layer to obtain the fused image corresponding to the spliced image; Splice the fused images corresponding to each of the spliced images to obtain a fused video.

3. The method for predicting the trajectory of dangerous rock masses based on the fusion of visible light and infrared thermal imaging according to claim 2, wherein The method further includes: Construct a first sample data set, a first loss function, and an initial first video fusion model, and pre-train the initial first video fusion model based on the first sample data set and the first loss function to obtain a trained first video fusion model. The sample data in the first sample data set includes infrared thermal imaging video sample data, visible light video sample data, and fused video sample data; The first loss function is as follows: , where N represents the total number of sample data in the first sample dataset, and L i represents the predicted value of the fused image, represents the label value of the fused image.

4. The method for predicting the trajectory of dangerous rock mass based on the fusion of visible light and infrared thermal imaging according to claim 1, wherein Obtaining a fused video based on the infrared thermal imaging video and visible light video includes: Perform image registration on the image frames in the infrared thermal imaging video and the image frames in the visible light video; and / or Obtain the infrared thermal imaging video and visible light video of the dangerous rock area in the first time period, including: Based on an infrared thermal imaging device and a visible light device with a preset distance from the dangerous rock area, respectively capture the infrared thermal imaging video and visible light video of the dangerous rock area. Both the infrared thermal imaging device and the visible light device are fixed on the device support rod.

5. The method for predicting the trajectory of dangerous rock masses based on the fusion of visible light and infrared thermal imaging according to claim 1, wherein The method includes: Construct a second sample data set, a second loss function, and an initial dangerous rock landslide identification model. Based on the second sample data set and the second loss function, pre-train the initial dangerous rock landslide identification model to obtain a trained dangerous rock landslide identification model. The sample data in the second sample data set includes fused video sample data and rockfall position data sample data corresponding to the next image frame; The second loss function is as follows: , where , ρ(p, g) represents the distance between the predicted value of the rockfall center point position data and the label value of the rockfall center point position data, Q represents the predicted value of the rockfall position data, R represents the label value of the rockfall position data, p represents the predicted value of the rockfall center point position data, and g represents the label value of the rockfall center point position data. represents the intersection over union loss, and Loss represents the second loss function.

6. A device for predicting the trajectory of dangerous rock masses based on the fusion of visible light and infrared thermal imaging, the device comprising a processor, a memory, and a computer program stored on the memory, characterized in that, The processor is used to execute the computer program, and when the computer program is executed, the device realizes the steps of the method according to any one of claims 1 to 5.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it realizes the steps of the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Road slope rockfall risk monitoring and early warning system

    CN116665422A

  • Real-time rockfall monitoring method, system and device based on machine vision and storage medium

    CN118675106A