A method for detecting people falling on the side of a bridge at night based on dual-light fusion

By using dual-light fusion technology in the detection of fallen personnel at night, combining visible light and infrared image data, and using V-I fusion codec detection model for feature extraction and fusion, the problem of detection difficulties in low-light environments at night is solved, accurate detection and state recognition are achieved, and rescue efficiency is improved.

CN119672645BActive Publication Date: 2025-06-06NANJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510181897.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-19
Publication Date
2025-06-06
Estimated Expiration
2045-02-19

AI Technical Summary

Technical Problem

The prior art is difficult to accurately detect fallen people on the bridge side in low light environments at night, and there is a problem of missing texture information when the infrared camera recognizes the motion state.

Method used

The night bridge side fallen personnel detection method based on dual-light fusion is adopted. By acquiring visible-infrared paired image data, the V-I fusion codec detection model is used to extract and fuse the images to generate multimodal information complement each other, thereby achieving accurate detection and state recognition.

Benefits of technology

In the low-light environment at night, people can accurately detect fallen people on the bridge and identify their status at different stages, which improves the timeliness and accuracy of rescue and improves the survival rate of fallen people on the bridge.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119672645B_ABST
    Figure CN119672645B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for detecting people falling on the side of a bridge at night based on dual-light fusion, which relates to the field of computer vision technology and aims to solve the problem that it is difficult to effectively detect people falling on the side of a bridge at night using only visible light cameras in the prior art. The method comprises the steps of obtaining real-time visible light-infrared paired image data of the side of a bridge at night; using the real-time visible light-infrared paired image data of the side of a bridge at night as input, detecting people falling on the side of the bridge based on a V-I fusion encoding and decoding detection model, and outputting a detection result, wherein the detection result includes the state information and position information of the people falling on the side of the bridge. The present invention combines the characteristics of rich texture information of visible light images and accurate mapping of contour information of infrared images through thermal radiation, forming a complementary advantage of multimodal information, which is helpful to accurately detect people falling on the side of a bridge and identify their states at different stages in a low-light environment at night.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to a method for detecting a person falling on the side of a bridge at night based on dual-light fusion, and belongs to the technical field of computer vision. Background Art

[0002] For existing research on the detection of people falling off bridges, researchers are often limited to daytime scenes with good lighting conditions, and only use visible light cameras for detection. Few researchers have conducted research on the detection of people falling off bridges under low-light conditions at night. Compared with people falling off bridges during the day, people falling off bridges in low-light conditions at night have a shorter survival time after falling into the water, and are more difficult to rescue. Therefore, it is very necessary to accurately detect people falling off bridges at night in a timely manner.

[0003] However, it is difficult to effectively detect people falling at night using only visible light cameras, because there is a lot of noise interference in low-light environments at night, and the contour information of the human body extracted by visible light cameras is blurred and difficult to distinguish. Infrared cameras use the thermal radiation of different objects for imaging, so they have strong anti-interference ability in low-light environments and can extract the contour information of the human body well. However, since infrared cameras ignore a lot of texture information, it will be difficult to accurately determine the specific movement status of people falling on the side of the bridge. Summary of the invention

[0004] The purpose of the present invention is to overcome the shortcomings of the prior art and provide a method for detecting people falling on the side of a bridge at night based on dual-light fusion, which combines the characteristics of rich texture information of visible light images and accurate mapping of contour information by thermal radiation of infrared images, forming a complementary advantage of multimodal information, which helps to accurately detect people falling on the side of a bridge and identify their status at different stages in low-light environments at night.

[0005] To achieve the above object, the present invention is implemented by adopting the following technical solutions:

[0006] On the one hand, the present invention provides a method for detecting people falling on the side of a bridge at night based on dual-light fusion, comprising:

[0007] Acquire real-time nighttime bridge-side visible light-infrared paired image data;

[0008] Input the real-time nighttime bridge-side visible light-infrared paired image data into the VI fusion codec detection model.

[0009] The VI fusion encoder of the VI fusion codec detection model is used to extract feature maps of visible light images and infrared images respectively and fuse them to obtain a VI feature fusion sequence;

[0010] The VI fusion decoder of the VI fusion codec detection model iteratively updates the input decoder embedding based on the VI feature fusion sequence and the input target query value to obtain the updated decoder embedding, and obtains the detection result of the person falling on the side of the bridge at night according to the updated decoder embedding;

[0011] The detection result includes the category information and location information of the person who fell on the bridge side.

[0012] Further, the VI fusion encoder includes two groups of identical processing branches, the output ends of the processing branches are connected to the input ends of multiple fusion encoding modules in a one-to-one correspondence, and the output ends of the fusion encoding modules are connected to the feature sequence combination modules;

[0013] Among them, one processing branch is used to perform feature processing on the visible light image to obtain visible light feature sequences at different scales, and the other processing branch is used to perform feature processing on the infrared image to obtain infrared feature sequences at different scales;

[0014] The fusion coding module is used to fuse the visible light feature sequences and infrared feature sequences at different scales to obtain fused feature sequences at different scales;

[0015] The feature sequence combination module is used to splice and combine fusion feature sequences of different scales to obtain a VI feature fusion sequence.

[0016] Further, the processing branch includes a channel expansion module, a downsampling module and a serialization module connected in sequence;

[0017] The channel expansion module is used to expand the image to the same number of channels;

[0018] The downsampling module includes multiple modules, which are used to downsample images with the same number of channels to obtain feature maps of different scales;

[0019] The serialization module is used to flatten feature maps of different scales into one-dimensional tensors to obtain feature sequences of different scales;

[0020] The downsampling module includes a convolutional layer, a maximum pooling layer and a dense residual connection unit connected in sequence.

[0021] Furthermore, the number of the fusion coding modules corresponds to the number of scales of the feature sequence, and the fusion coding module includes two processing branches, and the outputs of the processing branches are added and transposed to obtain a fusion feature sequence;

[0022] The two processing branches each include a self-attention extraction module, a first cross-attention fusion module, a transposition module, and a second cross-attention fusion module that are sequentially connected, and the self-attention extraction module in one processing branch is connected to the first cross-attention fusion module in the other processing branch, and the transposition module in one processing branch is connected to the second cross-attention fusion module in the other processing branch;

[0023] The expression of the self-attention extraction module is:

[0024]

[0025] in, represents the query parameter matrix, represents the key parameter matrix, represents the value parameter matrix, represents the visible light mode or infrared mode, Representing modality The input feature sequence is Representing modality The query matrix, Representing modality The key matrix of Representing modality The value matrix of represents normalization processing, Representing modality The residual self-attention extraction matrix of It means processed by the fully connected layer. Representing modality The self-attention extraction output of

[0026] The first cross attention fusion module and the second cross attention fusion module are the same, and their expressions are:

[0027]

[0028] in, Indicates that the mode is removed Another mode besides Representing modality The input feature sequence is Representing modality The key matrix of Representing modality The value matrix of represents the input dimension, represents the bimodal residual fusion attention extraction matrix, Represents the bimodal fusion output.

[0029] Further, the VI fusion decoder includes a plurality of VI decoding modules connected in sequence, and the VI decoding module is used to iteratively update the input decoder embedding using the VI feature fusion sequence and the input target query value to obtain an updated decoder embedding;

[0030] The output end of the VI decoding module at the end is connected to two fully connected layers, wherein one of the fully connected layers is used to obtain the category information of the person falling on the bridge side according to the updated decoder embedding, and the other fully connected layer is used to obtain the position information of the person falling on the bridge side according to the updated decoder embedding.

[0031] Furthermore, the VI decoding module is specifically used for:

[0032] The input target query value is trigonometrically encoded so that the dimension of the target query value is consistent with the dimension of the input decoder embedding and VI feature fusion sequence, and the encoded target query value is obtained, which is expressed as follows:

[0033]

[0034] in, represents the set of k-th encoded target query values ​​in the n-th group located at an even position, represents the set of k-th encoded target query values ​​in the n-th group located at odd positions, represents the kth encoded target query value of the nth group, b represents the number of encoded bits, Represents a preset natural number set;

[0035] Add the encoded target query value as an auxiliary supervision signal to the self-attention extraction process embedded in the input decoder to obtain a supervised self-attention extraction sequence;

[0036] The supervised self-attention extraction sequence is cross-attended with the VI feature fusion sequence to obtain the updated decoder embedding;

[0037] The updated decoder embedding is processed using a fully connected layer to obtain a target query value offset, and the target query value offset is multiplied by the input target query value to obtain an updated target query value;

[0038] The input target query value and the input decoder embedding of the VI decoding module are respectively the target query value and the updated decoder embedding updated by the previous VI decoding module; the target query value updated by the previous VI decoding module is obtained by the previous VI decoding module processing the decoder embedding updated by the previous VI decoding module using a fully connected layer to obtain a target query value offset, and the target query value offset is multiplied by the input target query value of the previous VI decoding module;

[0039] The input target query value of the VI decoding module at the head end is set to a matrix initialized to zero, the input decoder embedding is set to random initialization, and the length of the input decoder embedding is consistent with the length of the fused feature sequence.

[0040] Furthermore, the method further includes pre-training the VI fusion codec detection model before inputting the real-time nighttime bridge side visible light-infrared paired image data into the VI fusion codec detection model, and the pre-training method includes:

[0041] S1. Construct an original visible light-infrared image dataset of people falling on the side of a bridge at night;

[0042] S2, annotating the original visible light-infrared image dataset of people falling on the side of a bridge at night, and obtaining a dataset of images of people falling on the side of a bridge at night containing annotation information;

[0043] S3, dividing the nighttime bridge side fall image dataset containing labeled information into a training set and a test set;

[0044] S4, using the training set data as input, training the VI fusion codec detection model, calculating the loss function during the training process, and using the loss function to update the parameters of the VI fusion codec detection model;

[0045] S5. Use the test set data as input, test the VI fusion codec detection model after parameter update to obtain test results, calculate the evaluation index according to the test results, if the evaluation index is higher than the preset value, the training is completed, and the pre-trained VI fusion codec detection model is obtained; if the evaluation index is lower than the preset value, repeat S4 to S5 until the evaluation index is higher than the preset value, and the pre-trained VI fusion codec detection model is obtained.

[0046] Furthermore, the construction of the original visible light-infrared image dataset of people falling on the side of the bridge at night includes:

[0047] Obtain visible light and infrared images of people falling on the side of the bridge at night;

[0048] Performing size transformation on the visible light image and the infrared image to obtain a visible light-infrared paired image;

[0049] Data enhancement is performed on the visible light-infrared paired images, including geometric transformation enhancement and pixel transformation enhancement. The enhanced visible light-infrared paired images are added to the dataset to obtain the original visible light-infrared image dataset of people falling on the side of the bridge at night.

[0050] Furthermore, the annotation information includes position information and category information of the target frame, the position information includes the horizontal coordinate and the vertical coordinate of the center point of the target frame, the width and the height of the target frame, and the category information includes crossing the boundary, falling and falling into water.

[0051] Furthermore, the loss function includes a position loss function, a classification loss function and a supervision loss function, and the expression of the loss function is:

[0052]

[0053] in, represents the loss function, represents the position loss function, represents the classification loss function, represents the supervision loss function, , , Respectively represent the weight coefficients of the position loss function, classification loss function, and supervision loss function;

[0054] The expression of the position loss function is:

[0055]

[0056] Among them, M represents the number of real target boxes, represents the i-th true target box, represents the i-th predicted target box, express and The intersection ratio of express and The square of the Euclidean distance between the center points, Indicates that it also contains and The diagonal distance of the minimum closure area;

[0057] The expression of the classification loss function is:

[0058]

[0059] Where N represents the total number of categories. represents the i-th category, Represents the value of the i-th category in the true label, which is 0 or 1. Represents the probability of the i-th category predicted by the model;

[0060] The expression of the supervised loss function is:

[0061]

[0062] Among them, L represents the number of target query boxes, Indicates A target query box.

[0063] Compared with the prior art, the present invention has the following beneficial effects:

[0064] The present invention combines the characteristics of rich texture information of visible light images and accurate mapping of contour information of infrared images through thermal radiation, forming a complementary advantage of multimodal information, which helps to accurately detect people falling on the side of the bridge and identify their status at different stages in low-light environments at night, thereby guiding rescue personnel to carry out rescue in a timely and accurate manner and improving the survival rate of people falling on the side of the bridge;

[0065] The VI fusion coding and decoding detection model provided by the present invention includes a VI fusion encoder and a VI decoding detector. The visible light-infrared image multi-scale fusion structure designed in the VI fusion encoder comprehensively represents target information of different sizes and enhances the generalization of the model; at the same time, the VI fusion encoding module can not only perform deep information extraction on a single visible light or infrared modality sequence, enhance the internal features of each modality sequence, but also perform cross-attention fusion on the two modality sequences, and comprehensively utilize the unrelated complementary information of the two, so that the fused feature sequence has both rich texture details in the visible light modality and infrared energy information that cannot be captured by visible light in the infrared modality; the VI decoding module in the VI decoding detector introduces a target query value to explicitly provide a position supervision signal for the decoder embedding, so as to guide the decoder embedding to gradually iteratively update the position information of the real target frame, which not only alleviates the problem that the position information is more difficult to learn than the category information, but also accelerates the convergence speed of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0066] Figure 1 It is a schematic flow chart of a method for detecting a person falling on the side of a bridge at night based on dual-light fusion in an embodiment of the present invention;

[0067] Figure 2 It is a schematic diagram of the structure of a VI fusion codec detection model in an embodiment of the present invention;

[0068] Figure 3 A schematic diagram of the structure of a VI fusion encoder in an embodiment of the present invention;

[0069] Figure 4 A schematic diagram of the structure of a VI fusion coding module in an embodiment of the present invention;

[0070] Figure 5 FIG. 4 is a schematic diagram of the structure of a VI fusion decoder in an embodiment of the present invention. DETAILED DESCRIPTION

[0071] The present invention will be further described below in conjunction with the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present invention, and cannot be used to limit the protection scope of the present invention.

[0072] Embodiment 1:

[0073] like Figure 1 As shown, an embodiment of the present invention provides a method for detecting a person falling on the side of a bridge at night based on dual-light fusion, comprising the following steps:

[0074] Acquire real-time nighttime bridge-side visible light-infrared paired image data. Paired images refer to visible light images and infrared images with the same time, space, and image size.

[0075] A VI fusion coding and decoding detection model is constructed, and the real-time nighttime visible light-infrared paired image data of the bridge side is taken as input. The fallen people on the bridge side are detected based on the VI fusion coding and decoding detection model, and the detection results are output.

[0076] In this embodiment, the detection result includes position information and category information. The position information includes the horizontal coordinate and the vertical coordinate of the center point of the target frame, the width and the height of the target frame, and the category information includes crossing the boundary, falling, and falling into water.

[0077] like Figure 2 As shown, the VI fusion encoding and decoding detection model includes a VI fusion encoder and a VI fusion decoder connected in sequence, wherein the VI fusion encoder is used to extract feature maps of visible light images and infrared images respectively and fuse them to obtain a VI feature fusion sequence, and the VI fusion decoder is used to iteratively update the input decoder embedding using the VI feature fusion sequence and the input target query value to obtain an updated decoder embedding, and obtain the category information and location information of the person falling on the bridge side according to the updated decoder embedding.

[0078] The structure of the VI fusion encoder is as follows Figure 3 As shown, the VI fusion encoder includes two groups of identical processing branches, one processing branch is used to perform feature processing on the visible light image to obtain visible light feature sequences at different scales, and the other processing branch is used to perform feature processing on the infrared image to obtain infrared feature sequences at different scales.

[0079] The processing branches all include a channel expansion module, a downsampling module and a serialization module connected in sequence. The channel expansion module is used to expand the image to the same number of channels. The downsampling module includes multiple modules, which are used to downsample images with the same number of channels to obtain feature maps of different scales. In this embodiment, the number of downsampling modules is 5, and each downsampling module includes a convolutional layer, a maximum pooling layer and a dense residual connection unit connected in sequence. The serialization module is used to flatten feature maps of different scales into one-dimensional tensors to obtain feature sequences of different scales.

[0080] The output ends of the serialization modules of the two processing branches are connected to the input ends of multiple fusion coding modules in a one-to-one correspondence. The fusion coding modules are used to fuse the visible light feature sequences and infrared feature sequences at different scales to obtain fusion feature sequences of different scales. The output ends of the fusion coding modules are connected to the feature sequence combination module, which is used to splice and combine the fusion feature sequences of different scales to obtain the VI feature fusion sequence.

[0081] The number of fusion coding modules corresponds to the number of scales of the feature sequence, and in this embodiment, the number of fusion coding modules is 3. The internal structure of the fusion coding module includes two processing branches, each of which includes a self-attention extraction module, a first cross-attention fusion module, a transposition module, and a second cross-attention fusion module connected in sequence, and the self-attention extraction module in one processing branch is connected to the first cross-attention fusion module in the other processing branch, and the transposition module in one processing branch is connected to the second cross-attention fusion module in the other processing branch, and the outputs of the two processing branches are added and transposed to obtain a fused feature sequence.

[0082] In this embodiment, the data processing process of the VI fusion encoder specifically includes:

[0083] The sizes of the input visible light image and infrared image are 320×320×3 and 320×320×1 respectively (because the visible light image has three channels of R, G, and B, and the infrared image has a single channel, so even paired images have different sizes. Pairing only means that the time and space image sizes are the same).

[0084] The visible light image and the infrared image are respectively expanded to the same number of channels through the channel expansion module to obtain the visible light feature map 0 and the infrared feature map 0, that is, the sizes of the visible light feature map 0 and the infrared feature map 0 are both 320×320×256.

[0085] Then, the downsampling module performs five consecutive downsampling operations to extract visible light feature maps and infrared feature maps of different scales. Each downsampling reduces the width and height of the input feature map by half, and the number of channels remains unchanged. Therefore, a total of five visible light feature maps and infrared feature maps of different sizes are extracted.

[0086] Since the number of fusion coding modules in this embodiment is 3, the feature maps of the last three times are taken for subsequent fusion operations, that is, the visible light and infrared feature maps of sizes 40×40×256, 20×20×256 and 10×10×256 are Figure 3 , 4, 5 are serialized and sent to the three fusion coding modules respectively, where serialization refers to flattening the two-dimensional tensor of width and height into a one-dimensional tensor to obtain the visible light feature sequence and the infrared feature sequence, that is, the sizes of the visible light and infrared feature maps (i.e., visible light feature sequence and infrared feature sequence) after serialization are 1600×256, 400×256 and 100×256 respectively.

[0087] Combination Figure 4 The data processing process inside the fusion coding module is described by taking the fusion coding module 1 as an example. The input of the fusion coding module 1 is a visible light feature sequence and an infrared feature sequence of size 1600×256. The process includes:

[0088] Firstly, the visible light feature sequence and infrared feature sequence are respectively subjected to self-attention extraction by the self-attention extraction module to enhance the internal features of their respective modal sequences.

[0089] The self-attention extraction process includes: multiplying the single modality (visible light modality or infrared modality) input sequence of size 1600×256 by three different parameter matrices of size 256×256 (the parameter matrix is ​​randomly set, and then converges according to the gradient descent algorithm as the model is trained), and obtaining the query matrix Q, key matrix K and value matrix V. Multiply the transpose of the query matrix Q and the key matrix K and pass it through the Softmax function to obtain the self-attention matrix; then multiply the self-attention matrix by the value matrix V to obtain the self-attention extraction matrix; add the residual of the self-attention extraction matrix to the original input sequence and normalize it to obtain the residual self-attention extraction matrix; pass the residual self-attention extraction matrix through the fully connected layer, and then add the residual to itself and normalize it to obtain a single-modality self-attention extraction output of size 1600×256.

[0090] Therefore, the expression of the self-attention extraction module can be described as:

[0091]

[0092] in, represents the query parameter matrix, represents the key parameter matrix, represents the value parameter matrix, represents the visible light mode or infrared mode, Representing modality The input feature sequence is Representing modality The query matrix, Representing modality The key matrix of Representing modality The value matrix of represents the input dimension, represents normalization processing, Representing modality The residual self-attention extraction matrix of It means that it is processed by the fully connected layer. Representing modality The self-attention extraction output.

[0093] Next, in order to enhance the unrelated complementary information between the visible light modality and the infrared modality and optimize the information fusion process between the two modalities, the visible light modality sequence and the infrared modality sequence that have completed self-attention extraction need to be cross-attended twice by the cross-attention fusion module respectively.

[0094] The process of the first cross-attention fusion is as follows: multiply the unimodal input sequence of size 1600×256 by a parameter matrix of size 256×256 to obtain the query matrix Q, and multiply the other modal input sequence of size 1600×256 by two different parameter matrices of size 256×256 to obtain the key matrix K and the value matrix V; multiply the query matrix Q and the transpose of the key matrix K and pass them through the Softmax function to obtain the fused attention matrix, and then multiply the fused attention matrix by the value matrix V to obtain the fused attention extraction matrix; add the residual of the fused attention extraction matrix to the other modal input sequence and normalize them to obtain the residual fused attention extraction matrix of size 1600×256; pass the residual fused attention extraction matrix through the fully connected layer, and then add the residual to itself and normalize it to obtain the bimodal fusion output.

[0095] Therefore, the expression of the cross-attention fusion module can be described as:

[0096]

[0097] in, Indicates that the mode is removed Another mode besides Representing modality The input feature sequence is Representing modality The key matrix of Representing modality The value matrix of represents the input dimension, represents the bimodal residual fusion attention extraction matrix, Represents the bimodal fusion output.

[0098] A transposition operation is introduced between the two cross-attention fusions (i.e., flipping the fused sequence horizontally and vertically) to further enhance the internal features of the fused sequence so that it contains more global information. Due to the introduction of the transposition operation, the input sequence size of the two modalities in the second cross-attention fusion process becomes 256×1600, and the corresponding parameter matrix size becomes 1600×1600. At the same time, a bimodal fusion output of size 256×1600 is obtained, and the remaining steps are consistent with the first cross-attention fusion process.

[0099] Finally, the two bimodal fusion feature sequences of size 256×1600 after two cross-attention fusions are added and then transposed to obtain a 1600×256 fusion feature sequence output with the same size as the input.

[0100] Similarly, the data processing process of fusion coding module 2 and fusion coding module 3 is similar to that of fusion coding module 1, except that the size of the visible light feature sequence and infrared feature sequence input to fusion coding module 2 is 400×256, and the size of the output fusion feature sequence is also 400×256; the size of the visible light feature sequence and infrared feature sequence input to fusion coding module 3 is 100×256, and the size of the output fusion feature sequence is also 100×256, which will not be elaborated here.

[0101] The multi-scale fusion feature sequences output by the fusion coding modules 1, 2, and 3 are concatenated to obtain a VI feature fusion sequence of size 2100×256, which is the output of the VI fusion encoder.

[0102] like Figure 5 As shown, the VI fusion decoder includes a plurality of VI decoding modules connected in sequence, and in this embodiment, the number of VI decoding modules is 3. The VI decoding module is used to iteratively update the decoder embedding using the VI feature fusion sequence and the target query value to obtain an updated decoder embedding.

[0103] The output end of the VI decoding module at the end is connected to two fully connected layers, which are used to obtain the category information and location information of the person falling on the bridge side according to the updated decoder embedding.

[0104] The data processing process of the VI decoding module specifically includes:

[0105] First, the input target query value is trigonometrically encoded so that the dimensions of the target query value, decoder embedding, and fusion feature sequence are aligned to obtain the encoded target query value. It should be noted that the size of the target query value is 100×4, where 100 represents the number of target boxes predicted by the model in this embodiment.

[0106] The expression for trigonometric encoding of the input target query value is as follows:

[0107]

[0108] in, represents the set of k-th encoded target query values ​​in the n-th group located at an even position, represents the set of k-th encoded target query values ​​in the n-th group located at odd positions, represents the kth encoded target query value of the nth group, b represents the number of encoded bits, Represents a preset set of natural numbers.

[0109] To ensure that the length of each target query value changes from 4 to 256, the four position data in each target query value need to be encoded into 64 bits. Therefore, d is set to 64. represents the set of even-numbered items in the coded bit number, ,so The value ranges from 0 to 31.

[0110] The encoded target query value is then added as an auxiliary supervisory signal to the self-attention extraction embedded in the input decoder. The adding process includes: adding the encoded target query value to the query matrix and key matrix embedded in the decoder to obtain a new query matrix Q and key matrix K. Then self-attention extraction is performed to obtain a supervised self-attention extraction sequence. The self-attention extraction process is the same as the self-attention extraction module, so no further details are given.

[0111] In this embodiment, the size of the input decoder embedding is 100×256, where 100 represents the number of target boxes predicted by the model in this embodiment, and the input decoder embedding for the first VI decoding module is set to random initialization.

[0112] The supervised self-attention extraction sequence is cross-attention fused with the VI feature fusion sequence to obtain a fused decoding sequence of size 100×256, which is the updated decoder embedding.

[0113] After the fused decoding sequence passes through the fully connected layer, a target query value offset of size 100×4 is obtained, and the input target query value is multiplied by the target query value offset to obtain an updated target query value of size 100×4.

[0114] The input target query value of the first VI decoding module is initialized to zero, and the input target query values ​​of the subsequent two VI decoding modules are the updated target query values ​​output by the previous VI decoding module. The fused decoding sequence output by the previous VI decoding module is used as the input encoder embedding of the next VI decoding module. By dynamically updating the target query value and decoder embedding, the decoder embedding is guided to gradually iterate toward the position of the true target frame.

[0115] The fused decoding sequence output by the last VI decoding module passes through two fully connected layers to obtain the category information and position information of 100 predicted target boxes. The 100 predicted target boxes include the background box, which is not displayed in the final image.

[0116] Embodiment 2:

[0117] Based on Example 1, this embodiment also provides a pre-training method for a VI fusion codec detection model, which includes the following steps:

[0118] S1. Construct an original visible light-infrared image dataset of people falling on the side of a bridge at night, including:

[0119] Obtain visible light and infrared images of people falling on the side of the bridge at night;

[0120] Performing size transformation on the visible light image and the infrared image to obtain a visible light-infrared paired image;

[0121] Data enhancement is performed on the visible light-infrared paired images, including geometric transformation enhancement and pixel transformation enhancement. The enhanced visible light-infrared paired images are added to the dataset to obtain the original visible light-infrared image dataset of people falling on the side of the bridge at night.

[0122] S2. Annotate the original visible light-infrared image dataset of people falling on the side of a bridge at night. The annotation information includes the location information and category information of the target frame. The location information includes the horizontal coordinate and vertical coordinate of the center point of the target frame, the width and height of the target frame, and the category information includes crossing the boundary, falling, and falling into water, so as to obtain a nighttime bridge side falling image dataset containing annotation information.

[0123] S3, dividing the nighttime bridge side fall image dataset containing labeled information into a training set and a test set;

[0124] S4. Use the training set data as input to train the VI fusion codec detection model, calculate the loss function during the training process, and use the loss function to update the model parameters.

[0125] The loss function includes position loss function, classification loss function and supervision loss function, and its expression is:

[0126]

[0127] in, represents the loss function, represents the position loss function, represents the classification loss function, represents the supervision loss function, , , They represent the weight coefficients of the position loss function, classification loss function, and supervision loss function respectively.

[0128] The expression of the position loss function is:

[0129]

[0130] Among them, M represents the number of real target boxes, represents the i-th true target box, represents the i-th predicted target box, express and The intersection ratio of express and The square of the Euclidean distance between the center points, Indicates that it also contains and The diagonal distance of the minimum closure area.

[0131] The expression of the classification loss function is:

[0132]

[0133] Where N represents the total number of categories. represents the i-th category, Represents the value of the i-th category in the true label, which is 0 or 1. Represents the probability of the i-th category predicted by the model.

[0134] The expression of the supervised loss function is:

[0135]

[0136] Among them, L represents the number of target query boxes, Indicates A target query box.

[0137] S5. Use the test set data as input to test the VI fusion codec detection model to obtain test results, and calculate evaluation indicators based on the test results. In this embodiment, the evaluation indicators include precision, recall, and average precision. If the evaluation indicator is higher than the preset value, the training is completed and the pre-trained VI fusion codec detection model is obtained. If the evaluation indicator is lower than the preset value, S4 to S5 are repeated until the evaluation indicator is higher than the preset value to obtain the pre-trained VI fusion codec detection model.

[0138] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.

Claims

1. A method for detecting people falling on the side of a bridge at night based on dual-light fusion, characterized in that: include: Acquire real-time nighttime bridge-side visible light-infrared paired image data; Input the real-time nighttime bridge-side visible light-infrared paired image data into the VI fusion encoding and decoding detection model; The VI fusion encoder of the VI fusion coding and decoding detection model is used to extract feature maps of visible light images and infrared images respectively and fuse them to obtain a VI feature fusion sequence. The VI fusion encoder includes two groups of identical processing branches, and the output ends of the processing branches are connected to the input ends of multiple fusion coding modules in a one-to-one correspondence. The output ends of the fusion coding modules are connected to the feature sequence combination modules, wherein one of the processing branches is used to perform feature processing on the visible light image to obtain a visible light feature sequence at different scales, and the other processing branch is used to perform feature processing on the infrared image to obtain an infrared feature sequence at different scales. The number of the fusion coding modules corresponds to the number of scales of the feature sequence, the fusion coding module includes two processing branches, and the outputs of the processing branches are added and transposed to obtain a fusion feature sequence; the two processing branches each include a self-attention extraction module, a first cross-attention fusion module, a transposition module, and a second cross-attention fusion module connected in sequence, and the self-attention extraction module in one processing branch is connected to the first cross-attention fusion module in the other processing branch, and the transposition module in one processing branch is connected to the second cross-attention fusion module in the other processing branch; The VI fusion decoder of the VI fusion codec detection model iteratively updates the input decoder embedding based on the VI feature fusion sequence and the input target query value to obtain an updated decoder embedding, and obtains the detection result of the person falling on the side of the bridge at night according to the updated decoder embedding, and the detection result includes the category information and location information of the person falling on the side of the bridge; The VI fusion decoder includes a plurality of VI decoding modules connected in sequence, and the VI decoding modules are used to iteratively update the input decoder embedding using the VI feature fusion sequence and the input target query value to obtain an updated decoder embedding; The output end of the VI decoding module at the end is connected to two fully connected layers, wherein one of the fully connected layers is used to obtain the category information of the person falling on the bridge side according to the updated decoder embedding, and the other fully connected layer is used to obtain the position information of the person falling on the bridge side according to the updated decoder embedding.

2. The method for detecting people falling on the side of a bridge at night based on dual light fusion according to claim 1 is characterized in that: The fusion coding module is used to fuse the visible light feature sequences and infrared feature sequences at different scales to obtain fused feature sequences at different scales; The feature sequence combination module is used to splice and combine fusion feature sequences of different scales to obtain a VI feature fusion sequence.

3. The method for detecting people falling on the side of a bridge at night based on dual light fusion according to claim 1 is characterized in that: The processing branch includes a channel expansion module, a downsampling module and a serialization module connected in sequence; The channel expansion module is used to expand the image to the same number of channels; The downsampling module includes multiple modules, which are used to downsample images with the same number of channels to obtain feature maps of different scales; The serialization module is used to flatten feature maps of different scales into one-dimensional tensors to obtain feature sequences of different scales; The downsampling module includes a convolutional layer, a maximum pooling layer and a dense residual connection unit connected in sequence.

4. The method for detecting people falling on the side of a bridge at night based on dual light fusion according to claim 1 is characterized in that: The expression of the self-attention extraction module is: ; in, represents the query parameter matrix, represents the key parameter matrix, represents the value parameter matrix, represents the visible light mode or infrared mode, Representing modality The input feature sequence is Representing modality The query matrix, Representing modality The key matrix of Representing modality The value matrix of represents normalization processing, Representing modality The residual self-attention extraction matrix of It means processed by the fully connected layer. Representing modality The self-attention extraction output of The first cross attention fusion module and the second cross attention fusion module are the same, and their expressions are: ; in, Indicates that the mode is removed Another mode besides Representing modality The input feature sequence is Representing modality The key matrix of Representing modality The value matrix of represents the input dimension, represents the bimodal residual fusion attention extraction matrix, Represents the bimodal fusion output.

5. The method for detecting people falling on the side of a bridge at night based on dual light fusion according to claim 1 is characterized in that: The VI decoding module is specifically used for: The input target query value is trigonometrically encoded so that the dimension of the target query value is consistent with the dimension of the input decoder embedding and VI feature fusion sequence, and the encoded target query value is obtained, which is expressed as follows: ; in, represents the set of k-th encoded target query values ​​in the n-th group located at an even position, represents the set of k-th encoded target query values ​​in the n-th group located at odd positions, represents the kth encoded target query value of the nth group, b represents the number of encoded bits, Represents a preset natural number set; Add the encoded target query value as an auxiliary supervision signal to the self-attention extraction process embedded in the input decoder to obtain a supervised self-attention extraction sequence; Cross-attention fusion is performed on the supervised self-attention extraction sequence and the VI feature fusion sequence to obtain an updated decoder embedding; wherein the input target query value and the input decoder embedding of the VI decoding module are respectively the updated target query value and the updated decoder embedding of the previous VI decoding module; the updated target query value of the previous VI decoding module is obtained by the previous VI decoding module processing the updated decoder embedding of the previous VI decoding module using a fully connected layer to obtain a target query value offset, and the target query value offset is multiplied by the input target query value of the previous VI decoding module; The input target query value of the VI decoding module at the head end is set to a matrix initialized to zero, the input decoder embedding is set to random initialization, and the length of the input decoder embedding is consistent with the length of the fused feature sequence.

6. The method for detecting people falling on the side of a bridge at night based on dual light fusion according to claim 1 is characterized in that: The method also includes pre-training the VI fusion codec detection model before inputting the real-time nighttime bridge side visible light-infrared paired image data into the VI fusion codec detection model. The pre-training method includes: S1. Construct an original visible light-infrared image dataset of people falling on the side of a bridge at night; S2, annotating the original visible light-infrared image dataset of people falling on the side of a bridge at night, and obtaining a dataset of images of people falling on the side of a bridge at night containing annotation information; S3, dividing the nighttime bridge side fall image dataset containing labeled information into a training set and a test set; S4, using the training set data as input, training the VI fusion codec detection model, calculating the loss function during the training process, and using the loss function to update the parameters of the VI fusion codec detection model; S5. Use the test set data as input, test the VI fusion codec detection model after parameter update to obtain test results, calculate the evaluation index according to the test results, if the evaluation index is higher than the preset value, the training is completed, and the pre-trained VI fusion codec detection model is obtained; if the evaluation index is lower than the preset value, repeat S4 to S5 until the evaluation index is higher than the preset value, and the pre-trained VI fusion codec detection model is obtained.

7. The method for detecting people falling on the side of a bridge at night based on dual light fusion according to claim 6 is characterized in that: The construction of the original visible light-infrared image dataset of people falling on the side of the bridge at night includes: Obtain visible light and infrared images of people falling on the side of the bridge at night; Performing size transformation on the visible light image and the infrared image to obtain a visible light-infrared paired image; Data enhancement is performed on the visible light-infrared paired images, including geometric transformation enhancement and pixel transformation enhancement. The enhanced visible light-infrared paired images are added to the dataset to obtain the original visible light-infrared image dataset of people falling on the side of the bridge at night.

8. The method for detecting people falling on the side of a bridge at night based on dual light fusion according to claim 6 is characterized in that: The annotation information includes position information and category information of the target frame, the position information includes the horizontal coordinate and the vertical coordinate of the center point of the target frame, the width and the height of the target frame, and the category information includes crossing the boundary, falling and falling into water.

9. The method for detecting people falling on the side of a bridge at night based on dual light fusion according to claim 6, characterized in that: The loss function includes a position loss function, a classification loss function and a supervision loss function, and the expression of the loss function is: ; in, represents the loss function, represents the position loss function, represents the classification loss function, represents the supervision loss function, , , Respectively represent the weight coefficients of the position loss function, classification loss function, and supervision loss function; The expression of the position loss function is: ; Among them, M represents the number of real target boxes, represents the i-th true target box, represents the i-th predicted target box, express and The intersection ratio of express and The square of the Euclidean distance between the center points, Indicates that it also contains and The diagonal distance of the minimum closure area; The expression of the classification loss function is: ; Where N represents the total number of categories. represents the i-th category, Represents the value of the i-th category in the true label, which is 0 or 1. Represents the probability of the i-th category predicted by the model; The expression of the supervised loss function is: ; Among them, L represents the number of target query boxes, Indicates A target query box.

Citation Information

Patent Citations

  • Infrared visible image fusion method based on self-supervised feature decoupling

    CN116468644A

  • Using scene dependent object queries to generate bounding boxes

    US20240127597A1