Obstacle detection method and system for road inspection robots based on RT-DETR-Sat

By adopting the RT-DETR-Sat object detection model in the road patrol robot, combining the ResNeSat backbone network and hybrid encoder, the real-time and accuracy problems of obstacle detection in complex road environments are solved, and efficient and accurate obstacle detection is achieved.

CN118864424BActive Publication Date: 2025-05-09EAST CHINA JIAOTONG UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411028358.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-30
Publication Date
2025-05-09
Estimated Expiration
2044-07-30

AI Technical Summary

Technical Problem

The prior art is difficult to realize real-time and high-precision obstacle detection in complex road environments, and the traditional patrol robot object detection algorithm has problems such as low real-time and poor accuracy.

Method used

The road patrol robot obstacle detection method based on RT-DETR-Sat is adopted, and the ResNeSat model is combined with a hybrid encoder and a decoder to realize real-time detection of obstacles through end-to-end training.

Benefits of technology

The road inspection robot has achieved accurate and efficient detection of obstacles such as trees, rocks, construction materials, road maintenance facilities, animals, etc., and improved the generalization ability of the model and adaptability to different environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118864424B_ABST
    Figure CN118864424B_ABST
Patent Text Reader

Abstract

The present invention is a road inspection robot obstacle detection method and system based on RT-DETR-Sat, the detection method comprises the following steps: obtaining road images of the inspection area during the road inspection process, annotating the road images frame by frame, and dividing the annotated obstacle images into a training set and a test set; constructing an RT-DETR-Sat target detection model, the RT-DETR-Sat target detection model comprises a backbone network ResNeSat, a hybrid encoder and a decoder with an auxiliary prediction head, inputting images into the backbone network ResNeSat, obtaining multi-layer feature maps of different scales, using a hybrid encoder to process multi-scale features, connecting the output of the hybrid encoder to an IoU-aware query selection, selecting a fixed number of features from the query as the initial target query of the decoder, the decoder generating a bounding box and a confidence score, and finally converting the output of the model into a probability distribution to obtain a final prediction result for road inspection robot obstacle detection. It takes into account both high real-time performance and high detection accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of road inspection in prefabricated factories, and in particular to an obstacle detection method and system for a road inspection robot based on RT-DETR-Sat. Background Art

[0002] In prefabricated factories, obstacles on factory roads will affect the transportation safety of prefabricated components. Inspections can not only identify and deal with potential safety hazards in a timely manner to avoid personal injury or property loss, but also help ensure that the functional performance of prefabricated roads meets the design requirements and improve the overall efficiency and comfort of vehicle driving. Conventional manual inspections have the disadvantages of high labor intensity, low work efficiency, inability to quickly cover large areas, and difficulty in rapid data processing and analysis; and the quality of manual inspections is unstable, and the results of inspections are easily affected by personal experience, skills, and the work status of the day, resulting in low stability and reliability of the inspection results. At the same time, manual inspections are also cost-effective.

[0003] Compared with traditional manual inspection, robot inspection can work in dangerous environments such as high temperature, high voltage, and strong electricity, avoiding the risk of personal injury to inspection personnel and reducing the safety risks caused by improper equipment operation. Traditional inspection robot target detection mainly uses the YOLO series of algorithms. Although it can identify obstacles to a certain extent, the robot's environmental adaptability in complex road environments still needs to be improved. The robot's detection ability when facing unknown obstacles or complex situations still needs to be further improved. Although YOLO also supports end-to-end training, its processing process is complicated. It usually requires region proposal first, and then classification and bounding box regression. It has the disadvantages of low real-time performance and poor accuracy.

[0004] Due to the complexity of the roads in the prefabricated factories in this application, the existing target detection algorithms cannot meet the dual requirements of real-time performance and accuracy. Therefore, a new road inspection robot obstacle detection system based on the ResNeSat model is proposed. Summary of the invention

[0005] The purpose of the present invention is to overcome the shortcomings of the above technologies and provide a road inspection robot obstacle detection method and system based on RT-DETR-Sat, which takes into account both high real-time performance and high detection accuracy and is suitable for road inspection.

[0006] To achieve the above object, the technical solution adopted by the present invention is:

[0007] In a first aspect, the present invention provides a road inspection robot obstacle detection method based on RT-DETR-Sat, the detection method comprising the following steps:

[0008] Obtain road images of the inspection area during the road inspection process, annotate the road images frame by frame, and divide the annotated obstacle images into a training set and a test set;

[0009] Construct an RT-DETR-Sat target detection model, which includes a backbone network ResNeSat, a hybrid encoder and a decoder with an auxiliary prediction head. The image is input into the backbone network ResNeSat, and feature extraction is performed to obtain multi-layer feature maps of different scales. The multi-scale features are then processed using the hybrid encoder. The output of the hybrid encoder is connected to the IoU-aware query selection, and a fixed number of features are selected as the initial target query of the decoder. The decoder then generates a bounding box and a confidence score. Finally, the output of the model is converted into a probability distribution, thereby obtaining a final prediction result, and then judging whether there is an obstacle.

[0010] The road image is used as the input of the backbone network ResNeSat. It first passes through a convolution layer with a 7*7 convolution kernel, and the number of output channels of the convolution layer is 64. Then it passes through a 3*3 pooling layer. After pooling, the data enters four convolution groups. These four convolution groups are composed of two 1*1 convolution kernels and one 3*3 convolution kernel. These convolution groups are repeated for different times. The first convolution group is repeated 3 times, and the number of output channels is 256. The second convolution group is repeated 4 times, and the number of output channels is 512. The third convolution group is repeated 6 times, and the number of output channels is 1024. The fourth convolution group is repeated 3 times, and the number of output channels is 2048. After that, the self-attention mechanism is used to obtain output features with step sizes of 8, 16, and 32, which are recorded as S3, S4, and S5 respectively. S3, S4, and S5 are used as input features of the hybrid encoder.

[0011] The hybrid encoder includes an intra-scale feature interaction module AIFI and a multi-scale feature fusion module CCFM. The input feature S5 is used as the input of the intra-scale feature interaction module AIFI, and F5 is obtained after intra-scale interaction.

[0012] S3, S4, and F5 are used as the input of the multi-scale feature fusion module CCFM to perform multi-scale feature fusion. The specific processing process of multi-scale feature fusion is:

[0013] The input features of F5 are processed by 1*1 convolution, IN normalization and activation function ReLU calculation, and the input features of S4 are processed by 1*1 convolution respectively, and then the elements are added together to obtain the first fusion result;

[0014] The first fusion result is processed by 1*1 convolution, IN normalization and activation function, and then concatenated with the S3 input feature by 1*1 convolution and element-wise addition to obtain the second fusion result. The output of the first fusion result is processed by activation function and then connected with the 1*1 convolution of the branch where the S4 input feature is located.

[0015] The second fusion result is processed by 3*3 convolution, IN normalization and activation function, and then connected with the 1*1 convolution of the branch where the F5 input feature is located in the first fusion result; the first fusion result is processed by 3*3 convolution, IN normalization and activation function, and then processed by 1*1 convolution with the activation function processing result of the branch where the F5 input feature is located, and then the feature splicing of element addition is obtained to obtain the third fusion result;

[0016] The first fusion result, the second fusion result and the third fusion result are concatenated to obtain the output of the hybrid encoder;

[0017] The decoder includes a group query attention mechanism, a residual connection and a normalization process, a group query attention mechanism, a residual connection and a normalization process, a multi-layer perceptron, a residual connection and a normalization process, which are connected in sequence;

[0018] The output of the decoder is used as the input of a feedforward neural network, wherein the feedforward neural network includes a linear combination, a BN normalization process, and a softmax activation function;

[0019] The training set is used to train the RT-DETR-Sat target detection model, which is then used for obstacle detection by road inspection robots.

[0020] Furthermore, the IoU-aware query selection constrains the model to produce high classification scores for features with high IoU scores and low classification scores for features with low IoU scores during training. The prediction boxes corresponding to the first K encoder features selected by the model according to the classification scores have high classification scores and high IoU scores. The optimization goal is:

[0021]

[0022] in and y represent the prediction and true value respectively, c and b represent the category and bounding box respectively, and Represent the predicted category and predicted bounding box respectively; L box , L cls , L are the bounding box and category loss, and the total loss respectively;

[0023] After optimization of the constrained model, the prediction boxes with high IOU scores and high classification scores enter the decoder.

[0024] Furthermore, the linear combination in the feedforward neural network is: the input is scaled by weights, summed, and then offset by biases. The specific formula is:

[0025] z=w1x1+w2x2+....+w n x n +b

[0026] Where w1 is the weight of data x1, w2 is the weight of data x2, and w n is the data x n The weight of ; b is the bias; z is the linear combination result; n represents the number of features output by the decoder;

[0027] The specific process of BN normalization is:

[0028] 1) Calculate the batch mean μ B and variance

[0029]

[0030] Where m is the number of samples, x i is the sample data;

[0031] 2) Normalization

[0032]

[0033] in, is the normalized data, ∈ is a small positive constant used to prevent numerical instability caused by the denominator being zero;

[0034] The calculation formula of the Softmax activation function is:

[0035]

[0036] Here, e is the base of the natural logarithm, and the sum in the denominator is the sum of the exponential functions of all input elements.

[0037] Furthermore, the FPSbs=1 of the RT-DETR-Sat target detection model is 120, and the average accuracy is above 55%.

[0038] Furthermore, the specific process of obtaining the road image of the inspection area during the road inspection, annotating the road image frame by frame, and dividing the annotated obstacle image into a training set and a test set is:

[0039] (1.1) Using the image acquisition module, the inspection robot obtains the road image of the inspection area during the inspection process;

[0040] (1.2) Convert the format of the road image into a standard format of 1280*1280 to obtain a standard format image;

[0041] (1.3) low-pass filtering the acquired standard format image using a Gaussian function to obtain a filtered image;

[0042] (1.4) Convert the acquired standard format image and the filtered image into the logarithmic domain and perform subtraction to obtain a reflection image in the logarithmic domain;

[0043] (1.5) Perform image summation on the pixel level on the logarithmic domain reflection image to obtain the result;

[0044] (1.6) Color restoration:

[0045] (1.6.1) At the channel level, for each channel, add the pixel values ​​of the channel in the original image to get the sum of the channel, calculate the maximum value of the sum of all channels, and use it as the denominator of the normalization factor; for each channel sum, divide it by the denominator of the normalization factor to get the normalization factor;

[0046] (1.6.2) Normalize the weight matrix so that the weight of each channel is equal and convert it to the logarithmic domain. Then multiply it by the nonlinear factor of color restoration, which is 2.0, and divide it by the normalization factor to get the image color gain result.

[0047] (1.6.3) The result obtained in (1.5) is recombined and multiplied with the result of the weight matrix and the image color gain to obtain the image result after color restoration;

[0048] (1.7) Image restoration: multiply the color restored image result obtained in (1.6.3) by the gain of the image pixel value change range, and add the offset of the image pixel value change range to obtain the final result;

[0049] (2.1) Image data preprocessing: Sliding window technology is introduced. A point is randomly selected as the center point of the square sliding window. A square sliding window of a constant size is used to slide from the center point on the image in the order from left to right and from top to bottom. Each time the sliding window is slid, the image area covered by the sliding window is saved as a sub-image. The size range of the square sliding window is 100-300 pixels, and the sliding step range of the square sliding window is set to 50%-70% of its size. If the side length of the remaining uncovered image in the current row or column is less than the sliding step of the square sliding window, the part exceeding the sliding step is filled with 0, and then the sliding step continues to slide downward with the original sliding step until the square sliding window covers all areas of the image, stops sliding, completes a single sliding, changes the size of the sliding window within the above range, and performs the next sliding. The sliding window operation is repeated at least three times, and all sub-images of different scales are collected.

[0050] (2.2) Preprocess all collected sub-images. First, label them. The sub-images with obstacles are labeled as 1, which is a positive sample; the sub-images without obstacles are labeled as 0, which is a negative sample. Secondly, perform data enhancement on the labeled data set. Finally, the enhanced data set is divided into training set and test set in a ratio of 7:3.

[0051] The process of data enhancement is as follows: the original image is subjected to three Gaussian blur operations, and three differential images are obtained. The difference is that the sigmoid parameters of each Gaussian blur are different, and they are set to 20, 100, and 180 respectively. Finally, the differential images obtained with three different sigmoid parameters are weighted and averaged to obtain the enhanced data set.

[0052] In a second aspect, the present invention provides a road inspection robot obstacle detection system based on RT-DETR-Sat, the system comprising:

[0053] An image acquisition module is used to obtain road images of the inspection area during the road inspection process;

[0054] The image processing module is used to label the obstacle images of the image acquisition module according to their categories, obtain a data set, and divide it into a training set and a test set;

[0055] The obstacle images include trees, rocks, construction materials, road maintenance facilities, and animals;

[0056] The RT-DETR-Sat target detection model is used to detect obstacles. It is connected with the image acquisition module and the image processing module, and the RT-DETR-Sat target detection model is trained using the acquired data set. The trained RT-DETR-Sat target detection model is used to detect obstacles.

[0057] The early warning and feedback module uploads the output results of the RT-DETR-Sat target detection model. If there are obstacles in the specified area, an alarm will be issued and the staff will be reminded to clear the obstacles in time.

[0058] Compared with the prior art, the present invention has the following beneficial effects:

[0059] The system of the present invention creatively applies the RT-DETR-Sat target detection model of real-time end-to-end detection to road detection, and realizes the detection of obstacles such as trees, rocks, construction materials, road maintenance facilities, animals, etc. by road inspection robots. It can accurately and efficiently detect obstacles during the inspection process and remind staff to deal with them in time. The introduction of the group query attention mechanism in the hybrid encoder and decoder helps the real-time performance and accuracy of road detection.

[0060] The present invention adopts an end-to-end training method and regards the detection task as a direct set prediction problem. It does not require a region candidate stage and can directly learn the position and category of the output object from the input image. It solves the limitations of existing real-time detectors and end-to-end detectors, improves the generalization ability of the model and the target detection ability under different environments and conditions, accelerates the model convergence process, and improves the overall accuracy of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] Figure 1 It is a schematic diagram of the model training process of the road inspection robot obstacle detection method based on RT-DETR-Sat of the present invention.

[0062] Figure 2 It is a structural schematic diagram of the RT-DETR-Sat target detection model in the present invention.

[0063] Figure 3 It is a schematic diagram of the structure of the backbone network ResNeSat in the present invention.

[0064] Figure 4 It is a flow chart of the intra-scale feature interaction module AIFI of the present invention.

[0065] Figure 5 It is a flow chart of the mechanism of IoU-aware query selection in the present invention.

[0066] Figure 6It is a structural schematic diagram of the decoder in the present invention.

[0067] Figure 7 It is a schematic diagram of the structure of the feedforward neural network in the present invention. DETAILED DESCRIPTION

[0068] In order to more clearly describe the technical problems, technical solutions and advantages of the present invention, the following will be described in detail with reference to the drawings and embodiments. It should be noted that these embodiments are only used to illustrate the principles and application scope of the present invention and should not be regarded as limiting the present invention.

[0069] The road inspection robot obstacle detection method based on RT-DETR-Sat of the present invention comprises the following steps:

[0070] The detection method comprises the following steps:

[0071] Obtain road images of the inspection area during the road inspection process, annotate the road images frame by frame, and divide the annotated obstacle images into a training set and a test set;

[0072] Construct an RT-DETR-Sat target detection model, which includes a backbone network ResNeSat, a hybrid encoder and a decoder with an auxiliary prediction head. The image is input into the backbone network ResNeSat, and feature extraction is performed to obtain multi-layer feature maps of different scales. The multi-scale features are then processed using the hybrid encoder. The output of the hybrid encoder is connected to the IoU-aware query selection, and a fixed number of features are selected as the initial target query of the decoder. The decoder then generates a bounding box and a confidence score. Finally, the output of the model is converted into a probability distribution, thereby obtaining a final prediction result, and then judging whether there is an obstacle.

[0073] The road image is used as the input of the backbone network ResNeSat. It first passes through a convolution layer with a 7*7 convolution kernel, and the number of output channels of the convolution layer is 64. Then it passes through a 3*3 pooling layer. After pooling, the data enters four convolution groups. These four convolution groups are composed of two 1*1 convolution kernels and one 3*3 convolution kernel. These convolution groups are repeated for different times. The first convolution group is repeated 3 times, and the number of output channels is 256. The second convolution group is repeated 4 times, and the number of output channels is 512. The third convolution group is repeated 6 times, and the number of output channels is 1024. The fourth convolution group is repeated 3 times, and the number of output channels is 2048. After that, the self-attention mechanism is used to obtain output features with step sizes of 8, 16, and 32, which are recorded as S3, S4, and S5 respectively. S3, S4, and S5 are used as input features of the hybrid encoder.

[0074] The hybrid encoder includes an intra-scale feature interaction module AIFI and a multi-scale feature fusion module CCFM. The input feature S5 is used as the input of the intra-scale feature interaction module AIFI (encoder), and F5 is obtained after intra-scale interaction.

[0075] S3, S4, and F5 are used as the input of the multi-scale feature fusion module CCFM to perform multi-scale feature fusion. The specific processing process of multi-scale feature fusion is:

[0076] The input features of F5 are processed by 1*1 convolution, IN normalization and activation function ReLU calculation, and the input features of S4 are processed by 1*1 convolution respectively, and then the elements are added together to obtain the first fusion result;

[0077] The first fusion result is processed by 1*1 convolution, IN normalization and activation function, and then concatenated with the S3 input feature by 1*1 convolution and element-wise addition to obtain the second fusion result. The output of the first fusion result is processed by activation function and then connected with the 1*1 convolution of the branch where the S4 input feature is located.

[0078] The second fusion result is processed by 3*3 convolution, IN normalization and activation function, and then connected with the 1*1 convolution of the branch where the F5 input feature is located in the first fusion result; the first fusion result is processed by 3*3 convolution, IN normalization and activation function, and then processed by 1*1 convolution with the activation function processing result of the branch where the F5 input feature is located, and then the feature splicing of element addition is obtained to obtain the third fusion result;

[0079] The first fusion result, the second fusion result and the third fusion result are concatenated to obtain the output of the hybrid encoder;

[0080] The decoder includes a group query attention mechanism, a residual connection and a normalization process, a group query attention mechanism, a residual connection and a normalization process, a multi-layer perceptron, a residual connection and a normalization process, which are connected in sequence;

[0081] The output of the decoder is used as the input of a feedforward neural network, wherein the feedforward neural network includes a linear combination, a BN normalization process, and a softmax activation function;

[0082] The training set is used to train the RT-DETR-Sat target detection model, which is then used for obstacle detection by road inspection robots.

[0083] Among them, the calculation formula of the convolution group is as follows:

[0084]

[0085] Among them, the parameter C represents the number of channels, T i Represents the result of three convolutions for each channel, x is the input vector, vector x is projected into an optional low-dimensional integration, and then transformed; F(x) represents the output result of a convolution group.

[0086] The backbone network ResNeSat of the present invention can achieve an accuracy comparable to or even better than that of a deeper ResNet network with the same number of parameters; in terms of computing resource requirements, the computational complexity of ResNeSat is only about half of that of ResNet with the same number of parameters, which makes ResNeSat more efficient in road inspections; at the same time, ResNeSat enables the model to have better generalization capabilities, helping the model to maintain good performance when facing different data sets; ResNeSat reduces the number of hyperparameters that need to be manually adjusted through the concepts of modularization and repeated layers, thereby simplifying model design.

[0087] S5 is obtained by the intra-scale feature interaction module AIFI to obtain F5 (see Figure 4 ), the specific transformation formula is as follows:

[0088] Q=K=V=flatten(S5)

[0089] F5=Reshape(Attn(Q,K,V))

[0090] Among them, flatten is the process of converting the multi-dimensional array of S5 image features into a one-dimensional array, that is, merging each row or column of the S5 image features into a separate sequence, which is the flattening feature operation; Attn refers to processing the vectors Q, K and V with a multi-head self-attention mechanism, and Reshape refers to the deformation operation, which is to reshape the result of the multi-head self-attention mechanism into a multi-dimensional F5 image feature, that is, to reconstruct the feature map.

[0091] S3, S4, and F5 are used as the input of the multi-scale feature fusion module CCFM to perform multi-scale feature fusion. The specific process is (see Figure 2 ):

[0092] After 1*1 convolution, IN normalization and activation function calculation, F5 is used together with S4 as the input of fusion block 1. After fusion, the first fusion result of fusion block 1 is obtained. The mechanism of fusion block is as follows:

[0093] Features of different scales are used as the input of the fusion block and are first convolved with a 1*1 convolution kernel. Then, the features are concatenated by element-wise addition to obtain a new feature map as the output of the fusion block.

[0094] The first fusion result is processed by 1*1 convolution, IN normalization and activation function, and then used together with S3 as the input of fusion block 2 to obtain the output of fusion block 2, which is recorded as the second fusion result. The output of fusion block 2 is processed again by image processing (3*3 convolution, IN normalization and activation function) and then used as the input of fusion block 1 again. After fusion, the second output of fusion block 1 is obtained. The second output of fusion block 1 is processed by image processing (3*3 convolution, IN normalization and activation function) and F5 is processed by image processing (1*1 convolution, IN normalization and activation function) as the input of fusion block 3. After fusion, the output of fusion block 3 is obtained, which is recorded as the third fusion result. Finally, the second output of fusion block 1, the output of fusion block 2 and the output of fusion block 3 are concatenated by element addition to obtain a new feature map, and the output output of the hybrid encoder is obtained. The formula is as follows:

[0095] Output = CCFM ({S3, S4, F5})

[0096] The calculation formula of IN normalization technology is as follows:

[0097] 1. Data standardization

[0098] x_normalized = (x-mean) / std

[0099] x_normalized is the data after standardization, x is the original data, mean is the mean of the data, and std is the standard deviation of the data.

[0100] 2. Data Scaling

[0101] x_scaled=x_normalized*scale+shift

[0102] Where x_scaled is the data obtained by data scaling, scale is the scaling factor, and shift is the offset. For the range of [-1,1], scale = 1 / 2, shift = 0; for the range of [0,1], scale = 1 / 2, shift = -1 / 2.

[0103] 3. Calculate the mean and standard deviation

[0104] mean_new=decay*mean_old+(1-decay)*mean_batch

[0105] std_new=decay*std_old+(1-decay)*std_batch

[0106] Among them, mean_new is the calculated mean, std_new is the calculated standard deviation, decay is a decay factor close to 1, mean_batch and std_batch are the mean and standard deviation of the current batch respectively.

[0107] 4. Apply normalization

[0108] x_normalized_new=(x_scaled-mean_new) / std_new

[0109] x_scaled_new=x_normalized_new*scale+shift;

[0110] Among them, mean_new is the calculated mean, std_new is the calculated standard deviation x_normalized_new, x_scaled is the data obtained by data scaling, x_normalized_new is the standardized data after scaling, and x_scaled_new is the data obtained after normalization.

[0111] The multiple use of IN normalization technology in the hybrid encoder of the present invention can well learn the characteristics of different obstacles, which is conducive to the inspection robot to identify different obstacles more efficiently.

[0112] After IN normalization technology, the activation function ReLU is used to calculate higher-level abstract features. The ReLU calculation formula is as follows:

[0113] f(x)=max(0,x)

[0114] Among them, x is the value of the neuron after linear transformation.

[0115] The CCFM module in the present invention does not occupy more computing resources due to the increase in batch size, and by normalizing on each instance, it helps to more stably propagate the gradient, and provides a constant gradient in the positive interval, which helps to avoid the gradient vanishing problem.

[0116] A fixed number of image features are selected from the hybrid encoder output sequence as the initial object query for the decoder. IoU-aware query selection is achieved by constraining the model to produce high classification scores for features with high IoU scores and low classification scores for features with low IoU scores during training. Therefore, the prediction boxes corresponding to the first K hybrid encoder features selected by the model based on the classification scores have high classification scores and high IoU scores. The optimization goal is as follows:

[0117]

[0118] in and y represent the prediction and true value respectively, c and b represent the category and bounding box respectively, and Represent the predicted category and predicted bounding box respectively; and y represent the prediction and true value respectively, c and b represent the category and bounding box respectively, and Represent the predicted category and predicted bounding box respectively; L box , L cls , L are the bounding box loss, category loss and total loss respectively.

[0119] After optimization of the constraint model, the prediction boxes with high IOU scores and high classification scores enter the decoder. The decoder first calculates the group query attention mechanism. The calculation formula is as follows:

[0120] score=torch.matmul(query_group,key.transpose(-2,-1))

[0121] attention_scores=torch.softmax(score,dim=-1)

[0122] Among them, query_group is the matrix of the query group, key is the matrix of the key, transpose(-2,-1) means transposing on the last two dimensions to ensure that the dimensions of the matrix multiplication match, torch.matmul means multiplying the query group matrix query_group and the key matrix key and assigning the result to the variable score. dim=-1 means applying the torch.softmax function on the last dimension to obtain the normalized attention score attention_scores.

[0123] The decoder in the present invention reduces the memory usage by reducing the number of key and value matrices in the self-attention mechanism, so that the model can run more efficiently under limited memory resources; at the same time, the improved decoder reduces the memory usage of the model weights and reduces the space usage of KV-Cache, thereby improving the read and write efficiency of the cache.

[0124] Then enter the Add&Norm layer, which is the residual connection and normalization processing. The processing formula is as follows:

[0125] LayerNorm(X+GroupedQueryAttention(X))

[0126] Where X represents the input of Grouped QueryAttention, GroupedQueryAttention(X) represents the output; LayerNorm represents normalization, and GroupedQueryAttention represents the grouped query attention mechanism.

[0127] The processed data is once again processed through the group query attention mechanism, residual connection and normalization, and then enters the multi-layer perceptron. The multi-layer perceptron is a two-layer fully connected layer. The activation function of the first layer is Relu, and the second layer does not use an activation function. The corresponding formula is as follows:

[0128] y=ω·x+b

[0129] Where ω is a learnable parameter, b is the bias parameter, and x is the input data.

[0130] The output of the multi-layer perceptron is then subjected to residual connection and normalization to obtain the output of the entire decoder, the precise bounding box and its corresponding confidence score.

[0131] The linear combination in the feedforward neural network is: the input is scaled by weights, summed, and then offset by biases. The specific formula is:

[0132] z=w1x1+w2x2+....+w n x n +b

[0133] Among them, w is the weight of data x, w1 is the weight of data x1, w2 is the weight of data x2, and w n is the data x n The weight of ; b is the bias; z is the linear combination result; n represents the number of elements in the data x (a feature of the decoder output);

[0134] The specific process of BN normalization is:

[0135] 1) Calculate the batch mean μ B and variance

[0136]

[0137] Where m is the number of samples, x i is the sample data;

[0138] 2) Normalization

[0139]

[0140] in, is the normalized data, ∈ is a small positive constant used to prevent numerical instability caused by the denominator being zero;

[0141] The calculation formula of the Softmax activation function is:

[0142]

[0143] Wherein, e is the base of the natural logarithm, the sum in the denominator is the sum of the exponential functions of all input elements, j = 1, 2, ..., n.

[0144] After adding BN normalization processing to the feedforward neural network in the present invention, the input data is normalized to a stable range, which helps to avoid the problem of gradient vanishing or gradient exploding, keeps the gradient within a suitable range, is conducive to the stability of the model, can process time-dependent data, and improve the training effect.

[0145] Example 1

[0146] The road inspection robot obstacle detection system of this embodiment based on RT-DETR-Sat includes:

[0147] An image acquisition module is used to obtain road images of the inspection area during the road inspection process;

[0148] The image processing module is used to label the obstacle images of the image acquisition module according to their categories, obtain a data set, and divide it into a training set and a test set;

[0149] The obstacle images include trees, rocks, construction materials, road maintenance facilities, animals and other obstacle images;

[0150] The RT-DETR-Sat target detection model is used to detect obstacles. It is connected with the image acquisition module and the image processing module, and the obtained data set is used to apply the RT-DETR-Sat target detection model for obstacle detection and train it. The trained RT-DETR-Sat target detection model can detect obstacles.

[0151] The early warning and feedback module uploads the output results of the RT-DETR-Sat target detection model. If there are obstacles in the specified area, an alarm will be issued and the staff will be reminded to clear the obstacles in time.

[0152] The training process of the RT-DETR-Sat target detection model is as follows: start with random initial RT-DETR-Sat network data, load the training set image, perform image preprocessing, and input the preprocessed image into the RT-DETR-Sat target detection model to predict the category of each pixel in the image and calculate the loss function error; the training gradient is close to 0, and the network parameters of the trained RT-DETR-Sat target detection model are obtained. If the gradient is not close to 0, the error back propagation is performed to adjust the model parameters, and the loaded training set image is returned.

[0153] The test set is input into the trained RT-DETR-Sat target detection model, and the trained RT-DET R-Sat target detection model can realize the detection of obstacles; the trained RT-DETR-Sat target detection model is used to obtain the category prediction of each pixel in the image, and the prediction results are output, including the location of the obstacle.

[0154] The video image to be identified is input into the trained RT-DETR-Sat target detection model to detect whether there are any obstacles in the 8-meter-long and 2-meter-wide area directly in front of the inspection robot. If there are obstacles, the inspection robot will issue an alarm to remind the staff to clear them in time until the inspection robot completes the road inspection.

[0155] Example 2

[0156] In this embodiment, the obstacle images acquired during the inspection process specifically include: images of trees, rocks, construction materials, road maintenance facilities, animals, etc.

[0157] (1) Using the image acquisition module, the inspection robot obtains the road image of the inspection area during the inspection process;

[0158] (1.1) Using the image acquisition module, the inspection robot obtains the road image of the inspection area during the inspection process;

[0159] (1.2) Convert the format of the road image into a standard format of 1280*1280 to obtain a standard format image;

[0160] (1.3) low-pass filtering the acquired standard format image using a Gaussian function to obtain a filtered image;

[0161] (1.4) Convert the acquired standard format image and the filtered image into the logarithmic domain and perform a difference (i.e., subtract the low-frequency components in the image) to obtain a reflection image in the logarithmic domain;

[0162] (1.5) Perform image summation on the pixel level on the logarithmic domain reflection image to obtain the result;

[0163] (1.6) Color restoration:

[0164] (1.6.1) At the channel level, for each channel, add the pixel values ​​of the channel in the original image to get the sum of the channel, calculate the maximum value of the sum of all channels, and use it as the denominator of the normalization factor; for each channel sum, divide it by the denominator of the normalization factor to get the normalization factor;

[0165] (1.6.2) Normalize the weight matrix (the weight matrix is ​​set manually according to different applications. Here, a uniform weight matrix is ​​used, that is, the weight of each channel is equal). The weight of each channel is equal and converted to the logarithmic domain. Then, it is multiplied by the nonlinear factor of color restoration, which is 2.0, and then divided by the normalization factor to obtain the result of image color gain.

[0166] (1.6.3) The result obtained in (1.5) is recombined and multiplied with the result of the weight matrix and the image color gain to obtain the image result after color restoration;

[0167] (1.7) Image restoration: multiply the color restored image result obtained in (1.6.3) by the gain of the image pixel value change range, and add the offset of the image pixel value change range to obtain the final result;

[0168] The gain of the image pixel value change range is a multiplication factor used to adjust the brightness and contrast of the image. It is applied to each pixel value of the image result after color restoration. The offset of the image pixel value change range is a constant value used to shift the image after the multiplication factor. It is added to each pixel value of the image result after color restoration. The final image restoration result is obtained by multiplying the multiplication factor by the image result after color restoration and adding the offset to the multiplication result.

[0169] (2.1) Image data preprocessing: Sliding window technology is introduced. A point in the image is randomly selected as the center point of the square sliding window. A square sliding window of constant size is used to slide from the center point on the image in the order from left to right and from top to bottom. Each time the sliding window is slid, the image area covered by the sliding window is saved as a sub-image. The size range of the square sliding window is 100-300 pixels, and the sliding step range of the square sliding window is set to 50%-70% of its size. If the side length of the remaining uncovered image in the current row or column is smaller than the square sliding window, the image is saved as a sub-image. If the sliding step size is larger than the sliding step size, the part exceeding the sliding step size is filled with 0, and then the sliding continues downward with the original sliding step size until the square sliding window covers all areas of the image, and the sliding stops, completing a single sliding. The size of the sliding window is changed within the above range, and the next sliding is performed. The sliding window operation is repeated three times, and all sub-images of different scales are collected. In this embodiment, the original image size is 1080*1920, and the window sizes of the three sliding window operations are 100, 150, and 200 pixels respectively. The sliding step size value range of the square sliding window is set to 60% of its size.

[0170] (2.2) Preprocess all collected sub-images. First, label them. The sub-images with obstacles are labeled as 1, which is a positive sample; the sub-images without obstacles are labeled as 0, which is a negative sample. Secondly, perform data enhancement on the labeled data set. Finally, the enhanced data set is divided into training set and test set in a ratio of 7:3.

[0171] The process of data enhancement is as follows: the original image is repeatedly processed by Gaussian blur operation three times to obtain three difference images. The difference is that the sigmoid parameters of each Gaussian blur are different, which are set to 20, 100, and 180 respectively; finally, the difference images obtained by the three different sigmoid parameters are weighted and averaged to obtain the enhanced data set.

[0172] (3) Read the annotated training set into the RT-DETR-Sat target detection model, train the real-time end-to-end target detector, and iterate all images in turn. After 40,000 iterations, the training model reaches convergence, that is, when the model training gradient is close to 0 (less than 0.01 can be considered close to 0), stop training, extract the optimal network parameters for prediction, and adjust the weight parameters if the training gradient is not close to 0. Model calibration stage: 1) Input the test set data into the automatic obstacle detection model to obtain the detection results; 2) Compare the actual parameters of the detection target used for the test with the preliminary detection results to obtain the calibrated RT-DETR-Sat model; When predicting, the model first loads the trained parameters and loads the input image from the test set. The trained RT-DETR-Sat target detection model is used to calculate the category of each pixel, thereby realizing obstacle detection;

[0173] (4) Implementation phase:

[0174] (4.1) The inspection robot is equipped with an industrial camera with a resolution of 1080P and a sensor pixel of 2 million. The camera is located in the middle of the top of the inspection robot, 1.6 meters above the ground. The camera is tilted toward the ground and the angle formed with the ground is 10°. During inspection, the industrial camera will capture images of the area 8 meters long and 2 meters wide directly in front of the inspection robot.

[0175] (4.2) Input the obstacle image of the inspection area to be detected during the inspection process into the trained RT-DETR-Sat target detection model;

[0176] (4.3) If the image is identified as one without obstacles, the inspection robot continues the inspection;

[0177] (4.4) If an obstacle image is identified, the inspection robot will sound an alarm to remind the staff to clear the obstacle in time until the inspection is completed.

[0178] Example 3

[0179] The hardware devices used in the detection system of this embodiment include the following components:

[0180] Processor: As the core component of the present invention, the processor is responsible for controlling and managing the operation of the entire system, including data acquisition, data processing, image recognition and other functions, and needs to have sufficient computing power and parallel processing capabilities to meet real-time requirements. The processor can be in different forms such as single-chip microcomputers, microprocessors, computers, etc. to meet the needs of different application scenarios;

[0181] Memory: Memory can be used to store collected data and historical data for subsequent processing and analysis. It has the characteristics of high speed, high reliability and scalability to meet the needs of long-term stable operation of the system;

[0182] Database: Use database to store and manage collected data, historical data, analysis results and other information;

[0183] Network interface: used for data exchange and communication, with the characteristics of high speed, high stability and high security, ensuring the reliability and security of data transmission.

[0184] The processor is configured to execute computer executable instructions. When the computer executable instructions are executed by the processor, the various steps of the above-mentioned road inspection robot obstacle detection method based on RT-DETR-Sat are implemented.

[0185] A computer program is stored in the memory, and the computer program can be executed by the processor to implement the various steps of the above-mentioned road inspection robot obstacle detection method based on RT-DETR-Sat.

[0186] The database is configured to store and manage computer application data, including various data types and structures, which are applied to the various steps of the above-mentioned RT-DETR-Sat-based road inspection robot obstacle detection method.

[0187] The network interface realizes communication and data transmission between computers. The network interface can provide various communication protocols and data transmission methods to meet the communication and data transmission requirements of different application scenarios and different needs, and is applied to each step of the above-mentioned RT-DETR-Sat-based road inspection robot obstacle detection method.

[0188] The present invention is mainly used for obstacle detection during the inspection process of an inspection robot in an assembled factory, and utilizes an industrial camera installed on the inspection robot to automatically identify obstacles during the inspection process.

[0189] Example 4

[0190] This embodiment is based on the road inspection robot obstacle detection system of the ResNeSat model. The RT-DETR-Sat neural network consists of a backbone network ResNeSat, a hybrid encoder and a decoder with an auxiliary prediction head. The image is input into the backbone network ResNeSat, and the output characteristics with step sizes of 8, 16, and 32 are obtained as the input of the hybrid encoder, and S3, S4, and S5 are used to mark them respectively. S5 is used as the input of the AIFI module, and F5 is obtained through intra-scale interaction. S3, S4, and F5 are used as the input of the CCFM module for multi-scale feature fusion. The processing mechanism of image features is as follows:

[0191] The image features are first convolved with the convolution kernel to further obtain new image features, normalized by the normalization technology IN, and then calculated by the activation function ReLU to obtain higher-level abstract features. The image feature F5 is convolved with 1*1 convolution, normalized, and activated together with S4 as the input of fusion block 1. After fusion, the first output of fusion block 1 (the first fusion result) is obtained. The mechanism of the fusion block is as follows: the features of different scales are first convolved with 1*1 convolution kernel as the input of the fusion block, and then fused by element addition to obtain a new feature map as the output of the fusion block.

[0192] The first fusion result is processed by 1*1 convolution, IN normalization and activation function, and then used together with S3 as the input of fusion block 2 to obtain the output of fusion block 2, which is recorded as the second fusion result. The output of fusion block 2 is processed again by image processing (3*3 convolution, IN normalization and activation function) and then used as the input of fusion block 1 again. After fusion, the second output of fusion block 1 is obtained. The second output of fusion block 1 is processed by image processing (3*3 convolution, IN normalization and activation function) and F5 is processed by image processing (1*1 convolution, IN normalization and activation function) as the input of fusion block 3. After fusion, the output of fusion block 3 is obtained, which is recorded as the third fusion result. Finally, the second output of fusion block 1, the output of fusion block 2 and the output of fusion block 3 are concatenated by element addition to obtain a new feature map.

[0193] A fixed number of image features are selected from the hybrid encoder output sequence as the initial object query for the decoder. IoU-aware query selection is achieved by constraining the model to produce high classification scores for features with high IoU scores and low classification scores for features with low IoU scores during training. Therefore, the prediction boxes corresponding to the first K encoder features selected by the model based on the classification score have high classification scores and high IoU scores. After optimization of the constrained model, the prediction boxes with high IOU scores and high classification scores enter the decoder. The decoder first undergoes the calculation of the group query attention mechanism, and then enters the Add&Norm layer, that is, residual connection and normalization processing. The processed data is once again processed by the group query attention mechanism and residual connection and normalization processing, and enters the multi-layer perceptron. The perceptron has a two-layer fully connected layer. The activation function of the first layer is Relu, and the second layer does not use an activation function. The output of the perceptron is then residually connected and normalized to obtain the output of the entire decoder, the precise bounding box and its corresponding confidence score. The output of the decoder is used as the input of the feedforward neural network. The data is first linearly combined in the feedforward neural network, and the normalized data is then nonlinearly activated. The output of the model is converted into a probability distribution through the activation function softmax, thereby obtaining the final prediction result.

[0194] The hybrid encoder consists of an intra-scale feature interaction module AIFI and a multi-scale feature fusion module CCFM. Intra-scale interaction is performed in the AIFI module, and multi-scale feature fusion is performed in the CCFM module.

[0195] In the decoder, the predicted box after the IOU perception query is input into the decoder, and firstly calculated by the group query attention mechanism. The calculation formula is as follows:

[0196] score=torch.matmul(query_group,key.transpose(-2,-1))

[0197] attention_scores=torch.softmax(score,dim=-1)

[0198] Among them, query_group is the matrix of the query group, key is the matrix of the key, transpose(-2,-1) means transposing on the last two dimensions to ensure that the dimensions of the matrix multiplication match, and torch.matmul means multiplying the query group matrix query_group and the key matrix key, and assigning the result to the variable score. dim=-1 means applying the torch.softmax function on the last dimension to get the normalized attention score attention_scores. Then enter the Add&Norm layer, which is the residual connection and normalization processing. The processing formula is as follows:

[0199] LayerNorm(X+GroupedQueryAttention(X))

[0200] Where X represents the input of Grouped QueryAttention, and GroupedQueryAttention(X) represents the output. The processed data is once again processed by the grouped query attention mechanism, residual connection, and normalization, and then enters the multi-layer perceptron. The multi-layer perceptron has a two-layer fully connected layer. The activation function of the first layer is Relu, and the second layer does not use an activation function. The corresponding formula is as follows:

[0201] y=ω·x+b

[0202] Where ω is a learnable parameter, b is the bias parameter, and x is the input data.

[0203] The output of the multi-layer perceptron is then subjected to residual connection and normalization to obtain the output of the entire decoder, the precise bounding box and its corresponding confidence score.

[0204] The RT-DETR-Sat network of the present invention greatly improves the obstacle detection rate, enabling the system to achieve real-time and high-precision obstacle detection.

[0205] The annotated images are distributed as a training set plus a validation set and a test set in a ratio of 7:3, which not only ensures the amount of training data but also improves the generalization ability of the model. The training set is used to train the model, and the test set is used to evaluate the model performance under different parameter settings, so as to select the best parameter settings and make the model have better generalization ability. The use of the k-fold cross-validation method in the training process can make full use of the limited data set, more accurately evaluate the model performance, select the best model and parameters, and provide estimates of the model variance and bias.

[0206] The difference from other algorithms is that the image data of the present invention is trained and tested on multi-scale images, and the mIOU index is significantly improved. The performance test of different models is carried out using the data set established in Example 2. The test results are shown in the table below:

[0207]

[0208] Among them, FPSbs=1 represents the number of propagations per second. Compared with the existing common target detection algorithms, the present invention not only improves the real-time performance but also ensures the accuracy, so that both real-time performance and accuracy are achieved. The average accuracy under different thresholds is improved, indicating that it has a high detection accuracy for obstacles of different sizes.

[0209] The present invention aims to solve the problems caused by the slow speed and low accuracy of current inspection robots, and automatically identify obstacles during the inspection process through the RT-DETR-Sat target detection model. For obstacles, the system can quickly identify them and automatically alarm to remind the staff to complete the obstacle clearance. Compared with the existing technology, this technical solution has the following advantages and application prospects: improving inspection efficiency and providing the possibility of pursuing more intelligent inspection robots. This is of great significance for automatic inspection and has broad application prospects.

[0210] In the description of this specification, the described specific features, structures, or characteristics may be combined in an appropriate manner in any one or more embodiments or examples.

[0211] It should be understood that each part of the present invention can be implemented by hardware, software or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or hardware stored in a memory and executed by a suitable instruction execution system.

[0212] A person skilled in the art may understand that all or part of the steps in the method for implementing the above-mentioned embodiment may be completed by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, which, when executed, includes one or a combination of the steps of the method embodiment.

[0213] Any matters not described in the present invention are applicable to the prior art.

Claims

1. A road inspection robot obstacle detection method based on RT-DETR-Sat, characterized in that: The detection method comprises the following steps: Obtain road images of the inspection area during the road inspection process, annotate the road images frame by frame, and divide the annotated obstacle images into a training set and a test set; Construct an RT-DETR-Sat target detection model, which includes a backbone network ResNeSat, a hybrid encoder and a decoder with an auxiliary prediction head. The image is input into the backbone network ResNeSat, and feature extraction is performed to obtain multi-layer feature maps of different scales. The multi-scale features are then processed using the hybrid encoder. The output of the hybrid encoder is connected to the IoU-aware query selection, and a fixed number of features are selected as the initial target query of the decoder. The decoder then generates a bounding box and a confidence score. Finally, the output of the model is converted into a probability distribution, thereby obtaining a final prediction result, and then judging whether there is an obstacle. The road image is used as the input of the backbone network ResNeSat. It first passes through a convolution layer with a 7*7 convolution kernel, and the number of output channels of the convolution layer is 64. Then it passes through a 3*3 pooling layer. After pooling, the data enters four convolution groups. These four convolution groups are composed of two 1*1 convolution kernels and one 3*3 convolution kernel. These convolution groups are repeated for different times. The first convolution group is repeated 3 times, and the number of output channels is 256. The second convolution group is repeated 4 times, and the number of output channels is 512. The third convolution group is repeated 6 times, and the number of output channels is 1024. The fourth convolution group is repeated 3 times, and the number of output channels is 2048. After that, the self-attention mechanism is used to obtain output features with step sizes of 8, 16, and 32, which are recorded as S3, S4, and S5 respectively. S3, S4, and S5 are used as input features of the hybrid encoder. The hybrid encoder includes an intra-scale feature interaction module AIFI and a multi-scale feature fusion module CCFM. The input feature S5 is used as the input of the intra-scale feature interaction module AIFI, and F5 is obtained after intra-scale interaction. S3, S4, and F5 are used as the input of the multi-scale feature fusion module CCFM to perform multi-scale feature fusion. The specific processing process of multi-scale feature fusion is: The input features of F5 are processed by 1*1 convolution, IN normalization and activation function ReLU calculation, and the input features of S4 are processed by 1*1 convolution respectively, and then the elements are added together to obtain the first fusion result; The first fusion result is processed by 1*1 convolution, IN normalization and activation function, and then concatenated with the S3 input feature by 1*1 convolution and element-wise addition to obtain the second fusion result. The output of the first fusion result is processed by activation function and then connected with the 1*1 convolution of the branch where the S4 input feature is located. The second fusion result is processed by 3*3 convolution, IN normalization and activation function, and then connected with the 1*1 convolution of the branch where the F5 input feature is located in the first fusion result; the first fusion result is processed by 3*3 convolution, IN normalization and activation function, and then processed by 1*1 convolution with the activation function processing result of the branch where the F5 input feature is located, and then the feature splicing of element addition is obtained to obtain the third fusion result; The first fusion result, the second fusion result and the third fusion result are concatenated to obtain the output of the hybrid encoder; The decoder includes a group query attention mechanism, a residual connection and a normalization process, a group query attention mechanism, a residual connection and a normalization process, a multi-layer perceptron, a residual connection and a normalization process, which are connected in sequence; The output of the decoder is used as the input of a feedforward neural network, wherein the feedforward neural network includes a linear combination, a BN normalization process and a softmax activation function; The training set is used to train the RT-DETR-Sat target detection model, which is then used for obstacle detection by road inspection robots.

2. The detection method according to claim 1, characterized in that: The IoU-aware query selection constrains the model to produce high classification scores for features with high IoU scores and low classification scores for features with low IoU scores during training. The prediction boxes corresponding to the first K encoder features selected by the model based on the classification scores have high classification scores and high IoU scores. The optimization goal is: in and y represent the prediction and true value respectively, c and b represent the category and bounding box respectively, and Represent the predicted category and predicted bounding box respectively; L box , L cls , L are the bounding box and category loss, and the total loss respectively; After optimization of the constrained model, the prediction boxes with high IOU scores and high classification scores enter the decoder.

3. The detection method according to claim 1, characterized in that: The linear combination in the feedforward neural network is: the input is scaled by weights, summed, and then offset by biases. The specific formula is: z=w1x1+w2x2+....+w n x n +b Among them, w1 is the weight of data x1, w2 is the weight of data x2, and w n is the data x n The weight of ; b is the bias; z is the linear combination result; n represents the number of features output by the decoder; The specific process of BN normalization is: 1) Calculate the batch mean μ B and variance Where m is the number of samples, x i is the sample data; 2) Normalization in, is the normalized data, ∈ is a small positive constant used to prevent numerical instability caused by the denominator being zero; The calculation formula of the Softmax activation function is: Here, e is the base of the natural logarithm, and the sum in the denominator is the sum of the exponential functions of all input elements.

4. The detection method according to claim 1, characterized in that: The FPSbs=1 of the RT-DETR-Sat target detection model is 120, and the average accuracy is above 55%.

5. The detection method according to claim 1, characterized in that: The specific process of obtaining the road image of the inspection area during the road inspection, annotating the road image frame by frame, and dividing the annotated obstacle image into a training set and a test set is: (1.1) Using the image acquisition module, the inspection robot obtains the road image of the inspection area during the inspection process; (1.2) Convert the format of the road image into a standard format of 1280*1280 to obtain a standard format image; (1.3) low-pass filtering the acquired standard format image using a Gaussian function to obtain a filtered image; (1.4) Convert the acquired standard format image and the filtered image into the logarithmic domain and perform subtraction to obtain a reflection image in the logarithmic domain; (1.5) Perform image summation on the pixel level on the logarithmic domain reflection image to obtain the result; (1.6) Color restoration: (1.6.1) At the channel level, for each channel, add the pixel values ​​of the channel in the original image to get the sum of the channel, calculate the maximum value of the sum of all channels, and use it as the denominator of the normalization factor; for each channel sum, divide it by the denominator of the normalization factor to get the normalization factor; (1.6.2) Normalize the weight matrix so that the weight of each channel is equal and convert it to the logarithmic domain. Then multiply it by the nonlinear factor of color restoration, which is 2.0, and divide it by the normalization factor to get the image color gain result. (1.6.3) The result obtained in (1.5) is recombined and multiplied with the result of the weight matrix and the image color gain to obtain the image result after color restoration; (1.7) Image restoration: multiply the color restored image result obtained in (1.6.3) by the gain of the image pixel value change range, and add the offset of the image pixel value change range to obtain the final result; (2.1) Image data preprocessing: Sliding window technology is introduced. A point is randomly selected as the center point of the square sliding window. A square sliding window of a constant size is used to slide from the center point on the image in the order from left to right and from top to bottom. Each time the sliding window is slid, the image area covered by the sliding window is saved as a sub-image. The size range of the square sliding window is 100-300 pixels, and the sliding step range of the square sliding window is set to 50%-70% of its size. If the side length of the remaining uncovered image in the current row or column is less than the sliding step of the square sliding window, the part exceeding the sliding step is filled with 0, and then the sliding step continues to slide downward with the original sliding step until the square sliding window covers all areas of the image, stops sliding, completes a single sliding, changes the size of the sliding window within the above range, and performs the next sliding. The sliding window operation is repeated at least three times, and all sub-images of different scales are collected. (2.2) Preprocess all collected sub-images. First, label them. The sub-images with obstacles are labeled as 1, which is a positive sample; the sub-images without obstacles are labeled as 0, which is a negative sample. Secondly, perform data enhancement on the labeled data set. Finally, the enhanced data set is divided into training set and test set in a ratio of 7:

3. The process of data enhancement is as follows: the original image is subjected to three Gaussian blur operations, and three differential images are obtained. The difference is that the sigmoid parameters of each Gaussian blur are different, and they are set to 20, 100, and 180 respectively. Finally, the differential images obtained with three different sigmoid parameters are weighted and averaged to obtain the enhanced data set.

6. A road inspection robot obstacle detection system based on RT-DETR-Sat, characterized in that: The system executes the method according to any one of claims 1 to 5, including: An image acquisition module is used to obtain road images of the inspection area during the road inspection process; The image processing module is used to label the obstacle images of the image acquisition module according to their categories, obtain a data set, and divide it into a training set and a test set; The obstacle images include trees, rocks, construction materials, road maintenance facilities, and animals; The RT-DETR-Sat target detection model is used to detect obstacles. It is connected with the image acquisition module and the image processing module, and the RT-DETR-Sat target detection model is trained using the acquired data set. The trained RT-DETR-Sat target detection model is used to detect obstacles. The early warning and feedback module uploads the output results of the RT-DETR-Sat target detection model. If there are obstacles in the specified area, an alarm will be issued and the staff will be reminded to clear the obstacles in time.

Citation Information

Patent Citations

  • Small target detection method based on attention and context awareness

    CN115937736A

  • Remote sensing image change detection method combining U-shaped network and self-attention mechanism

    CN116740527A