Obstacle Avoidance Method and System for Road Inspection Robot Based on Instance Segmentation

By adopting a real-time end-to-end object detection model based on instance segmentation in the inspection robot, the problem of being unable to accurately identify and segment overlapping obstacles in the prior art is solved, and rapid and high-precision detection and automatic obstacle avoidance of obstacles in road inspection are achieved.

CN119672330BActive Publication Date: 2025-06-13EAST CHINA JIAOTONG UNIVERSITY +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411620897.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-14
Publication Date
2025-06-13
Estimated Expiration
2044-11-14

AI Technical Summary

Technical Problem

The existing patrol robot obstacle avoidance methods cannot accurately identify and divide overlapping or contact obstacles, making it difficult to accurately adjust the route obstacle avoidance.

Method used

The road patrol robot obstacle avoidance method based on instance segmentation is adopted. By building a real-time end-to-end object detection model, using the VIM model and CBAM attention mechanism, real-time instance segmentation and width measurement of obstacles are realized, and the inspection route is automatically adjusted.

Benefits of technology

It realizes rapid and high-precision detection of obstacles such as trees, rocks, construction materials, road maintenance facilities, etc. during road inspections, especially real-time detection and obstacle avoidance can be carried out more accurately and efficiently for moving obstacles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119672330B_ABST
    Figure CN119672330B_ABST
Patent Text Reader

Abstract

The present invention relates to an obstacle avoidance method and system for a road inspection robot based on instance segmentation. The method includes the following steps: obtaining a real-time video of an inspection area during road inspection, and performing obstacle annotation frame by frame on the road images in the real-time video, and dividing the annotated obstacle images into a training set and a test set; constructing a real-time end-to-end object detection model, the real-time end-to-end object detection model includes an image encoder module, a prompt encoder module, a mask decoder module, a memory attention module, a memory decoder module, and a memory bank; using the training set to train the real-time end-to-end object detection model to obtain a trained real-time end-to-end object detection model for performing frame-by-frame real-time detection on each frame image in the real-time video of the inspection area during road inspection. The present invention can more accurately and efficiently detect obstacles in real time during the inspection process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of prefabricated factories and instance segmentation, and particularly to an obstacle avoidance method and system for road inspection robots based on instance segmentation Background Art

[0002] In prefabricated factories, obstacles on factory roads can affect the transportation safety of prefabricated components. Inspection can not only timely identify and handle potential safety hazards, avoid causing personal injuries or property losses, but also help ensure that the functional performance of prefabricated roads meets the design requirements, improve the overall usage efficiency and the comfort of vehicle driving. Conventional manual inspection has the characteristics of low work efficiency and high cost. At the same time, manual detection has low accuracy and cannot be quantified. The results of inspection are easily affected by personal experience, skills and the working state of the day, resulting in low stability and reliability of the detection results. At the same time, manual inspection also has the problem of low cost-effectiveness. At the same time, there are safety risks in manual inspection in dangerous areas

[0003] Compared with traditional manual inspection, inspection robots can work in dangerous environments such as high temperature, high pressure and strong electricity, avoiding the risk of causing personal injuries to inspection personnel and reducing the production safety risks caused by improper equipment operation. The traditional obstacle avoidance of inspection robots uses object detection algorithms, which only provide rectangular bounding boxes to locate objects, that is, only the target position can be located through the detection box, and the object dimensions cannot be accurately determined, and the specific shape information of the object cannot be provided. The method of the present invention can more accurately segment the contour of the object through a new instance segmentation, can distinguish and separately mark the overlapping or contacting objects in the image, and avoids the deficiency that object detection may regard them as a whole when dealing with overlapping objects and it is difficult to distinguish individual instances Summary of the Invention

[0004] The object of the present invention is to overcome the above-mentioned technical drawbacks and provide an obstacle avoidance method and system for road inspection robots based on instance segmentation with stronger real-time performance and higher accuracy

[0005] To solve the above-mentioned technical problems, the technical solution of the present invention is as follows

[0006] In the first aspect, the present invention provides an obstacle avoidance method for a road inspection robot based on instance segmentation, and the method includes the following steps

[0007] Step 1, obtain the real-time video of the inspection area during the road inspection process, and frame by frame label the obstacles in the road images in the real-time video, and divide the labeled obstacle images into a training set and a test set

[0008] Step 2, construct a real-time end-to-end object detection model

[0009] The real-time end-to-end object detection model includes an image encoder module, a prompt encoder module, a mask decoder module, a memory attention module, a memory decoder module, and a memory bank. The image encoder module uses the VIM model, which has an identified condition. The VIM model is used to segment the input first image into blocks and extract feature vectors for each block of the image. The image encoder outputs a first image embedding and a first identified condition.

[0010] The prompt encoder module includes a convolutional block and CLIP. The mask prompt is input into the convolutional block, and the text prompt is input into CLIP. The prompt encoder module finally outputs an object token and a prompt token, and the object token and the prompt token are combined to form a prompt embedding.

[0011] The prompt embedding is input into the mask decoder. The mask decoder includes a first CBAM attention mechanism layer, a second CBAM attention mechanism layer, a first MLP, a third CBAM attention mechanism layer, a fourth CBAM attention mechanism layer, a convolutional operation, and a second MLP, and outputs an instance segmentation map.

[0012] The instance segmentation map output by the mask decoder is then input into the memory decoder module. After three convolutional operations, a feature map and an object point vector are obtained, and then summed with the first image embedding to combine the feature map and the object point output by the three convolutional operations with the first image embedding to become the first memory feature. The first memory feature is input into the memory bank for storage. At this time, the processing of the first image is completed.

[0013] Then, the processing of the second image begins. The image encoder module outputs a second image embedding and a second identified condition. The second image embedding is input into the memory attention module. The memory attention module includes a CBAM attention mechanism and a cross-attention mechanism connected in sequence. The output result of the CBAM attention mechanism is used as the Query and input into the cross-attention mechanism. At the same time, the first memory feature stored in the memory bank is used as the Key and Value and input into the cross-attention mechanism. After the operation of the cross-attention mechanism, an updated second image embedding is output.

[0014] Finally, the updated second image embedding is respectively input into the second CBAM attention mechanism layer and the third CBAM attention mechanism layer of the mask decoder. The mask decoder outputs an instance segmentation map, which is then input into the memory decoder module. In the memory decoder module, the instance segmentation map output by the mask decoder is combined with the second image embedding after convolutional operation processing to become the second memory feature, and then input into the memory bank for storage. At this time, the processing of the second image is completed.

[0015] Process all the images according to the above-described process;

[0016] The working process of the VIM model is as follows: First, the input image F is divided into multiple non-overlapping blocks of a fixed size. Then, the position information of each image block is embedded into its corresponding pixel vector representation information, and the pixel vector representation information of each image block is flattened and arranged into a sequence to obtain the original pixel sequence information of each image block and the corresponding position information; the original pixel sequence information of each image block is linearly transformed through a linear layer to adjust its data dimension, obtaining the combined embedding features of the primary pixel sequence information of each image block and the corresponding position information; the linear transformation is represented as Y = TW Y +b Y , where T is the original pixel sequence information of the image block, W Y is the weight matrix, b Y is the bias vector, and Y is the primary pixel sequence information obtained after the transformation;

[0017] Then, the primary pixel sequence information of the image block is batch-normalized through a normalization layer;

[0018] After that, the results obtained by processing the primary pixel sequence information Y(B, M, D) of the image block through the normalization layer are respectively projected into sequences X(B, M, E) and Z(B, M, E) with a dimension of E using two different projection layers;

[0019] where B is the batch size; M is the sequence length; D is the original feature dimension; E is the extended state dimension, which is set to 384;

[0020] The two different projection layers are two different linear layers, and these two linear layers project the input sequence to generate sequences X and Z; the linear layer consists of a weight matrix and a bias term. For sequences X and Z, there are respectively weight matrices Wx and Wz and bias vectors bx and bz; multiply the input sequence by the weight matrix Wx and add the bias bx to obtain sequence X,

[0021] X = Wx × Input + bx,

[0022] Input is the input sequence, and Wx is a D×E matrix; the method for obtaining sequence Z is the same. The dimensions of the weight matrix and bias vector of the corresponding linear layer are the same as those of sequence X, but the element values are different;

[0023] Then, sequence X is processed in both the forward and backward directions to capture the local dependencies in the vector. For the forward processing, first apply a one-dimensional forward convolution operation to sequence X to obtain X' o , and then X'o Linearly project it onto three different linear layers to obtain sequence B respectively o :(B, M, N), o:(B, M, N) and Δ o , where sequence Δ o Use the softplus activation function to ensure it is positive;

[0024] Δ o = log(1 + exp(X' o ))

[0025] where N is the dimension of the hidden state;

[0026] Then apply the Hadamard product to sequence Δ o and multiply it with A o and B o respectively according to the following formula to obtain the conversion parameters and

[0027]

[0028] Finally, use sequence C o , conversion parameter conversion parameter and the output X' of the forward convolution o Through the spatio-temporal state model SSM, calculate the forward processing result y of X forward ;

[0029] Apply a one-dimensional backward convolution operation and the spatio-temporal state model to obtain the backward processing result y of sequence X backward ;

[0030] Combine the y forward obtained from sequence X and y backward with the result obtained after the sequence Z is processed by the activation function, and then add the two obtained results together. After the obtained result is processed by the projection layer, perform a residual connection with the primary pixel sequence information Y of the image patch, and then generate the output final encoded pixel sequence information T l , T l will be sent to the multi-layer perceptron MLP for processing, combine T l with the corresponding image patch position information to obtain the image embedding of the image patch, which is the output of the image encoder;

[0031] Step 3, Use the training set to train the real-time end-to-end object detection model. During training, use the mask as the prompt input to the prompt encoder module to obtain the trained real-time end-to-end object detection model, which is used to perform frame-by-frame real-time detection on each frame of the real-time video in the inspection area during road inspection.

[0032] Further, the mask decoder module is used to combine the image embedding output by the image encoder and the target tokens and prompt tokens output by the prompt encoder as inputs, and finally output the predicted segmentation mask. Its specific working process is as follows:

[0033] First, combine the target tokens and prompt tokens into a prompt embedding F 0 , and use F 0 as Query, key, and value to input into the first CBAM attention mechanism layer. The first CBAM attention mechanism layer outputs one-dimensional first output information F 1 ;

[0034] The first CBAM attention mechanism layer will update the representations of these tokens internally. Each token will be updated according to the information of other tokens. For each token, calculate its attention scores with all other tokens (including itself) in the set, normalize the attention scores, and use the normalized attention scores as weights to perform a weighted sum of the representations of other tokens. The updated token representation is added to the original token representation through a residual connection. Then, use layer normalization to normalize the updated token representation to obtain one-dimensional first output information F 1 ;

[0035] Then, use the first output information F 1 as Query to input into the second CBAM attention mechanism layer. At the same time, use the encoded sequence information in the updated image embedding feature information output by the memory attention module as Key and Value to input into the second CBAM attention mechanism layer. Combine Query, Key, and Value into three-dimensional matrix feature information. This feature information is processed by the second CBAM attention mechanism layer to obtain one-dimensional second output information F 2 ;

[0036] Input the second output information F 2 into the first MLP layer. After being processed by the first MLP layer, obtain the third output information F 3 ; Among them, the part of the third output information F 3 corresponding to the first output information F 1 is the mask prediction token related to the image segmentation task, and the part corresponding to the encoded pixel sequence information in the image embedding feature information is the specific image feature representation for predicting the segmentation mask;

[0037] Then, input the third output information F 3As a Query, it is input into the third CBAM attention mechanism layer. The output of the cross-attention mechanism of the memory attention is used as the Key and Value and input into the third CBAM attention mechanism layer. Then, after processing and updating the information, a one-dimensional fourth output information F is obtained. 4 ;

[0038] The fourth output information F 4 is used as the key and Value, and the masked prediction token in the third output information F 3 is used as the Query and input into the fourth CBAM attention mechanism layer again. After processing, the mask of each target point is output, and a one-dimensional fifth output information F is obtained. 5 ; Among them, the part of F 5 corresponding to the first output information F 1 is the masked prediction token related to the image segmentation task, and the part corresponding to the encoded pixel sequence information in the image embedding feature information is the specific image feature representation for predicting the segmentation mask;

[0039] At the same time, the fourth output information F 4 undergoes two convolutions to obtain an output feature F with the same size as the original image F. 6 ;

[0040] The part of the fifth output information F 5 corresponding to the masked prediction token is input into the second MLP layer. The output of the second MLP layer is dot-multiplied with the output feature F 6 to obtain the final instance segmentation map.

[0041] Furthermore, a width threshold of the obstacle, a first distance threshold between the obstacle and the robot, and a second distance threshold are set. The inspection robot uses the trained real-time end-to-end target detection model to obtain in real time whether there is an obstacle in each frame of the real-time video within the acquisition range of the camera. If there is an obstacle, the width of the obstacle is obtained. If the width of the obstacle is greater than the set width threshold of the obstacle and the distance between the obstacle and the robot is less than the first distance threshold, or if the width of the obstacle is not greater than the set width threshold of the obstacle and the distance between the obstacle and the robot is less than the second distance threshold, the inspection robot automatically adjusts the route to avoid the obstacle;

[0042] The first distance threshold is greater than the second distance threshold.

[0043] Furthermore, the width threshold of the obstacle is 3m, the first distance threshold between the obstacle and the robot is 2m, and the second distance threshold is 1m.

[0044] In a second aspect, the present invention provides a road inspection robot obstacle avoidance system based on instance segmentation, which executes the obstacle avoidance method described above. The system includes:

[0045] An image acquisition module for obtaining real-time videos of the inspection area during road inspection;

[0046] An image processing module for frame-by-frame obstacle annotation of road images in the real-time videos acquired by the image acquisition module and annotating the categories to which the obstacles belong;

[0047] The obstacle images include at least one of tree obstacle images, rock obstacle images, construction material obstacle images, road maintenance facility obstacle images, and animal obstacle images;

[0048] A real-time end-to-end object detection model for realizing real-time instance segmentation of obstacles and obtaining the widths of the obstacles;

[0049] A distance detection unit: for measuring the real-time distance between the obstacle and the robot;

[0050] An obstacle avoidance unit: according to the width of the obstacle and the real-time distance measured by the distance detection unit, timely adjust the inspection robot to automatically adjust the route to avoid obstacles.

[0051] Compared with the prior art, the beneficial effects of the present invention are:

[0052] The present invention creatively uses a real-time end-to-end object detector to achieve fast and high-precision detection of obstacles such as trees, rocks, construction materials, and road maintenance facilities by the road inspection robot. Especially for moving obstacles such as animals, it can more accurately and efficiently detect the obstacles in the inspection process in real time and remind the staff to deal with them in time. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 is a schematic structural diagram of the real-time end-to-end object detection model in the present invention.

[0054] Figure 2 is a schematic structural diagram of the VIM model in the present invention.

[0055] Figure 3 is a schematic structural diagram of the CBAM attention mechanism. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0056] In order to more clearly describe the technical problems, technical solutions, and advantages of the present invention, the following will be described in detail with reference to the drawings and embodiments. It should be noted that these embodiments are only used to illustrate the principles and application scope of the present invention and should not be regarded as a limitation of the present invention.

[0057] The obstacle avoidance method for a road inspection robot based on instance segmentation according to the present invention includes the following steps:

[0058] Step 1: Obtain the real-time video of the inspection area during road inspection, label the road images in the real-time video frame by frame, and divide the labeled obstacle images into a training set and a test set;

[0059] Step 2: Construct a real-time end-to-end object detection model,

[0060] The real-time end-to-end object detection model (see Figure 1 ) includes an image encoder module, a prompt encoder module, a mask decoder module, a memory attention module, a memory decoder module, and a memory bank. The image encoder module uses the VIM model, which has an identification condition "conditioned". The VIM model is used to divide the input first image into blocks and extract feature vectors for each block image. The image encoder outputs the first image embedding and the first identification condition "conditioned";

[0061] The prompt encoder module includes a convolutional block and CLIP. The mask prompt is input into the convolutional block, and the text prompt is input into CLIP. The prompt encoder module finally outputs a target token and a prompt token, and the target token and the prompt token are combined to form a prompt embedding;

[0062] The prompt embedding is input into the mask decoder. The mask decoder includes a first CBAM attention mechanism layer, a second CBAM attention mechanism layer, a first MLP, a third CBAM attention mechanism layer, a fourth CBAM attention mechanism layer, a convolutional operation, and a second MLP, and outputs an instance segmentation map;

[0063] The instance segmentation map output by the mask decoder is then input into the memory decoder module. After three convolutional operations, a feature map and an object point vector "object point" are obtained, and then summed with the first image embedding to combine the feature map and "object point" output by the three convolutional operations with the first image embedding to become the first memory feature; The first memory feature is input into the memory bank for storage; At this time, the processing of the first image is completed;

[0064] Then, start processing the second image. The image encoder module will output the second image embedding and the second recognition condition conditioned. The second image embedding is input into the memory attention module, which includes a CBAM attention mechanism and a cross-attention mechanism connected in sequence. The output result of the CBAM attention mechanism is used as the Query and input into the cross-attention mechanism. At the same time, the first memory feature stored in the memory bank is used as the Key and Value and input into the cross-attention mechanism. After the operation of the cross-attention mechanism, an updated second image embedding is output.

[0065] Finally, the updated second image embedding is respectively input into the second CBAM attention mechanism layer and the third CBAM attention mechanism layer of the mask decoder. The mask decoder outputs an instance segmentation map, which is then input into the memory decoder module. In the memory decoder module, the instance segmentation map output by the mask decoder is combined with the second image embedding after convolution operation processing to become the second memory feature, and then input into the memory bank for storage. At this time, the processing of the second image is completed.

[0066] Process all images according to the above-mentioned process.

[0067] The working process of the VIM model (see Figure 2 ) is as follows: First, the input image F is divided into multiple non-overlapping blocks of a fixed size. Then, the position information of each image block is embedded into its corresponding pixel vector representation information, and the pixel vector representation information of each image block is flattened and arranged into a sequence to obtain the original pixel sequence information of each image block and the corresponding position information. The original pixel sequence information of each image block is linearly transformed through a linear layer to adjust its data dimension, and the combined embedding feature of the primary pixel sequence information of each image block and the corresponding position information is obtained. The linear transformation is expressed as Y = TW Y + b Y , where T is the original pixel sequence information of the image block, W Y is the weight matrix, b Y is the bias vector, and Y is the primary pixel sequence information obtained after transformation.

[0068] Then, the primary pixel sequence information of the image block is batch-normalized through a normalization layer.

[0069] After that, the results obtained by processing the primary pixel sequence information Y(B, M, D) of the image block through the normalization layer are respectively projected using two different projection layers into sequences X(B, M, E) and Z(B, M, E) with a dimension of E.

[0070] Where B is the batch size; M is the sequence length; D is the dimension of the original features; E is the dimension of the extended state, set to 384;

[0071] The two different projection layers are two different linear layers, which project the input sequence to generate sequences X and Z; The linear layer consists of a weight matrix and a bias term. For sequences X and Z, there are weight matrices Wx and Wz and bias vectors bx and bz respectively; Multiply the input sequence by the weight matrix Wx and add the bias bx to obtain sequence X.

[0072] X = Wx × Input + bx,

[0073] Input is the input sequence, and Wx is a D×E matrix; The method for obtaining sequence Z is the same. The dimensions of the weight matrix and bias vector of the corresponding linear layer are the same as those of sequence X, but the element values are different.

[0074] Then, process sequence X in both the forward and backward directions to capture local dependencies in the vector. For forward processing, first apply a one-dimensional forward convolution operation to sequence X to obtain X'. o , and then project X' o linearly onto three different linear layers to obtain sequences B o : (B, M, N), C o : (B, M, N) and Δ o , where sequence Δ o uses the softplus activation function to ensure it is positive;

[0075] Δ o = log(1 + exp(X' o ))

[0076] where N is the dimension of the hidden state;

[0077] Then apply the Hadamard product to sequence Δ o and multiply it by A o and B o respectively according to the following formula to obtain the transformation parameters and

[0078]

[0079] Finally, use sequence C o , transformation parameter transformation parameter and the output X' of the forward convolution o to calculate the forward processing result y of X through the spatio-temporal state model SSM forward ;

[0080] The backward processing result y of sequence X is obtained by applying a one-dimensional backward convolution operation and a spatial state model backward ;

[0081] y is obtained from sequence X forward and y backward are respectively added to the results obtained after the sequence Z is processed by the activation function, and then the two obtained results are added together. The obtained result is processed by the projection layer and then residually connected to the primary pixel sequence information Y of the image patch, and then the output final encoded pixel sequence information T is generated l , T l will be sent to the multi-layer perceptron MLP for processing. T l is combined with the corresponding image patch position information to obtain the image embedding of the image patch, which is the output of the image encoder;

[0082] Step 3: Use the training set to train the real-time end-to-end object detection model. During training, the mask is used as the prompt input to the prompt encoder module to obtain the trained real-time end-to-end object detection model, which is used to perform frame-by-frame real-time detection on each frame of the real-time video in the inspection area during road inspection.

[0083] During actual detection and use, first, the real-time video will enter the image encoder module. The image encoder module includes a VIM model. The VIM model is used to divide the first frame image in the input video into blocks and extract feature vectors for each block image. The image encoder will finally output the first image embedding and a conditioned (recognition condition) (this conditioned is a recognition point, included in the image embedding, which can locate the object to be segmented and can avoid objects being recognized as one object due to overlap). The image embedding captures the high-level representation of all important visual features in the first frame image.

[0084] At this time, the prompt encoder module converts the input prompt into a vector sequence that the model can understand. In this implementation example, the input to the prompt encoder is a text prompt, and the prompt encoder module will output a vector sequence through CLIP as the encoder. The prompt encoder module outputs target tokens and prompt tokens, and the target tokens and prompt tokens are combined to form a prompt embedding.

[0085] Next, the prompt embedding is input into the Mask decoder. The Mask decoder includes a prompt embedding operation, a first CBAM attention mechanism layer, a second CBAM attention mechanism layer, a first MLP, a third CBAM attention mechanism layer, a fourth CBAM attention mechanism layer, a convolution operation, and a second MLP. Through attention mechanisms, MLP layers, and convolution operations, an instance segmentation map is output.

[0086] Then, the instance segmentation map output by the Mask decoder is input into the Memory decoder module. After three convolution operations, a feature map and an object point vector (essentially a vector used to predict the movement of the segmented part in the next frame) are obtained. Then, it is summed with the result of the first image embedding of the image embedding, combining the feature map and object point output by the three convolution operations with the first image embedding to form the first memory feature. The first memory feature is input into the Memory Bank for storage. At this time, the processing of the first frame of the image is completed.

[0087] Then, the processing of the second frame of the image begins. The image encoder module outputs a second image embedding, the second conditioned. The second image embedding and the conditioned are input into the Memory attention module. The Memory attention module includes a CBAM attention mechanism and a cross-attention mechanism. First, it passes through the CBAM attention mechanism and then enters the cross-attention mechanism as a Query. At the same time, the first memory feature stored in the Memory Bank is also input into the cross-attention mechanism as a Key and a Value. After the operation of the cross-attention mechanism, an updated second image embedding is output.

[0088] Finally, the updated second image embedding is input into the Mask decoder. Through attention mechanisms, MLP layers, and convolution operations, an instance segmentation map is output. Then, it is input into the Memory decoder module. In the Memory decoder module, the instance segmentation map output by the Mask decoder is combined with the second image embedding after convolution operation processing to form the second memory feature, and then input into the Memory Bank for storage. At this time, the processing of the second frame of the image is completed.

[0089] According to the above-mentioned process, until the last frame of the real-time video is processed, the real-time detection of the real-time video is completed.

[0090] In the present invention, the VIM (vision mamba) model is used as an image encoder to extract the feature representation of the input image. The VIM (vision mamba) model includes an image segmentation operation, batch normalization, two projection layers, forward convolution and linear layers, backward convolution and linear layers, a spatial state module, and an MLP:

[0091] Image segmentation operation: The input image F is segmented into multiple non-overlapping patches of a fixed size. Then, the position information of each image patch is embedded into its corresponding pixel vector representation information, and the pixel vector representation information of each image patch is flattened and arranged into a sequence to obtain the original pixel sequence information of each image patch and the corresponding position information. The original pixel sequence information of each image patch is linearly transformed through a linear layer to adjust its data dimension, obtaining the combined embedding features of the primary pixel sequence information of each image patch and the corresponding position information. The above linear transformation is defined by a weight matrix and a bias term, where the input vector is multiplied by the weight matrix and then the bias term is added. The linear transformation can be expressed as Y = TW Y +b Y , where T is the original pixel sequence information of the image patch, W Y is the weight matrix, b Y is the bias vector, and Y is the primary pixel sequence information obtained after the transformation.

[0092] Batch normalization: The primary pixel sequence information of the image patches is normalized through a normalization layer, and batch normalization is adopted.

[0093] Two projection layers: The results obtained by processing the primary pixel sequence information Y(B, M, D) of the image patches through the normalization layer are respectively projected using two different projection layers into sequences X(B, M, E) and Z(B, M, E) with a dimension of E (expanded state dimension, set to 384), where B is the batch size, M is the sequence length, and D is the original feature dimension.

[0094] The two different projection layers are two different linear layers (or called fully connected layers), and these two linear layers project the input sequence to generate sequences X and Z. The linear layer consists of a weight matrix and a bias term. For sequences X and Z, there are respectively a weight matrix Wx and Wz, and bias vectors bx and bz. Multiply the input sequence by the weight matrix Wx and add the bias vector bx to obtain sequence X,

[0095] X = Wx × Input + bx,

[0096] Input is the input sequence, Wx is a D×E matrix, E is the extended state dimension, and D is the original feature dimension. The sequence Z is obtained in the same way, and the dimensions of the weight matrix and bias vector of the corresponding linear layer are the same as those of the sequence X, but the element values are different.

[0097] Process the sequence X in both the forward and backward directions to capture local dependencies in the vector: For the forward processing, first apply a one-dimensional forward convolution operation to the sequence X to obtain X' o , and then project X' o linearly onto three different linear layers to obtain the sequences B o : (B, M, N), C o : (B, M, N) and Δ o , where the sequence Δ o uses the softplus activation function to ensure it is positive.

[0098] Δ o = log(1 + exp(X' o ))

[0099] The sequence Δ o is multiplied by the sequence A o (E, N) and the bias vector bx and then subjected to a Hadamard product to obtain the transformation parameter The sequence Δ o is multiplied by the sequence B o (B, M, N) and the bias vector bz and then subjected to a Hadamard product to obtain the transformation parameter

[0100]

[0101] where N is the dimension of the hidden state; the sequence A o (E, N) is a learnable parameter.

[0102] The state space model SSM is a system that maps a one-dimensional function or sequence x(t) (x(t) ∈ R) to y(t) (y(t) ∈ R) through a hidden state h(t) (h(t) ∈ R N ).

[0103] Using the parameter C o , and the output X' of the forward convolution o in combination with the evolution parameter α (α ∈ R N*N ), the projection parameter β (β ∈ R N*1 ), the forward processing result y of the sequence X is calculated through the state space model forward, and its specific process is expressed by the formula as follows:

[0104] h(X) = α + β × X

[0105]

[0106] y forward = C o × h’(X)

[0107] In the above formula, X is the sequence X; α is the evolution matrix of the hidden state, and β is the influence matrix of the input signal x(t) on the hidden state. Both are trainable parameter matrices.

[0108] The backward processing result y of the sequence X backward The process of obtaining is similar to the process of obtaining the forward processing result y of the sequence X forward The difference is that a one-dimensional backward convolution operation is applied.

[0109] The y obtained from the sequence X forward 、y backward and the results obtained after the sequence Z is processed by the activation function are respectively added, and then the two obtained results are added together. The obtained result is processed by the projection layer and then residually connected (added) with the primary pixel sequence information Y of the image patch, and then the output final encoded pixel sequence information T is generated l , T l will be sent to the MLP (Multi-Layer Perceptron) for processing. Combine T l with the corresponding image patch position information to obtain the embedded feature information of the image patch, which is the image embedding.

[0110] In the present invention, the prompt encoder module is used to convert the input prompt into a vector sequence that the model can understand. There are two types of prompts, text and mask. The mask prompt is used during training. After convolution processing, the mask obtains the corresponding target token and indication token. In actual use, the mask prompt does not work, and only the text prompt is input. The input text prompt is processed by the ClIP model to obtain the target token and the prompt instruction.

[0111] The Memory attention module is composed of the CBAM attention mechanism and the Cross attention mechanism. When an image embedding is input into the Memory attention module, it first enters the CBAM attention mechanism to achieve self-adaptive refinement of its own features, and then enters the Cross attention mechanism together with the memory features stored in the memory bank of the previous frame. At this time, the features output by CBAM are used as Query, and the memory features of the previous frame are used as Key and Value. Finally, the Memory attention module will input an updated image embedding into the second and third CBAM attention mechanism layers of the Mask decoder.

[0112] The CBAM attention mechanism (see Figure 3 ) includes a channel attention module and a spatial attention module. The output of the previous level is used as the input of Query, Key, and Value, and the three are arranged into a matrix and input into the Channel Attention Module to emphasize the features in different channels. In the channel attention module, average pooling and max pooling are used to compress the features in the spatial dimension to generate two different spatial context descriptors F c-avg and F c-max . Then, the two descriptors are input into a shared multi-layer perceptron (MLP) to generate the channel attention feature M c , and the formula is as follows:

[0113] M c =σ(W 1 (W 0 (F c-avg ))+W 1 (W 0 (F c-max ))) +b

[0114] where σ refers to the sigmoid function, W 0 ∈R C / r*C , W 1 ∈R C*C / r ; R is the hidden activation size, r is the reduction rate, and C is the number of channels. W 0 , W 1 are the trainable parameters of the shared multi-layer perceptron MLP, and b is the parameter;

[0115] Then, the spatial regions in the features are emphasized through the Spatial Attention Module. The spatial attention module applies average pooling and max pooling in the channel dimension to generate two 2D maps F s-avg and F s-max . These two 2D maps are then concatenated in the channel dimension and passed through a convolutional layer to generate the spatial attention feature M s . The formula is as follows:

[0116] M s =σ(f 7*7 ([F s-avg ; F s-max )) +b

[0117] σ refers to the sigmoid function, f 7*7 represents the convolutional operation of a 7*7 filter, b is a parameter, and the b parameter of M s is the same as that of M c to keep them consistent.

[0118] First, multiply the channel attention feature M c by the input feature of the channel attention module to obtain the refined feature F'. Take the feature F' as the input of the spatial attention module, and then multiply the spatial attention feature M s by F' to obtain the final adaptively refined feature F1. The original feature is refined using the channel attention feature and the spatial attention feature, and the refinement process is completed through element-wise multiplication.

[0119] The Mask decoder module is used to combine the image embedding output by the image encoder and the output tokens and prompt tokens output by the prompt encoder as inputs, and finally output the predicted segmentation mask. Its specific workflow is as follows:

[0120] First, combine the output tokens and prompt tokens into a prompt embedding F 0 , and use F 0 as the Query, key, and value to input the first CBAM attention mechanism layer. The first CBAM attention mechanism layer outputs the one-dimensional first output information F 1 .

[0121] The first CBAM attention mechanism layer updates the representations of these tokens internally, and each token is updated based on the information of other tokens. For each token, the attention scores between it and all other tokens (including itself) in the set are calculated. This is achieved by computing the dot product between tokens. The softmax function of the channel attention module in the CBAM attention mechanism is used to normalize the attention scores, ensuring that the sum of all scores is 1. This step ensures that the model can weigh the relative importance between different tokens. The normalized attention scores are used as weights to perform a weighted sum of the representations of other tokens. In this way, the updated representation of each token not only contains its own information but also integrates the information of all other tokens. The updated token representation is added to the original token representation through a residual connection to enhance the learning ability of the model and avoid the vanishing gradient problem in deep networks. Then, layer normalization is used to normalize the updated token representation to obtain the one-dimensional first output information F 1 ;

[0122] Then, the first output information F 1 is used as the Query and input into the second CBAM attention mechanism layer. At the same time, the encoded sequence information in the updated image embedding feature information output by the memory attention module is used as the Key and Value and input into the second CBAM attention mechanism layer. The above Query, Key, and Value are combined into the feature information of a three-dimensional matrix. After being processed by the second CBAM attention mechanism layer, the one-dimensional second output information F 2 ;

[0123] The second output information F 2 is input into the first MLP layer. After being processed by the first MLP layer, the third output information F 3 is obtained; among them, the part of the third output information F 3 corresponding to the first output information F 1 is the mask prediction token related to the image segmentation task, and the part corresponding to the encoded pixel sequence information in the image embedding feature information is the specific image feature representation for predicting the segmentation mask.

[0124] Then, the third output information F 3 is used as the Query and input into the third CBAM attention mechanism layer. The output of the cross-attention mechanism of the memory attention is used as the Key and Value and input into the third CBAM attention mechanism layer. Then, after processing and updating the information, the one-dimensional fourth output information F 4 ;

[0125] The fourth output information F 4As the key and Value and the third output information F 3 The masked prediction tokens in 3 are input into the fourth CBAM attention mechanism layer again as Query. After processing, the masks of each target point are output, and the one-dimensional fifth output information F is obtained. 5 ; where F 5 The part in 5 that corresponds to the first output information F 1 The corresponding part is the masked prediction tokens related to the image segmentation task, and the part corresponding to the encoded pixel sequence information in the image embedding feature information is the specific image feature representation for predicting the segmentation mask.

[0126] Meanwhile, the fourth output information F 4 After two convolutions, the output feature F with the same size as the original image F (the initially input image) is obtained. 6 ;

[0127] The part in the fifth output information F 5 that corresponds to the masked prediction tokens is input into the second MLP layer. The output of the second MLP layer is dot-multiplied with the output feature F 6 to obtain the final instance segmentation map. The output of the second MLP layer after processing is also an instance segmentation map with each detected object.

[0128] After training, the mask decoder will finally output the instance segmentation image of the original image F. Taking the output of the mask decoder as the output of the real-time end-to-end object detection model, and combining the output result of the model with the pixel block size of the segmented instance, the width of the obstacle is accurately calculated. In the present invention, the labeled image of 1024×1024 pixels is used during training, and it is converted according to the actual ratio of 1:20, that is, 1 pixel is 0.05m.

[0129] The memory decoder module consists of 3 3x3 convolution operations and a summation layer. The instance segmentation image output by the mask decoder will be input into the memory decoder module again. In the memory decoder module, it first enters the convolution layer for convolution to obtain an object point. Then, the object point is input into the summation layer, and in the summation layer, the object point is summed with the image embedding output by the image encoder module to obtain a memory feature. Finally, this memory feature is input into the memory bank module for storage.

[0130] Embodiment 1

[0131] This embodiment is an obstacle avoidance system for a road inspection robot based on instance segmentation, which executes the described obstacle avoidance method. The system includes:

[0132] An image acquisition module for obtaining real-time videos of the inspection area during road inspection;

[0133] An image processing module for frame-by-frame obstacle annotation of road images in the real-time videos acquired by the image acquisition module and annotating the obstacle category;

[0134] The obstacle images include at least one of tree obstacle images, rock obstacle images, construction material obstacle images, road maintenance facility obstacle images, animal obstacle images, etc.;

[0135] A real-time end-to-end object detection model for realizing real-time instance segmentation of obstacles and obtaining the width of obstacles;

[0136] A distance detection unit: for measuring the real-time distance between the obstacle and the robot;

[0137] An obstacle avoidance unit: according to the width of the obstacle and the real-time distance measured by the distance detection unit, timely adjust the inspection robot to automatically adjust the route to avoid obstacles.

[0138] The specific steps are as follows:

[0139] (1) Obtain the video of the inspection road during inspection and perform image color enhancement processing on each frame of the video image;

[0140] (1.1) Use the image acquisition module to obtain about 20,000 road images of the inspection area during the inspection of the inspection robot;

[0141] (1.2) Convert the format of the road images to the standard format 1024*1024 to obtain standard format images;

[0142] (1.3) Use the Gaussian function to perform low-pass filtering on the obtained standard format images to obtain filtered images;

[0143] (1.4) Convert the obtained standard format images and the filtered images to the logarithmic domain and subtract (i.e., subtract the low-frequency components in the images) to obtain the reflected images in the logarithmic domain;

[0144] (1.5) Perform image summation on the reflected images in the logarithmic domain at the pixel level to obtain the result;

[0145] (1.6) Color restoration:

[0146] (1.6.1) At the channel level, for each channel, add the pixel values of this channel in the original image to obtain the sum of this channel, calculate the maximum value of the sums of all channels as the denominator of the normalization factor. For the sum of each channel, divide it by the denominator of the normalization factor to obtain the normalization factor;

[0147] (1.6.2) Normalize the weight matrix (the weight matrix is manually set according to different applications. Here, a uniform weight matrix is adopted, that is, the weights of each channel are equal), convert it to the logarithmic domain, then multiply by the non-linear factor for color restoration (here taken as 2.0), and then divide by the normalization factor to obtain the result of image color gain.

[0148] (1.6.3) Recombine (multiply) the result obtained in (1.5) according to the weight matrix and the result of image color gain to obtain the image result after color restoration;

[0149] (1.7) Image restoration: Multiply the image result after color restoration obtained in (1.6.3) by the gain of the image pixel value change range, and then add the offset of the image pixel value change range to obtain the final result.

[0150] Among them, the gain of the image pixel value change range is a multiplication factor used to adjust the brightness and contrast of the image. It is applied to each pixel value of the image result after color restoration. The offset of the image pixel value change range is a constant value used to translate the image after the multiplication factor. It is added to each pixel value of the image result after color restoration. By multiplying the multiplication factor by the image result after color restoration and adding the offset to the multiplication result, the final image restoration result is obtained.

[0151] (2) First, preprocess the acquired image, and then label the obstacle type of the preprocessed obstacle image.

[0152] (2.1) Image data preprocessing: Introduce the sliding window technique. Randomly select a point in the image as the center point of a square sliding window. Use a square sliding window with a constant size to slide on the image in the order from left to right and from top to bottom starting from the center point position. Each time it slides, the image area covered by the sliding window is used as a sub-image and saved. Among them, the size range of the square sliding window is 100 - 300 pixels, and the sliding step size of the square sliding window is set to range from 50% to 70% of its size. If the side length of the remaining uncovered image in the current row or column is less than the sliding step size of the square sliding window, the part exceeding the sliding step size is filled with 0, and then continue to slide downward with the original sliding step size until the square sliding window covers all areas of the image, stop sliding, and collect all sub-images.

[0153] (2.2) Preprocess all the collected sub-images. First, label them. Label the sub-images containing obstacles as 1, that is, positive samples; label the sub-images without obstacles as 0, that is, negative samples.

[0154] (2.3) Perform data augmentation on the labeled dataset. Repeat the original image three times for operations ①②③ to obtain three differential images, with the difference being that the sigmoid parameters for Gaussian blur are different each time, set to 20, 100, and 180 respectively. Finally, weighted average the three differential images obtained with different sigmoid parameters to obtain the augmented dataset.

[0155] Finally, randomly allocate the augmented dataset and the labeled dataset into a training set and a test set according to a ratio of 7 to 3.

[0156] (3) Construct and train a real-time end-to-end object detection model

[0157] Construct a real-time end-to-end object detection model as described above and perform training. During the training process, transfer learning is used: in this embodiment, select a dataset that has been pre-trained on 11 million natural images and 1.1 billion masks, and use the masks as prompts at this time.

[0158] After that, based on the pre-trained parameters, use the aforementioned training set to train the real-time end-to-end object detection model again. Iteratively train all images in turn. After 100,000 iterations, calculate the pixel-level training loss using the mean squared error loss function, and then update the trainable parameters of the network by backpropagation of the training loss. Iteratively train all images in turn. Completing the training of all the images in the training set completes one round of training. Use the trainable parameters of the network that completed the previous round as the initial trainable parameters for the next round. When the training gradient value is within the range of ±1e -5 or the number of training iteration rounds reaches 100,000, stop the training, and extract the network parameters with the minimum training loss, that is, obtain the trained real-time end-to-end object detection model.

[0159] Then input the images in the test set and the corresponding annotations into the trained real-time end-to-end object detection model, and calculate the accuracy of the segmentation result. When the pixel accuracy is 95%-100%, it proves that the trained real-time end-to-end object detection model is the optimal segmentation model. Otherwise, adjust the initialization method of the trainable parameters of the real-time end-to-end object detection model, adjust the values of the fixed parameters of the network, and repeat the training process and the testing process of the segmentation model until the optimal segmentation model is obtained.

[0160] Train on the same training set, and at the same time, perform performance testing on the same test set. The performance comparisons of different network models are shown in Table 1.

[0161] Table 1 Performance comparison of different models

[0162]

[0163] The mean PixelAccuracy (mPA for short) is a commonly used evaluation metric in image segmentation tasks. It measures the consistency between the mask predicted by the segmentation model and the ground truth mask. The pixel accuracy is the proportion of correctly predicted pixels for each pixel point, and it is an intuitive metric for measuring the quality of image segmentation. The obstacle recognition measurement accuracy usually refers to the accuracy of instance segmentation of obstacles in image analysis or computer vision tasks. The average FPS is the number of images that can be processed per second.

[0164] (4) Input the video image to be recognized into the trained real-time end-to-end object detection model for obstacle detection and obstacle avoidance.

[0165] (4.1) The inspection robot is equipped with a binocular camera with a resolution of 1080P and 2 million sensor pixels. The camera is located in the middle of the top of the inspection robot, 1.6 meters above the ground. The camera is tilted towards the ground at an angle of 10° with the ground. When inspecting, the industrial camera will capture images of an area 8 meters long and 5 meters wide directly in front of the inspection robot.

[0166] A binocular camera is an imaging system equipped with two cameras. It simulates human binocular vision and obtains depth information by capturing two slightly different views of the same scene. By comparing the disparity of the images captured by the two cameras, it can automatically obtain the depth information of the objects in the scene. According to the obtained distance and the built-in measurement scale, combined with the width of the target obstacle output by the real-time end-to-end object detection model, it can accurately measure the size of the target.

[0167] (4.2) If the recognized image shows no obstacles, the inspection robot continues its inspection.

[0168] (4.3) If the recognized image shows obstacles, the inspection robot issues an alarm to remind the staff to clear the obstacles in time until the inspection is completed. If there is an obstacle with a width greater than 3m and a distance of 2 meters from the robot, the inspection robot will automatically adjust its route to avoid the obstacle. Or if the width of the obstacle is less than 3 meters and the inspection robot is less than 1 meter away from the obstacle, the inspection robot will automatically adjust its route to avoid the obstacle.

[0169] Obstacle images specifically include: images of trees, rocks, construction materials, road maintenance facilities, animals, etc.

[0170] The present invention creatively uses a real-time end-to-end object detector to achieve the detection of obstacles such as trees, rocks, construction materials, road maintenance facilities, etc. by the road inspection robot. Especially for moving obstacles such as animals, it can more accurately and efficiently detect the obstacles in the inspection process in real time and avoid obstacles automatically. The method of the present invention can more precisely segment the contour of the object, can distinguish and separately mark the overlapping or contacting objects in the image, and can better distinguish individual instances.

[0171] The method in the present invention is applicable to high-resolution image processing and meets the training requirements of large-scale visual datasets.

[0172] In the present invention, the Mask decoder module uses the CBAM attention mechanism to independently process the channel and spatial dimensions, reduces the computational complexity and the amount of computation, avoids complex attention weight calculations and high-dimensional matrix multiplication operations, enables the network to better capture the key information in the image, and helps to improve the accuracy of inspection obstacle detection.

[0173] What is not described in the present invention is applicable to the prior art.

Claims

1. A road inspection robot obstacle avoidance method based on instance segmentation, characterized in that: The method comprises the following steps: Step 1: obtain real-time video of the inspection area during the road inspection process, and mark obstacles frame by frame on the road images in the real-time video, and divide the marked obstacle images into a training set and a test set; Step 2: Build a real-time end-to-end object detection model. The real-time end-to-end object detection model includes an image encoder module, a prompt encoder module, a mask decoder module, a memory attention module, a memory decoder module and a memory bank. The image encoder module adopts a VIM model and has recognition conditional conditioning. The VIM model is used to divide the input first image into blocks and extract feature vectors for each block of the image. The image encoder outputs a first image embedding and a first recognition conditional conditioning. The prompt encoder module includes a convolution block and a CLIP, the mask prompt is input into the convolution block, the text prompt is input into the CLIP, and the prompt encoder module finally outputs a target token and a prompt token, and the target token and the prompt token are combined to form a prompt embedding; The hint embedding is input into a mask decoder, which includes a first CBAM attention layer, a second CBAM attention layer, a first MLP, a third CBAM attention layer, a fourth CBAM attention layer, a convolution operation, a second MLP, and outputs an instance segmentation map; The instance segmentation map output by the mask decoder is then input into the memory decoder module, and after three convolution operations, a feature map and an object point vector are obtained, which are then summed with the first image embedding to realize the combination of the feature map and object point output by the three convolution operations with the first image embedding to become the first memory feature; the first memory feature is input into the memory bank for storage; at this point, the processing of the first image is completed; Then, the second image is processed, and the image encoder module outputs the second image embedding and the second recognition condition conditioned; the second image embedding is input into the memory attention module, and the memory attention module includes a CBAM attention mechanism and a cross attention mechanism connected in sequence. The output result of the CBAM attention mechanism is input into the cross attention mechanism as a query. At the same time, the first memory feature stored in the memory bank is input into the cross attention mechanism as a key and a value. After the operation of the cross attention mechanism, an updated second image embedding is output; Finally, the updated second image embedding is input into the second CBAM attention mechanism layer and the third CBAM attention mechanism layer of the mask decoder respectively. The mask decoder outputs an instance segmentation map, which is then input into the memory decoder module. In the memory decoder module, the instance segmentation map output by the mask decoder is processed by the convolution operation and combined with the second image embedding to become the second memory feature, which is then input into the memory bank for storage. At this point, the processing of the second image is completed; Following the above process, all images are processed; The workflow of the VIM model is as follows: first, the input image F is divided into multiple non-overlapping blocks of fixed size, and then the position information of each image block is embedded into its corresponding pixel vector representation information, and the pixel vector representation information of each image block is flattened and arranged into a sequence to obtain the original pixel sequence information and the corresponding position information of each image block; the original pixel sequence information of each image block is linearly transformed through a linear layer to adjust its data dimension, and the combined embedding features of the primary pixel sequence information and the corresponding position information of each image block are obtained; the linear transformation is expressed as Y=TW Y +b Y , where T is the original pixel sequence information of the image block, W Y is the weight matrix, b Y is the bias vector, and Y is the primary pixel sequence information obtained after transformation; Then, the primary pixel sequence information of the image block is batch normalized through the normalization layer; Afterwards, the primary pixel sequence information Y(B, M, D) of the image block is processed by the normalization layer and the results are projected into a sequence X(B, M, E) and a sequence Z(B, M, E) with a dimension of E using two different projection layers respectively; Where B is the batch size; M is the sequence length; D is the original feature dimension; E is the expanded state dimension, which is set to 384; The two different projection layers are two different linear layers, which project the input sequence to generate sequences X and Z; the linear layer is composed of a weight matrix and a bias term, and for sequences X and Z, there are weight matrices Wx and Wz and bias vectors bx and bz respectively; the input sequence is multiplied by the weight matrix Wx, and the bias bx is added to obtain the sequence X, X=Wx×Input+bx, Input is the input sequence, Wx is a D×E matrix; the sequence Z is obtained in the same way. The dimensions of the weight matrix and bias vector of the corresponding linear layer are the same as those of the sequence X, but the element values ​​are different; Then, the sequence X is processed in both forward and backward directions to capture local dependencies in the vector; for forward processing, a one-dimensional forward convolution operation is first applied to the sequence X to obtain X′ o , and then X' o Linear projection to three different linear layers, respectively, to obtain sequence B o :(B, M, N), C o :(B, M, N) and Δ o , where the sequence Δ o Use a soft-add activation function to ensure it is positive; Δ o =log(1+exp(X' o )) Where N is the dimension of the hidden state; Then the sequence Δ o Apply the Hadamard product according to the following formula and A o , B o Multiply to get the conversion parameter and Finally, use sequence C o , conversion parameters Conversion parameters And the output of the forward convolution X' o Through the spatial state model SSM, the forward processing result y of X is calculated forward ; Apply a one-dimensional backward convolution operation and a spatial state model to obtain the backward processing result y of the sequence X backward ; Will get y from the sequence X forward ,y backward The results obtained after the activation function processing of the sequence Z are added separately, and then the two results are added together. After the result is processed by the projection layer, it is residually connected with the primary pixel sequence information Y of the image block, and then the final encoded pixel sequence information T is generated. l , T l It will be sent to the multi-layer perceptron MLP for processing. l Combined with the corresponding image block position information, the image embedding of the image block is obtained, which is the output of the image encoder; Step 3: Use the training set to train a real-time end-to-end object detection model. During training, the mask is used as a prompt input to prompt the encoder module to obtain a trained real-time end-to-end object detection model, which is used to perform real-time detection frame by frame on each frame of the real-time video of the inspection area during the road inspection process; The mask decoder module is used to combine the image embedding output by the image encoder and the target token and prompt token output by the prompt encoder as input, and finally output the predicted segmentation mask. The specific workflow is as follows: First, the target token and the prompt token are combined into the prompt embedding F0, and F0 is input into the first CBAM attention mechanism layer as the query, key, and value. The first CBAM attention mechanism layer outputs the one-dimensional first output information F1; The first CBAM attention mechanism layer will update the representations of these tokens internally, and each token will be updated according to the information of other tokens; for each token, its attention score with all other tokens in the set is calculated, the attention score is normalized, and the normalized attention score is used as a weight to perform a weighted sum of the representations of other tokens; the updated token representation is added to the original token representation through a residual connection, and then the updated token representation is normalized using layer normalization to obtain the one-dimensional first output information F1; Then, the first output information F1 is input as Query into the second CBAM attention mechanism layer. At the same time, the encoding sequence information in the updated image embedding feature information output by the memory attention module is input as Key and Value into the second CBAM attention mechanism layer. The Query, Key and Value are combined into feature information of a three-dimensional matrix. The feature information is processed by the second CBAM attention mechanism layer to obtain a one-dimensional second output information F2. The second output information F2 is input into the first MLP layer, and after being processed by the first MLP layer, the third output information F3 is obtained; wherein the part of the third output information F3 corresponding to the first output information F1 is a mask prediction token related to the image segmentation task, and the part corresponding to the encoded pixel sequence information in the image embedding feature information is a specific image feature representation used to predict the segmentation mask; Then, the third output information F3 is input as Query to the third CBAM attention mechanism layer, and the output of the cross attention mechanism of memory attention is input as Key and Value to the third CBAM attention mechanism layer, and then after processing and updating the information, the one-dimensional fourth output information F4 is obtained; The fourth output information F4 is used as the key and value and the mask prediction token in the third output information F3 is used as the query to be input into the fourth CBAM attention mechanism layer again. After processing, the mask of each target point is output to obtain the one-dimensional fifth output information F5; wherein the part of F5 corresponding to the first output information F1 is the mask prediction token related to the image segmentation task, and the part corresponding to the encoded pixel sequence information in the image embedding feature information is the specific image feature representation used to predict the segmentation mask; At the same time, the fourth output information F4 undergoes two convolutions to obtain an output feature F6 of the same size as the original image F; The part of the fifth output information F5 corresponding to the mask prediction token is input into the second MLP layer, and the output of the second MLP layer is dot-multiplied with the output feature F6 to obtain the final instance segmentation map.

2. The method according to claim 1, characterized in that The width threshold of the obstacle, the first distance threshold and the second distance threshold between the obstacle and the robot are set. The inspection robot uses the trained real-time end-to-end target detection model to obtain in real time whether there is an obstacle in the real-time video frame by frame within the camera acquisition range. If there is an obstacle, the width of the obstacle is obtained. If the obstacle width is greater than the set obstacle width threshold and the distance between the obstacle and the robot is less than the first distance threshold, or if the obstacle width is not greater than the set obstacle width threshold and the distance between the obstacle and the robot is less than the second distance threshold, the inspection robot automatically adjusts the route to avoid obstacles. The first distance threshold is greater than the second distance threshold.

3. The method according to claim 1, characterized in that The width threshold of the obstacle is 3m, the first distance threshold between the obstacle and the robot is 2m, and the second distance threshold is 1m.

4. A road inspection robot obstacle avoidance system based on instance segmentation, characterized in that: The obstacle avoidance method according to any one of claims 1 to 3 is implemented, wherein the system comprises: Image acquisition module, used to obtain real-time video of the inspection area during road inspection; An image processing module is used to mark obstacles frame by frame on the road image in the real-time video collected by the image acquisition module, and mark the category to which the obstacles belong; The obstacle image includes at least one of a tree obstacle image, a rock obstacle image, a construction material obstacle image, a road maintenance facility obstacle image, and an animal obstacle image; Real-time end-to-end object detection model, used to achieve real-time instance segmentation of obstacles and obtain the width of obstacles; Distance detection unit: used to measure the real-time distance between obstacles and the robot; Obstacle avoidance unit: According to the width of the obstacle and the real-time distance measured by the distance detection unit, the inspection robot automatically adjusts its route to avoid obstacles.

Citation Information

Patent Citations

  • Water surface target detection method and system based on improved Deformable DETR

    CN118015255A

  • Mask correction-based semi-supervised video target segmentation method and system

    CN118097507A