Multi-modal Target Following Method and System for a Follow-up Vehicle under Night Conditions
Through the YOLOv5-RTFT network, the RGB and infrared image features are integrated with FPN and PAN structures, the problem of low night object detection accuracy of follow-up cars is solved, efficient target follow-up is achieved, cost is reduced and the applicability of night application scenarios is improved.
Patent Information
- Application Number
- CN202310695587.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-13
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2043-06-13
AI Technical Summary
Traditional follower cars have reduced target detection accuracy in environments with poor night light, resulting in limited application scenarios and requires a lot of manpower and material resources.
The YOLOv5-RTFT target detection network is adopted, combined with RGB images and infrared images, and the characteristics are integrated through the Transformer architecture, multi-scale information is enhanced using FPN and PAN structures, and the gimbal structure is equipped for active search to achieve target follow-up.
It improves the accuracy of night target detection, enhances the car's follow-up ability in complex lighting scenarios, and reduces manpower and material consumption.
Smart Images

Figure CN116934794B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision, and in particular to a multi-modal target following method and system for a follow-up vehicle under night conditions. Background Art
[0002] In recent years, more and more technologies have been integrated into the field of functional vehicles, making functional vehicles more and more diverse. Currently, intelligent follow-up vehicles are a popular field. Traditional following methods include laser following technology, GPS following technology, Bluetooth following technology, vision following technology, etc. Laser following technology has high power consumption and is often blocked by indoor walls or objects, with poor practicability; GPS following technology has low positioning accuracy in places with poor signals; although Bluetooth following technology is less affected by the environment, it has a short operating range and weak communication capabilities, and is not convenient to integrate into other systems; for the vision target following method, it is completed with the help of a vision sensor. The robot obtains images with the help of other external devices such as a monocular camera, a binocular camera, a depth camera, a video signal digitization device, or a fast signal processor based on DSP. At the same time, with the rapid development of deep learning, the convolutional neural network can automatically discover the features required for detecting and classifying targets. At the same time, through the convolutional neural network, the original input information can be transformed into more abstract and higher-dimensional features. This high-dimensional feature has strong feature expression ability and generalization ability, and performs well in complex scenarios.
[0003] The follow-up method based on target detection uses global information and has a fast detection speed, which can meet the processing requirements of specific occasions. Based on the characteristics of their algorithm flow, the target detection algorithms based on deep learning can be roughly divided into two categories: two-stage (Two-Stage) target detection algorithms and one-stage (One-Stage) target detection algorithms. The main representative of the two-stage target detection algorithm is the Regions with Convolutional Neural Networks Features (R-CNN) series. Although such detection algorithms have high detection accuracy, their detection speed is slow. The representative of the One-Stage target detection algorithm is the YouOnly Look Once (YOLO) series. Such detection algorithms have general accuracy, but their detection speed is very fast, and they have the advantages of high efficiency, flexibility, and good generalization performance, and are widely used in the industrial field.
[0004] Since the development of the YOLO algorithm series to date, it has included algorithms from YOLOV1 to YOLOV7 and various object detection algorithms based on improved YOLO. Redmon J et al. proposed in 2016 to directly use the method of regression to detect and classify the viewing frame, transforming object detection into a regression problem to solve, and based on a single end-to-end network, completing the output from the original input image to the object position and category, greatly improving the speed of object detection. After a series of optimizations, it has become the mainstream algorithm in the field of object detection and is widely used in the industrial community. Batch Normalization is used in YOLOV2 to preprocess the data, greatly improving the training speed and the training effect, and introducing the Anchor mechanism generated by the K-means clustering method of the standard Euclidean distance, greatly improving the recall rate of the algorithm. YOLOV3 uses a method of fusing multiple scales and becomes more adaptable to objects of different scales. YOLOV4 adds the SPP structure to solve the problem of multi-scale detection, introduces the PAN structure to enable the shallow feature map to have the semantic information of the deep feature map and the deep feature map to have the semantic information of the shallow feature map, and uses Mosaic data augmentation to increase the generalization of the network. YOLOV5 uses the CSP structure with residual modules in the BackBone stage to strengthen the network's ability to fuse features.
[0005] The core problem of functional vehicle detection and tracking targets is that during the following process, complex lighting scenarios and street environments may affect the accuracy of its target detection. YOLOv5 has relatively high accuracy, is widely applied, and has high compatibility. It uses DarkNet as the backbone network and applies the CSP module and the PAN structure. However, since the application scenarios of the follow-up vehicle are not only in the case of good daylight, but also may be in places with poor lighting at night. Although YOLOv5 has good results in the daytime scenario, its performance will be greatly reduced at night, greatly affecting the application scenarios of the follow-up vehicle. To expand the application scenarios of the follow-up vehicle, more information needs to be extracted and the network structure needs to be improved. Summary of the Invention
[0006] The purpose of the present invention is to overcome the defects of the existing technology that a traditional cleaning vehicle requires a dedicated person to drive, resulting in high costs and low efficiency, and consuming a large amount of manpower and material resources, and to provide a multi-modal target following method and system for follow-up vehicles under night conditions that can assist the work efficiency of staff.
[0007] The purpose of the present invention can be achieved by the following technical solutions:
[0008] A multi-modal target following method for a follow-up vehicle under night conditions, comprising the following steps:
[0009] During the movement of the vehicle, various attitude images of the target to be detected are collected by a camera and an infrared imager, and image annotation is performed to produce a training data set.
[0010] The data in the training data set is input into a pre-constructed YOLOv5-RTFT target detection network for training to obtain a trained target detection model. The YOLOv5-RTFT target detection network is a dual-path network structure based on YOLOv5, and an RTFT structure is introduced. On the basis of the Transformer architecture, the Decoder structure is deleted, and the image information is segmented into multiple patches, so as to fuse RGB image features and infrared image features.
[0011] During the target following process of the vehicle, an RGB image and a depth image are collected by a camera, and an infrared image is collected by an infrared imager.
[0012] The collected RGB image and infrared image are input into the trained target detection model to obtain a detection result.
[0013] According to the detection result, the difference between the center coordinates of the tracking target and the center coordinates of the viewfinder frame is obtained, so as to judge the steering angle of the vehicle, so that the center point of the tracking target remains the center point of the viewfinder frame.
[0014] The coordinates of the tracking target on the RGB image are mapped to the depth image to obtain the distance between the tracking target and the vehicle, which is used to judge whether to move forward to achieve target following.
[0015] Further, the YOLOv5-RTFT target detection network includes an input end module, a Backbone module, a Neck module, and a Prediction module.
[0016] Further, the input end module uses Mosaic technology for data augmentation and adopts an adaptive Anchor calculation method to adjust the calculated Anchor.
[0017] The process of using Mosaic technology for data augmentation includes: using Mosaic technology to randomly crop, scale four pictures in the input data set and then randomly splice them into one picture to achieve data set expansion.
[0018] The process of the described adaptive Anchor calculation method includes: before the start of training, calculating the width and height of all targets in the dataset input to the network, thereby calculating the best recall rate of the annotation information of this dataset for the default Anchor. If the best recall rate meets the preset recall rate requirement, the Anchor is not updated; otherwise, the Anchor of this dataset is recalculated.
[0019] Furthermore, the Backbone module includes a CBS structure, an RTFT structure, and a BottleNeck structure;
[0020] The CBS structure includes a Conv layer, a Batch Normalization layer, and a SiLU layer connected in series in sequence; the Conv layer includes a 1×1 convolutional layer and a 3×3 convolutional layer, and both the 1×1 convolutional layer and the 3×3 convolutional layer are used to expand the feature maps of RGB images and depth images; the Batch Normalization layer is used to perform normalization processing on an entire feature map as a neuron by using the weight sharing strategy; the SiLU layer is an activation function layer based on SiLU.
[0021] The RTFT structure is used to fuse the features of RGB images and depth images based on the Vision Transformer structure;
[0022] The BottleNeck structure is a BottleNeckTrue structure or a BottleNeckFalse structure. The BottleNeckTrue structure first performs convolution through a 1×1 CBS structure, then performs convolution through a 3×3 CBS structure, and finally adds the result to the initial input of the BottleNeckTrue structure through a residual structure; the BottleNeckFalse structure first performs convolution through a 1×1 CBS structure, and then performs convolution through a 3×3 CBS structure.
[0023] Furthermore, the RTFT structure includes an image chunk processing sub-structure, an image patch embedding sub-structure, a position encoding sub-structure, a Transformer encoder, and a multi-layer perceptron;
[0024] The image chunk processing sub-structure is used for image preprocessing. After the feature maps of RGB images and depth images are processed by CBS, the feature maps are uniformly changed into , and the image chunk processing sub-structure is used to divide the feature map into pieces of patches, and then flatten each patch, and the resulting data dimension is , where N is the sequence length input to the Transformer encoder, C is the number of channels of the input feature map, and P is the size of the image patch;
[0025] The image patch embedding substructure is used to convert the vector dimension of to a two-dimensional input of size
[0026] and perform image patch embedding;
[0027] The Transformer encoder includes a connected Layer Normalization substructure and Multi-Head Attention substructure. The Layer Normalization substructure is used to calculate the mean and variance on each sample, subtract the mean of each column from each element in the column, and then divide by the standard deviation of the column to obtain a normalized value that satisfies the standard normal distribution;
[0028] The Multi-Head Attention substructure includes the stacking of multiple self-attention layers. Each self-attention layer calculates three new vectors Query, Key, and Value. All three vectors are the results of multiplying the embedding vector by a randomly initialized matrix, which is updated during backpropagation; by calculating the dot product of the vectors Query and Key, and then dividing the result of the dot product by the length of the vector Query, the attention is calculated to obtain the weight distribution of the vector Value;
[0029] The multi-layer perceptron is a fully connected layer, which is used to separate the fused features again to obtain the fused RGB features and infrared features.
[0030] Further, the Neck module includes an FPN structure and a PAN structure;
[0031] The FPN structure is a top-down feature pyramid, which is used to transfer and fuse the high-level feature information through upsampling to obtain the feature layer for prediction and enhance the semantic information of the feature pyramid;
[0032] The PAN structure is a bottom-up feature pyramid, which is used to supplement the FPN structure and transfer the low-level features to the high level.
[0033] Further, the Lableme software is used to complete the annotation of the tracking target in the collected images.
[0034] Further, the camera is the D4535i depth camera, and the infrared imager is the FLIR infrared imager.
[0035] Further, the trolley is also provided with a pan-tilt structure, and both the camera and the infrared imager are mounted on the pan-tilt structure;
[0036] During the process of judging the steering angle of the trolley, if the tracking target is lost, active search is performed by rotating the pan-tilt.
[0037] The present invention also provides a multi-modal target following system for a follow-up trolley under night conditions, including a trolley, on which a camera and an infrared imager are provided. The trolley also includes a memory and a processor. The memory stores a computer program, and the processor calls the computer program to execute the steps of the method described above.
[0038] Compared with the prior art, the present invention has the following advantages:
[0039] (1) The present invention selects the YOLOv5-RTFT algorithm improved for the current task. By introducing the Transformer structure commonly used when processing text sequences in natural language processing and making improvements, in the Transformer architecture, the Decoder structure is deleted, and the image information is segmented into multiple patches to obtain an RTFT structure that can fuse RGB image features and Thermal image features. Through training and migration, the accuracy of the model in night recognition is improved;
[0040] (2) The present invention uses the FPN and PAN structures to fuse multi-scale information, so that both small objects and large objects can achieve better results;
[0041] (3) After the present invention identifies the selected target, it determines the moving direction of the trolley by judging whether the selected target is at the center of the camera's field of view, and determines whether the trolley moves according to the depth sensor. The accuracy of the result obtained by this method is higher than that of the traditional YOLOv5 target detection method. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 It is a schematic flow chart of a multi-modal target following method for a follow-up trolley under night conditions provided in an embodiment of the present invention;
[0043] Figure 2 It is a schematic diagram of an RGB-Thermal Fusion Transformer module improved based on Vision Transformer provided in an embodiment of the present invention;
[0044] Figures 3a - 3dIt is a network schematic diagram of each part of YOLOv5-RTFT improved based on YOLOv5 provided in the embodiments of the present invention;
[0045] Figure 4 It is a schematic diagram of the calculation method of a target following method provided in the embodiments of the present invention. Specific embodiments
[0046] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Usually, the components of the embodiments of the present invention described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations.
[0047] Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed present invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0048] It should be noted that similar reference numerals and letters indicate similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.
[0049] In the description of the present invention, it should be noted that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, or the orientation or positional relationship in which the product of the invention is usually placed during use. It is only for the convenience of describing the present invention and simplifying the description, and does not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation of the present invention.
[0050] It should be noted that the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present application, "a plurality" means two or more unless otherwise specifically defined.
[0051] In addition, terms such as "horizontal" and "vertical" do not require the components to be absolutely horizontal or hanging vertically, but can be slightly inclined. For example, "horizontal" only means that its direction is more horizontal relative to "vertical", and does not mean that the structure must be completely horizontal, but can be slightly inclined.
[0052] Embodiment 1
[0053] As Figure 1 shown, this embodiment provides a multi-modal target following method for a follow-up vehicle under night conditions, including the following steps:
[0054] S1: During the movement of the vehicle, various attitude images of the target to be detected are collected through a camera and an infrared imager, and image annotation is performed to produce a training data set;
[0055] S2: The data in the training data set is input into a pre-constructed YOLOv5-RTFT target detection network for training to obtain a trained target detection model. The YOLOv5-RTFT target detection network is a dual-channel network structure based on YOLOv5, and an RTFT structure is introduced. This RTFT structure deletes the Decoder structure on the basis of the Transformer architecture, divides the image information into multiple patches, so as to fuse RGB image features and infrared image features;
[0056] S3: During the target following process of the vehicle, an RGB image and a depth image are collected through a camera, and an infrared image is collected through an infrared imager;
[0057] S4: The collected RGB image and infrared image are input into the trained target detection model to obtain a detection result;
[0058] S5: According to the detection result, the difference between the center coordinates of the tracking target and the center coordinates of the viewfinder frame is obtained, so as to judge the steering angle of the vehicle, so that the center point of the tracking target remains the center point of the viewfinder frame, as Figure 4 shown;
[0059] S6: The coordinates of the tracking target on the RGB image are mapped to the depth image to obtain the distance between the tracking target and the vehicle, which is used to judge whether to move forward to achieve target following.
[0060] The following is a specific description of each step.
[0061] In step S1, the Lableme software is used to complete the annotation of the tracking target in the collected images.
[0062] As Figures 3a - 3cAs shown, the backbone network of the YOLOv5-RTFT object detection network described in step S2 includes an input end module, a Backbone module, a Neck module, and a Prediction module. The backbone network is used to extract features from RGB and Thermal images to obtain features that are shallow and deep fused; the object detection branch layer is used to perform object, classification, and regression predictions based on the detection branch feature map to obtain the object detection branch result.
[0063] 1. Input end module
[0064] The input end module uses the Mosaic technology for data augmentation, enriches the data set and reduces GPU usage, and adopts an adaptive Anchor calculation method to adjust the calculated Anchor;
[0065] The process of using the Mosaic technology for data augmentation includes: using the Mosaic technology to randomly crop, scale four pictures in the input data set and then randomly splice them into one picture to achieve data set expansion;
[0066] Specifically, since obtaining a neural network model with good performance often requires a large amount of data as support, and obtaining a large data set often requires a lot of time and labor costs. By using the Mosaic data augmentation technology, the computer can be fully utilized to generate data, increase the data volume. Mosaic uses four pictures, randomly crops and scales the four pictures and then randomly splices them into one picture. While enriching the data set, it increases small-sample objects, improves the training speed of the network, and calculates 4 pictures at one time during the normalization operation, reducing CPU usage;
[0067] The process of the adaptive Anchor calculation method includes: before the start of training, calculate the width and height of all objects in the data set input to the network, so as to calculate the best recall rate of the annotation information of this data set for the default Anchor. If the best recall rate meets the preset recall rate requirement, the Anchor is not updated, otherwise, recalculate the Anchor of this data set.
[0068] Specifically, in YOLOv5-RTFT, the default Anchor is not only used. Before each training starts, the Anchor will be adaptively calculated according to different data sets. First, obtain the width and height of all objects in the data set, and calculate the best recall rate of the annotation information of this data set for the default Anchor. When the best recall rate is greater than or equal to 0.98, the Anchor does not need to be updated. If the best recall rate is less than 0.98, the Anchor that conforms to this data set needs to be recalculated.
[0069] 2. Backbone module
[0070] The Backbone module includes a CBS structure, an RTFT structure, and a BottleNeck structure;
[0071] 2.1. CBS Structure
[0072] The CBS structure includes a Conv layer, a Batch Normalization layer, and a SiLU layer connected in series in sequence; the Conv layer includes a 1×1 convolutional layer and a 3×3 convolutional layer, and both the 1×1 convolutional layer and the 3×3 convolutional layer are used to expand the feature maps of RGB images and depth images; the Batch Normalization layer is used to perform normalization processing on an entire feature map as a single neuron using a weight sharing strategy; the SiLU layer is an activation function layer based on SiLU.
[0073] Specifically, the CBS structure is composed of a Conv layer, a Batch Normalization layer, and a SiLU layer connected in series. The Conv layer used in CBS includes two types: a 1×1 convolutional layer and a 3×3 convolutional layer. The 1×1 convolutional layer expands the RGB and Thermal feature maps into feature maps , and the 3×3 convolutional layer expands the RGB and Thermal feature maps into feature maps ; once the network is trained, the parameters inside will definitely be updated. Except for the data in the input layer (because for the input layer data, we have manually normalized each sample), the distribution of the input data for each subsequent layer of the network is changing, greatly reducing the training speed of the network. Therefore, we need to perform normalization processing on the data. The Batch Normalization layer performs normalization processing on each neuron, or even only needs to normalize a single neuron, rather than normalizing all the neurons in an entire layer of the network. Using Batch Normalization in convolution utilizes the weight sharing strategy and treats an entire feature map as a single neuron; , different from LeakyReLU, the activation of SiLU is not monotonically increasing and has self-stabilizing characteristics.
[0074] 2.2. RTFT Structure
[0075] As Figure 2 shown, the RTFT structure is used to fuse the features of RGB images and depth images based on the Vision Transformer structure;
[0076] The RTFT structure includes an image chunk processing sub-structure, an image patch embedding sub-structure, a position encoding sub-structure, a Transformer encoder, and a multi-layer perceptron;
[0077] The image block processing sub-structure is used for image preprocessing. After the feature maps of the RGB image and the depth image are processed by CBS, the feature maps are uniformly transformed into , and the image block processing sub-structure is used to divide the feature map into pieces of patch, and then each patch is flattened, and the data dimension obtained is , where N is the sequence length input to the Transformer encoder, C is the number of channels of the input feature map, and P is the size of the image patch;
[0078] The image patch embedding sub-structure is used to convert the vector dimension into a two-dimensional input of size and perform image patch embedding;
[0079] The position encoding sub-structure is used to append position information during the image patch embedding process;
[0080] The Transformer encoder includes a connected Layer Normalization sub-structure and a Multi-Head Attention sub-structure. The Layer Normalization sub-structure is used to calculate the mean and variance on each sample, subtract the mean of each column from each element of each column, and then divide by the standard deviation of the column, so as to obtain the normalized value that satisfies the standard normal distribution;
[0081] The Multi-Head Attention sub-structure includes the stacking of multiple self-attention layers. Each self-attention layer calculates three new vectors Query, Key, and Value. The three vectors are all the results of multiplying the embedding vector by a randomly initialized matrix, and the matrix is updated during the backpropagation process; by calculating the dot product of the vectors Query and Key, and then dividing the result of the dot product by the length of the vector Query, the attention is calculated to obtain the weight distribution of the vector Value;
[0082] The multi-layer perceptron is a fully connected layer, which is used to separate the fused features again to obtain the fused RGB features and infrared features.
[0083] Specifically, the RTFT (RGB-Thermal Fusion Transformer) structure is based on the Vision Transformer structure and fuses RGB image features and Thermal image features. The core processes include image chunking (make patches), patch embedding, position encoding, the Transformer encoder, and a multi-layer perceptron;
[0084] Image chunking can be regarded as an image preprocessing step. In this patent, after the RGB image and Thermal image are respectively processed by CBS, the feature maps are uniformly transformed into , and now they are divided into patches, so there are actually patches. The dimension of all patches can be written as . Then each patch is flattened, and the corresponding data dimension can be written as . N can be understood as the sequence length input to the Transformer, C is the number of channels of the input feature map, and P is the size of the image patch;
[0085] Patch embedding can be regarded as a preprocessing process that transforms the vector dimension into a two-dimensional input of size . An operation of patch embedding is also required, similar to word embedding in natural processing. Patch embedding is also a way to transform a high-dimensional vector into a low-dimensional vector. In fact, it performs a linear transformation on each flattened patch vector, that is, a fully connected layer, and the dimension after dimensionality reduction is D;
[0086] The position encoding is set because, different from convolutional neural networks, in RTFT, the spatial position information of the pathes in the sequence data is not known, and a position information needs to be appended to the patch embedding, that is, the Figure 2 Position vector in. In RTFT, learnable position encoding vectors are used and added to the corresponding output patch embeddings;
[0087] The Transformer encoder mainly includes Layer Normalization and Multi-Head Attention. The function of Layer Normalization is also to normalize the hidden layer in the neural network to a standard normal distribution, so as to accelerate the training speed and convergence. However, different from Batch Normalization, Layer Normalization calculates the mean and variance on each sample. Each element in each column is subtracted by the mean of this column, and then divided by the standard deviation of this column to obtain the normalized value. BN is the scaling of the internal features of the sample, while LN is the scaling of all features between samples;
[0088] For Multi-Head Attention, it is actually the superposition of multiple self-attention. First, self-attention will calculate three new vectors Query, Key, and Value. These three vectors are the results of multiplying the embedding vector by a matrix. This matrix is randomly initialized and its value will be continuously updated during the backpropagation process. Then, by calculating the dot product of Query and Key, the self-attention score value is obtained. This score value determines the degree of attention to other parts of the input patches when we encode a patch. Next, divide the result of the dot product by a , and the attention calculation formula can be written as . This method of determining the weight distribution of value through the similarity degree of query and key is called scaled dot-product attention, and Multi-Head Attention means that not only one group of Q, K, and V matrices are initialized, but multiple groups are initialized.
[0089] The multi-layer perceptron is actually a fully connected layer that plays a role in classification. It separates the fused features again to obtain the fused RGB features and the fused Thermal features .
[0090] 2.3. BottleNeck Structure
[0091] The Bottleneck structure is the BottleneckTrue structure or the BottleneckFalse structure. The BottleneckTrue structure first performs convolution through a 1×1 CBS structure, then performs convolution through a 3×3 CBS structure, and finally adds the result to the initial input of the BottleneckTrue structure through a residual structure; the BottleneckFalse structure first performs convolution through a 1×1 CBS structure, and then performs convolution through a 3×3 CBS structure.
[0092] Specifically, the Bottleneck structure is different from the Bottleneck in ResNet. There are two types of Bottlenecks in YOLOv5-RTFT: one is BottleneckTrue, which first performs convolution through a 1×1 CBS structure, then through a 3×3 CBS structure, and finally adds the result to the initial input through a residual structure; the other is BottleneckFalse, which first performs convolution through a 1×1 CBS structure and then through a 3×3 CBS convolution, but without adding a residual structure.
[0093] 3. Neck module
[0094] The Neck module includes an FPN structure and a PAN structure;
[0095] The FPN structure is a top-down feature pyramid, which is used to transfer and fuse high-level feature information through upsampling to obtain a feature layer for prediction and enhance the semantic information of the feature pyramid;
[0096] Specifically, the FPN structure is top-down. It transfers and fuses high-level feature information through upsampling to obtain a feature layer for prediction, enhances the semantic information, and enlarges the small feature map at the top layer to the size of the feature map of the previous stage through upsampling. Figure 1 That is, it utilizes the strong semantic features at the top layer (beneficial for classification) and the high-resolution information at the bottom layer (beneficial for localization). The upsampling method is implemented using nearest neighbor interpolation. The combination method uses a residual structure similar to that in ResNet for lateral connection. Different from FPNNet, here the connection uses Concat for dimension expansion;
[0097] The PAN structure is a bottom-up feature pyramid, which is used to supplement the FPN structure and transfer the bottom layer features to the high layer.
[0098] Specifically, the PAN structure is bottom-up. High-level features have a large receptive field and obtain rich context information, allowing small candidates to obtain these features and better use this information for prediction. Low-level features have many tiny details and high localization accuracy, allowing large candidates to obtain these features. The feature pyramid conveys strong localization features from bottom to top.
[0099] In this embodiment, the camera is the D4535i depth camera. The D435i is used to obtain a video sequence, from which an RGB image and a depth image are obtained. The infrared imager is a FLIR infrared imager.
[0100] Preferably, the trolley is further provided with a pan-tilt structure, and both the camera and the infrared imager are mounted on the pan-tilt structure;
[0101] If the target is lost during tracking in step S5, the functional trolley uses the rotation of the pan-tilt and re-uses the target detection algorithm to determine whether the target is within a set threshold distance. This is a re-binding process to ensure that the trolley can always stably track the target.
[0102] In step S6, the distance from the tracking target to the functional vehicle is obtained based on the depth image, and it is determined whether to move forward according to the set threshold.
[0103] In summary, the multi-modal follow-up trolley target following method based on YOLOv5-RTFT of the present invention has strong robustness and high accuracy. Compared with traditional target detection methods, it can achieve better results both during the day and at night. And for the problem of possible tracking loss during the tracking process, the present invention adds a pan-tilt to actively search and bind the selected target. The multi-modal follow-up trolley target following method based on YOLOv5-RTFT proposed by the present invention mainly features:
[0104] 1) To improve the accuracy of trolley tracking at night, the RTFT module is designed to deeply fuse RGB features and Thermal features;
[0105] 2) To simultaneously input the RGB image and the Thermal image, a dual-channel DarkNet network structure is designed, adopting a symmetric structure for feature-level fusion;
[0106] 3) To enable the trolley to stably track the selected target, during the target following process, a pan-tilt structure mounted on the trolley is designed. When the target is lost, the pan-tilt rotates to actively search for the selected target.
[0107] The present invention also provides a multi-modal target following system for a follow-up vehicle under night conditions, including a vehicle, on which a camera and an infrared imager are provided. The vehicle further includes a memory and a processor. The memory stores a computer program, and the processor calls the computer program to execute the steps of the multi-modal target following method for the follow-up vehicle under night conditions as described above.
[0108] The preferred specific embodiments of the present invention have been described in detail above. It should be understood that those of ordinary skill in the art can make many modifications and variations based on the concept of the present invention without creative efforts. Therefore, all technical solutions that can be obtained by those skilled in the art in the technical field of the present invention through logical analysis, reasoning or limited experiments based on the concept of the present invention on the basis of the prior art shall fall within the protection scope determined by the claims.
Claims
1. A multi-modal target following method for a servo car under night conditions, characterized in that, It includes the following steps: During the movement of the trolley, various attitude images of the target to be detected are collected by a camera and an infrared imager, and image annotation is performed to make a training data set; The data in the training data set is input into a pre-constructed YOLOv5-RTFT object detection network for training to obtain a trained object detection model. The YOLOv5-RTFT object detection network is a dual-path network structure based on YOLOv5, and an RTFT structure is introduced. Based on the Transformer architecture, the Decoder structure is deleted, and the image information is segmented into multiple patches, so as to fuse RGB image features and infrared image features; During the target following process of the trolley, an RGB image and a depth image are collected by a camera, and an infrared image is collected by an infrared imager; The collected RGB image and infrared image are input into the trained object detection model to obtain a detection result; According to the detection result, the difference between the center coordinates of the tracking target and the center coordinates of the viewfinder frame is obtained, so as to judge the steering angle of the trolley, so that the center point of the tracking target remains the center point of the viewfinder frame; The coordinates of the tracking target on the RGB image are mapped to the depth image to obtain the distance between the tracking target and the trolley, which is used to judge whether to move forward to achieve target following; The YOLOv5-RTFT object detection network includes an input end module, a Backbone module, a Neck module, and a Prediction module; The Backbone module includes a CBS structure, an RTFT structure, and a BottleNeck structure; The RTFT structure includes an image block processing sub-structure, an image block embedding sub-structure, a position encoding sub-structure, a Transformer encoder, and a multi-layer perceptron; The image block processing sub-structure is used for image preprocessing. After the feature maps of the RGB image and the depth image are processed by CBS, the feature maps are uniformly changed to , and the image block processing sub-structure is used to divide the feature map into pieces of patch, and then flatten each patch, and the data dimension obtained is , where N is the sequence length input to the Transformer encoder, C is the number of channels of the input feature map, and P is the size of the image patch; The image block embedding substructure is used to convert the vector dimension of into a two-dimensional input of the size of and perform image block embedding; The position encoding sub-structure is used to append position information during the image block embedding process; The Transformer encoder includes a connected Layer Normalization sub-structure and a Multi-HeadAttention sub-structure. The Layer Normalization sub-structure is used to calculate the mean and variance on each sample, subtract the mean of each column from each element of each column, and then divide by the standard deviation of each column, so as to obtain a normalized value that satisfies the standard normal distribution; The Multi-Head Attention sub-structure includes the superposition of multiple self-attention layers. Each self-attention layer calculates three new vectors Query, Key, and Value. The three vectors are all the results of multiplying the embedding vector by a randomly initialized matrix, and the matrix is updated during the backpropagation process; by calculating the dot product of the vectors Query and Key, and then dividing the result of the dot product by the length of the vector Query, the attention is calculated to obtain the weight distribution of the vector Value; The multi-layer perceptron is a fully connected layer, which is used to separate the fused features again to obtain the fused RGB features and infrared features.
2. The multi-modal target following method for a servo car under night conditions according to claim 1, wherein The input end module uses the Mosaic technology for data augmentation and adopts an adaptive Anchor calculation method to adjust the calculated Anchor; The process of using the Mosaic technology for data augmentation includes: randomly cropping and scaling four pictures in the input data set by using the Mosaic technology and then randomly splicing them into one picture to achieve data set expansion; The process of the adaptive Anchor calculation method includes: before the start of training, calculating the width and height of all targets in the data set input to the network, thereby calculating the best recall rate of the annotation information of this data set for the default Anchor. If the best recall rate meets the preset recall rate requirement, the Anchor is not updated, otherwise the Anchor of this data set is recalculated.
3. The multi-modal target following method for a servo car under night conditions according to claim 1, characterized in that The CBS structure includes a Conv layer, a Batch Normalization layer, and a SiLU layer connected in series in sequence; the Conv layer includes a 1×1 convolutional layer and a 3×3 convolutional layer, and both the 1×1 convolutional layer and the 3×3 convolutional layer are used to expand the feature maps of the RGB image and the depth image; the Batch Normalization layer is used to normalize an entire feature map as a neuron by using a weight sharing strategy; the SiLU layer is an activation function layer based on SiLU; The RTFT structure is used to fuse the features of the RGB image and the depth image based on the Vision Transformer structure; The BottleNeck structure is a BottleNeckTrue structure or a BottleNeckFalse structure. The BottleNeckTrue structure first performs convolution through a 1×1 CBS structure, then performs convolution through a 3×3 CBS structure, and finally adds the result to the initial input of the BottleNeckTrue structure through a residual structure; the BottleNeckFalse structure first performs convolution through a 1×1 CBS structure, and then performs convolution through a 3×3 CBS structure.
4. A multi-modal target following method for a servo car under night conditions according to claim 1, characterized in that The Neck module includes an FPN structure and a PAN structure; The FPN structure is a top-down feature pyramid, which is used to transfer and fuse the high-level feature information through upsampling to obtain the feature layer for prediction and enhance the semantic information of the feature pyramid; The PAN structure is a bottom-up feature pyramid, which is used to supplement the FPN structure and transfer the low-level features to the high level.
5. A multi-modal target following method for a follow-up vehicle under night conditions according to claim 1, characterized in that Use the Lableme software to complete the annotation of the tracking target in the collected images.
6. A multi-modal target following method for a servo car under night conditions according to claim 1, characterized in that The camera is the D4535i depth camera, and the infrared imager is the FLIR infrared imager.
7. A multi-modal target following method for a servo car under night conditions according to claim 1, characterized in that The trolley is also provided with a pan-tilt structure, and both the camera and the infrared imager are installed on the pan-tilt structure; During the process of judging the steering angle of the trolley, if the tracking target is lost, active search is performed by rotating the pan-tilt.
8. A multi-modal target following system for a servo car under night conditions, including a car, on which a camera and an infrared imager are provided, characterized in that, The trolley further includes a memory and a processor, the memory stores a computer program, and the processor calls the computer program to execute the steps of the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Transform-based optical remote sensing target detection method
CN114821357A
Adaptive target detection method in strong / weak illumination and fog environment
CN115375991A