Track small target detection method and device, electronic equipment and storage medium

The small target detection model constructed by the improved Yolo11 algorithm uses the feature extraction, alignment and fusion methods of the backbone network, neck network and head network to solve the problem of inaccurate detection of small targets in the track and achieve higher detection accuracy.

CN120807890APending Publication Date: 2025-10-17BEIJING CENTURY DONGFANG COMMUNICATION EQUIPMENT CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510944962.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-09
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing technologies have difficulty accurately detecting small targets in tracks, such as pedestrians or fallen rocks, resulting in inaccurate detection results.

Method used

The small target detection model constructed by the improved Yolo11 algorithm performs feature extraction, alignment and fusion through the combination of backbone network, neck network and head network to generate fusion features, and then performs small target detection based on the fusion features.

Benefits of technology

It improves the detection accuracy of small targets, reduces the spatial offset problem caused by feature downsampling, and can more accurately detect pedestrians or fallen rocks on the track.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120807890A_ABST
    Figure CN120807890A_ABST
Patent Text Reader

Abstract

The invention provides a track small target detection method and device, electronic equipment and a storage medium, and relates to the technical field of artificial intelligence, and the method comprises the steps: obtaining a to-be-detected image; inputting the to-be-detected image into a pre-trained small target detection model to obtain a small target detection result output by the small target detection model; the small target detection model is obtained by pre-training an initial model constructed based on an improved yo11 algorithm based on a sample detection image and a sample small target detection result corresponding to the sample detection image; and the small target detection model is used for carrying out feature extraction on the to-be-detected image, generating image features, carrying out feature alignment and feature fusion on the image features, generating fusion features, and carrying out small target detection based on the fusion features. Through the mode, small targets such as pedestrians or rockfall and the like in the to-be-detected image can be accurately detected, and the accuracy of a track small target detection result is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence, and in particular to a track small target detection method and device, electronic equipment and storage medium. BACKGROUND

[0002] With the development of deep learning, the visual target detection problem has been well solved. The deep learning algorithm automatically learns the multi-level features of the image through the convolutional neural network (CNN), breaks away from the limitation of relying on manually designed features in traditional methods, significantly improves the feature expression ability in complex scenes, and is conducive to the detection and recognition of image targets.

[0003] Generally, the target detection of an image needs to complete target positioning (i.e. coordinate regression) and target classification at the same time. The mainstream methods can be divided into two categories: one is a two-stage detection method, that is, first generate a candidate region, and then perform target classification and coordinate regression; the other is a single-stage detection method, which directly predicts the position and classification of the target through an end-to-end neural network. The two-stage detection method is slower due to the existence of the candidate region generation link, while the single-stage detection method can achieve more balanced performance in terms of detection accuracy and speed.

[0004] However, when the target in the image is small, the existing target detection method is difficult to accurately realize the detection and recognition of small targets. Due to the characteristics of small target size, lack of feature information, and easy to be disturbed by the background, the following key problems exist in the prior art: first, the model has insufficient feature extraction capability for small targets, that is, the traditional target detection network extracts the semantic features of small targets through deep layer convolution downsampling, but multiple downsampling operations will cause the serious loss of the detailed features of small targets. The shallow features of small targets can retain high-resolution information, but their semantic abstraction ability is weak, while the deep features of small targets have the problem of insufficient feature map resolution, which ultimately makes it difficult for the model to distinguish small targets from noise; second, the model does not sufficiently fuse the multi-scale features of small targets, and the shallow features and deep features of small targets are difficult to effectively combine, which further leads to the difficulty of the model in accurately detecting small targets in the image.

[0005] Due to the above-mentioned defects of the existing small target detection method, when this small target detection method is applied to track target detection, although the model can accurately recognize large targets such as tracks or trains, it is difficult to accurately detect small targets such as pedestrians or falling rocks, resulting in inaccurate track small target detection results. SUMMARY

[0006] The present application provides a track small target detection method, device, electronic equipment and storage medium, to solve the defects that the prior art cannot accurately detect small targets such as pedestrians or falling stones, resulting in inaccurate track small target detection results.

[0007] The present application provides a track small target detection method, comprising: acquiring a to-be-detected image; the to-be-detected image is obtained after photographing a track area containing small targets, and the small targets include pedestrians and / or falling stones; inputting the to-be-detected image into a pre-trained small target detection model to obtain a small target detection result output by the small target detection model; wherein the small target detection model is obtained by pre-training an initial model based on an improved yolo11 algorithm based on a sample detection image and a sample small target detection result corresponding to the sample detection image; the small target detection model is used for feature extraction of the to-be-detected image, generating image features, feature alignment and feature fusion of the image features, generating fused features, and small target detection based on the fused features.

[0008] According to the track small target detection method provided by the present application, the small target detection model comprises a backbone network, a neck network and a head network connected in sequence; wherein the backbone network is used for feature extraction of the to-be-detected image to generate image features; the neck network is used for feature alignment and feature fusion of the image features to generate fused features; and the head network is used for small target detection based on the fused features to generate a small target detection result.

[0009] According to the track small target detection method provided by the present application, the backbone network comprises a first network layer, a second network layer, a third network layer, a fourth network layer and a fifth network layer connected in sequence, and the image features comprise a first image feature, a second image feature, a third image feature and a fourth image feature; wherein the first network layer is used for feature extraction of the to-be-detected image to generate initial image features; the second network layer is used for feature extraction of the initial image features to generate the first image features; the third network layer is used for feature extraction of the first image features to generate the second image features; the fourth network layer is used for feature extraction of the second image features to generate the third image features; and the fifth network layer is used for feature extraction of the third image features to generate the fourth image features.

[0010] The neck network comprises an MLP module, a feature alignment network and a feature fusion network connected in sequence, the feature alignment network and the feature fusion network are deformable convolution networks, and the fusion features comprise first fusion features, second fusion features, third fusion features and fourth fusion features; wherein the MLP module is used for performing channel screening processing on the first image features, the second image features, the third image features and the fourth image features respectively to generate first feature maps, second feature maps, third feature maps and fourth feature maps; the feature alignment network is used for performing feature alignment on the first feature maps, the second feature maps, the third feature maps and the fourth feature maps to obtain first to-be-fused feature maps, second to-be-fused feature maps, third to-be-fused feature maps and fourth to-be-fused feature maps; and the feature fusion network is used for performing feature fusion based on the first to-be-fused feature maps, the second to-be-fused feature maps, the third to-be-fused feature maps and the fourth to-be-fused feature maps to generate the first fusion features, the second fusion features, the third fusion features and the fourth fusion features respectively.

[0011] The head network comprises a first prediction head, a second prediction head, a third prediction head and a fourth prediction head; wherein the second prediction head, the third prediction head and the fourth prediction head are used for performing target detection based on the second fusion features, the third fusion features and the fourth fusion features to generate track target detection results; and the first prediction head is used for performing small target detection based on the first fusion features to generate small target detection results; the small target detection results comprise small target position information and small target category information.

[0012] The small target detection model is pre-trained based on a preset loss function, and the preset loss function is determined based on a positioning loss, a classification loss, a confidence loss and a self-supervised contrast loss.

[0013] Before the image to be detected is input into the pre-trained small target detection model to obtain small target detection results output by the small target detection model, the method further comprises: constructing a data set; the data set comprises a plurality of sample detection images and sample small target detection results corresponding to each sample detection image collected in a test line scene and an actual running line scene; the data set is divided into a training set, a test set and a verification set; an initial model constructed based on an improved yolo11 algorithm is pre-trained based on the training set and the verification set to obtain the small target detection model; and performance index testing is performed on the small target detection model based on the test set.

[0014] The application further provides a track small target detection device, comprising: an acquisition module, configured to acquire a to-be-detected image; the to-be-detected image is obtained by photographing a track area containing small targets, and the small targets include pedestrians and / or falling rocks; a small target detection module, configured to input the to-be-detected image into a pre-trained small target detection model to obtain a small target detection result output by the small target detection model; wherein the small target detection model is obtained by pre-training an initial model based on an improved yolo11 algorithm based on a sample detection image and a sample small target detection result corresponding to the sample detection image; the small target detection model is configured to perform feature extraction on the to-be-detected image to generate image features, perform feature alignment and feature fusion on the image features to generate fused features, and perform small target detection based on the fused features.

[0015] The application further provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements any of the above track small target detection methods when executing the computer program.

[0016] The application further provides a non-transitory computer-readable storage medium having a computer program stored thereon, wherein the computer program is executable on a processor to implement any of the above track small target detection methods.

[0017] The track small target detection method, device, electronic device, and storage medium provided by the application acquire a to-be-detected image; the to-be-detected image is obtained by photographing a track area containing small targets, and the small targets include pedestrians and / or falling rocks; the to-be-detected image is input into a pre-trained small target detection model to obtain a small target detection result output by the small target detection model; wherein the small target detection model is obtained by pre-training an initial model based on an improved yolo11 algorithm based on a sample detection image and a sample small target detection result corresponding to the sample detection image; the small target detection model is configured to perform feature extraction on the to-be-detected image to generate image features, perform feature alignment and feature fusion on the image features to generate fused features, and perform small target detection based on the fused features. In this way, the fusion method based on feature alignment is introduced into the small target detection model based on the improved yolo11 algorithm, the small target detection model can perform feature extraction on the to-be-detected image to generate image features, and then perform feature alignment and feature fusion on the image features to generate fused features. Since the fusion method based on feature alignment can effectively fuse image features of different levels and reduce the spatial offset problem caused by feature down-sampling, it is beneficial for the model to accurately locate small targets, and small target detection is performed based on the fused features, so that small targets such as pedestrians or falling rocks in the to-be-detected image can be accurately detected, and the accuracy of the track small target detection result is improved. BRIEF DESCRIPTION OF DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the present application or the prior art, the drawings needed to be used in the embodiments or prior art description will be briefly introduced. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative effort.

[0019] Figure 1 is a flowchart of the track small target detection method provided by the present application.

[0020] Figure 2 is a structural diagram of the small target detection model provided by the present application.

[0021] Figure 3 is a structural diagram of the feature fusion network provided by the present application.

[0022] Figure 4 is a structural diagram of the feature alignment network provided by the present application.

[0023] Figure 5 is a training flowchart of the small target detection model provided by the present application.

[0024] Figure 6 is a structural diagram of the track small target detection device provided by the present application.

[0025] Figure 7 is a structural diagram of the electronic device provided by the present application. DETAILED DESCRIPTION

[0026] In order to make the purpose, technical solutions and advantages of the present application clearer, the technical solutions in the present application will be described clearly and completely in combination with the drawings in the present application. Obviously, the described embodiments are some embodiments of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative effort belong to the protection scope of the present application.

[0027] Please refer to Figures 1 to 5 , Figure 1 is a flowchart of the track small target detection method provided by the present application, Figure 2 is a structural diagram of the small target detection model provided by the present application, Figure 3 is a structural diagram of the feature fusion network provided by the present application, Figure 4 is a structural diagram of the feature alignment network provided by the present application, Figure 5 is a training flowchart of the small target detection model provided by the present application.

[0028] As Figure 1As shown, in the embodiment, the track small target detection method includes steps S110 to S120, and each step is specifically as follows: S110: Obtain a to-be-detected image.

[0029] The to-be-detected image is obtained after photographing a track area containing small targets, and the small targets include pedestrians and / or falling rocks.

[0030] S120: Input the to-be-detected image into a pre-trained small target detection model to obtain a small target detection result output by the small target detection model.

[0031] The small target detection model is obtained by pre-training an initial model based on an improved yolo11 algorithm based on a sample detection image and a sample small target detection result corresponding to the sample detection image.

[0032] The small target detection model is used for feature extraction on the to-be-detected image to generate image features, feature alignment and feature fusion on the image features to generate fused features, and small target detection based on the fused features.

[0033] The track small target detection method provided in the embodiment obtains a to-be-detected image, the to-be-detected image is obtained after photographing a track area containing small targets, and the small targets include pedestrians and / or falling rocks, inputs the to-be-detected image into a pre-trained small target detection model to obtain a small target detection result output by the small target detection model, and the small target detection model is obtained by pre-training an initial model based on an improved yolo11 algorithm based on a sample detection image and a sample small target detection result corresponding to the sample detection image. The small target detection model is used for feature extraction on the to-be-detected image to generate image features, feature alignment and feature fusion on the image features to generate fused features, and small target detection based on the fused features. In this way, the fusion method based on feature alignment is introduced into the small target detection model constructed based on the improved yolo11 algorithm, the small target detection model can perform feature extraction on the to-be-detected image to generate image features, and then perform feature alignment and feature fusion on the image features to generate fused features. Since the fusion method based on feature alignment can effectively fuse image features of different levels and reduce the spatial offset problem caused by feature down-sampling, it is beneficial for the model to accurately locate small targets and perform small target detection based on the fused features, so that small targets such as pedestrians or falling rocks in the to-be-detected image can be accurately detected, and the accuracy of the track small target detection result is improved.

[0034] In some embodiments, the small target detection model comprises a backbone network, a neck network and a head network connected in sequence; wherein the backbone network is configured to perform feature extraction on the to-be-detected image to generate image features; the neck network is configured to perform feature alignment and feature fusion on the image features to generate fused features; and the head network is configured to perform small target detection based on the fused features to generate a small target detection result.

[0035] As shown in Figure 2 , the small target detection model of the present embodiment is improved on the basis of an existing yolo11 network model, and comprises a backbone network, a neck network and a head network connected in sequence.

[0036] The backbone network is configured to perform multi-scale feature extraction on the to-be-detected image to generate a plurality of image features of different scales, and the image features comprise a first image feature, a second image feature, a third image feature and a fourth image feature.

[0037] Specifically, the backbone network comprises a first network layer , a second network layer , a third network layer , a fourth network layer and a fifth network layer connected in sequence; wherein the first network layer is configured to perform feature extraction on the to-be-detected image to generate initial image features; the second network layer is configured to perform feature extraction on the initial image features to generate the first image features; the third network layer is configured to perform feature extraction on the first image features to generate the second image features; the fourth network layer is configured to perform feature extraction on the second image features to generate the third image features; and the fifth network layer is configured to perform feature extraction on the third image features to generate the fourth image features.

[0038] The neck network is configured to perform feature alignment and feature fusion on the plurality of image features of different scales to generate a plurality of fused features, and the fused features comprise a first fused feature, a second fused feature, a third fused feature and a fourth fused feature.

[0039] Specifically, the first image feature and the second image feature are adjacent image features, the second image feature and the third image feature are adjacent image features, and the third image feature and the fourth image feature are adjacent image features; the neck network can perform feature fusion operations on adjacent image features based on feature alignment, that is, before performing feature fusion, feature alignment is performed on adjacent image features respectively, a feature offset prediction network is constructed by using a deformable convolution network (DCNv4 module) to perform offset prediction between high-level features (two adjacent image features), then the predicted offset is used to correct the high-level features layer by layer, and then the corrected high-level features are added to the low-level features to obtain fused features; after the feature alignment and the feature fusion of all adjacent image features are completed, the first fused feature , the second fused feature , the third fused feature , and the fourth fused feature are generated.

[0040] The head network of the existing YOLOv1 network model includes three detection branches, which are used to perform target detection based on the second fused feature , the third fused feature , and the fourth fused feature to generate track target detection results, and the track target detection results include position information and target category information of a track or a train and the like.

[0041] In this embodiment, a high-resolution detection branch is additionally added to the head network of the existing YOLOv1 network model to improve the response to small targets. Since each detection branch includes one detection head, the head network of the small target detection model includes a first prediction head, a second prediction head, a third prediction head, and a fourth prediction head.

[0042] The second prediction head, the third prediction head, and the fourth prediction head are used to perform target detection based on the second fused feature , the third fused feature , and the fourth fused feature to generate track target detection results; and the first prediction head is used to perform small target detection based on the first fused feature to generate small target detection results, and the small target detection results include small target position information and small target category information.

[0043] The track small target detection method provided in this embodiment can effectively improve the response of the model to small target detection by adding a high-resolution feature detection branch to the head network.

[0044] In some embodiments, the backbone network comprises a first network layer, a second network layer, a third network layer, a fourth network layer and a fifth network layer connected in sequence, and the image features comprise a first image feature, a second image feature, a third image feature and a fourth image feature; wherein the first network layer is configured to perform feature extraction on the to-be-detected image to generate an initial image feature; the second network layer is configured to perform feature extraction on the initial image feature to generate the first image feature; the third network layer is configured to perform feature extraction on the first image feature to generate the second image feature; the fourth network layer is configured to perform feature extraction on the second image feature to generate the third image feature; and the fifth network layer is configured to perform feature extraction on the third image feature to generate the fourth image feature.

[0045] As shown in Figure 2 , the backbone network is configured to perform multi-scale feature extraction on the to-be-detected image to generate a plurality of image features of different scales, and the image features comprise a first image feature, a second image feature, a third image feature and a fourth image feature.

[0046] Specifically, the backbone network comprises a first network layer , a second network layer , a third network layer , a fourth network layer and a fifth network layer connected in sequence.

[0047] The first network layer is configured to perform feature extraction on the to-be-detected image to generate an initial image feature; the second network layer is configured to perform feature extraction on the initial image feature to generate the first image feature; the third network layer is configured to perform feature extraction on the first image feature to generate the second image feature; the fourth network layer is configured to perform feature extraction on the second image feature to generate the third image feature; and the fifth network layer is configured to perform feature extraction on the third image feature to generate the fourth image feature.

[0048] Optionally, the step convolution module of each network layer in the backbone network is replaced by an SPD-conv (spatial depth conversion convolution) module, so that the backbone network completes down-sampling while retaining all the information of the image features, which is conducive to subsequent detection of track small targets and reduces the missed detection rate and the false detection rate.

[0049] In some embodiments, the neck network includes an MLP module, a feature alignment network and a feature fusion network connected in sequence, the feature alignment network and the feature fusion network are deformable convolutional networks, and the fusion features include a first fusion feature, a second fusion feature, a third fusion feature and a fourth fusion feature; wherein the MLP module is used to perform channel screening processing on the first image feature, the second image feature, the third image feature and the fourth image feature, respectively, to generate a first feature map, a second feature map, a third feature map and a fourth feature map; the feature alignment network is used to perform feature alignment on the first feature map, the second feature map, the third feature map and the fourth feature map, to obtain a first feature map to be fused, a second feature map to be fused, a third feature map to be fused and a fourth feature map to be fused; the feature fusion network is used to perform feature fusion based on the first feature map to be fused, the second feature map to be fused, the third feature map to be fused and the fourth feature map to be fused, respectively, to generate a first fusion feature, a second fusion feature, a third fusion feature and a fourth fusion feature.

[0050] Specifically, the neck network includes an MLP module, a feature alignment network, and a feature fusion network connected in sequence, and the feature alignment network and the feature fusion network are deformable convolutional networks.

[0051] Among them, the MLP (Multilayer Perceptron) module can perform channel dimension screening on the four image features of different scales (i.e., the first image feature, the second image feature, the third image feature, and the fourth image feature) extracted by the backbone network based on the channel mixed attention mechanism, reduce the channel dimensions of the four image features of different scales, remove redundant features, reduce the computational complexity of the neck network, and finally generate the first feature map, second feature map, third feature map, and fourth feature map after removing redundant features.

[0052] like Figure 4 As shown, the feature alignment network is used to perform feature alignment on the first feature map, the second feature map, the third feature map and the fourth feature map to obtain the first feature map to be fused, the second feature map to be fused, the third feature map to be fused and the fourth feature map to be fused.

[0053] Specifically, the first feature map and the second feature map are adjacent feature maps, the second feature map and the third feature map are adjacent feature maps, and the third feature map and the fourth feature map are adjacent feature maps; for any two adjacent feature maps, the feature alignment network can perform feature alignment on the two feature maps. Since the feature alignment network is a feature offset prediction network constructed using a deformable convolutional network, the feature alignment network can predict the feature offset of the two adjacent feature maps, and then use the predicted offset to perform feature correction to achieve feature alignment.

[0054] For example Figure 4In the middle, for the feature map , the feature alignment network can perform an upsample operation on the feature map , and perform feature offset prediction according to the low-level feature map adjacent thereto, and then perform feature correction on the feature map using the predicted offset to achieve feature alignment and generate an aligned feature map to be fused .

[0055] As shown in Figure 3 , the feature fusion network is configured to perform feature fusion based on the first feature map to be fused, the second feature map to be fused, the third feature map to be fused, and the fourth feature map to be fused, and generate a first fused feature, a second fused feature, a third fused feature, and a fourth fused feature, respectively.

[0056] Specifically, the first feature map to be fused and the second feature map to be fused are adjacent feature maps, the second feature map to be fused and the third feature map to be fused are adjacent feature maps, and the third feature map to be fused and the fourth feature map to be fused are adjacent feature maps; for any two adjacent feature maps to be fused, the feature fusion network can perform addition fusion on the two feature maps to be fused, and generate a first fused feature , a second fused feature , a third fused feature , and a fourth fused feature .

[0057] Among them, the second fused feature is obtained by addition fusion of the first feature map to be fused and the second feature map to be fused, the third fused feature is obtained by addition fusion of the second feature map to be fused and the third feature map to be fused, the fourth fused feature is obtained by addition fusion of the third feature map to be fused and the fourth feature map to be fused, and the first feature map to be fused can be directly used as the first fused feature .

[0058] The track small target detection method provided in the embodiment adopts a multi-scale feature fusion network based on feature alignment, which can better effectively fuse features of different scales, reduce spatial offset problems caused by feature down-sampling, and help the model better locate small targets; at the same time, the neck network adopts an attention mechanism based on channel mixing, which can effectively reduce the feature dimension, reduce redundant features, and reduce the computational amount of the model.

[0059] In some embodiments, the head network comprises a first prediction head, a second prediction head, a third prediction head, and a fourth prediction head; the second prediction head, the third prediction head, and the fourth prediction head are configured to perform target detection based on the second fused feature, the third fused feature, and the fourth fused feature to generate a track target detection result; and the first prediction head is configured to perform small target detection based on the first fused feature to generate a small target detection result; the small target detection result comprises small target position information and small target category information.

[0060] The head network of the existing YOLOv1 network model comprises three detection branches, which are configured to perform target detection based on the second fused feature , the third fused feature , and the fourth fused feature to generate a track target detection result, and the track target detection result comprises position information and target category information of a track or a train or other large-sized target.

[0061] In the present embodiment, a high-resolution detection branch is additionally added to the head network of the existing YOLOv1 network model to improve the response to small targets. Since each detection branch comprises one detection head, the head network of the small target detection model comprises a first prediction head, a second prediction head, a third prediction head, and a fourth prediction head.

[0062] The second prediction head, the third prediction head, and the fourth prediction head are configured to perform target detection based on the second fused feature , the third fused feature , and the fourth fused feature to generate a track target detection result; and the first prediction head is configured to perform small target detection based on the first fused feature to generate a small target detection result, and the small target detection result comprises small target position information and small target category information.

[0063] The track small target detection method provided in the present embodiment can effectively improve the response of the model to small target detection by adding a high-resolution feature detection branch to the head network.

[0064] In some embodiments, the small target detection model is pre-trained based on a preset loss function, and the preset loss function is determined based on a positioning loss, a classification loss, a confidence loss, and a self-supervised contrast loss.

[0065] Specifically, the small target detection model is pre-trained based on a preset loss function, which is determined based on positioning loss, classification loss, confidence loss and self-supervised contrast loss. During the pre-training process, the model can perform supervised learning on the predicted results based on positioning loss, classification loss, confidence loss and self-supervised contrast loss, and continuously update the network weights until the model pre-training is completed.

[0066] Optionally, the network weights are initialized with pre-trained weights, the main training parameters batch is set to 8, epoch is set to 100, the optimizer is optimized using the stochastic gradient descent (SGD) algorithm, the initial learning rate is set to 0.01, and the network scale is set to .

[0067] In some embodiments, before the image to be detected is input into a pre-trained small target detection model and the small target detection result output by the small target detection model is obtained, it also includes: constructing a data set; the data set includes multiple sample detection images collected in the test line scene and the actual operation line scene and the sample small target detection result corresponding to each sample detection image; dividing the data set into a training set, a test set and a validation set; based on the training set and the validation set, pre-training the initial model constructed based on the improved yolo11 algorithm to obtain a small target detection model; based on the test set, performing a performance index test on the small target detection model.

[0068] It is understandable that before using the small target detection model to detect small targets, the small target detection model needs to be trained.

[0069] like Figure 5 As shown, multiple sample detection images are collected in the test line scenario and the sample small target detection results corresponding to each sample detection image are labeled to obtain a long-distance small target dataset; multiple sample detection images are collected in the actual operation line scenario and the sample small target detection results corresponding to each sample detection image are labeled to obtain an actual operation line small target dataset; the long-distance small target dataset and the actual operation line small target dataset are merged to obtain a dataset.

[0070] It should be noted that the sample detection images collected in the test line scenario are images that simulate the real scene. These data are difficult to obtain in the actual operation line scenario and the sample size is small. Therefore, by simulating the real scene, the scarcity of real scene data can be compensated.

[0071] Specifically, for the test line scenario, pedestrians at different distances (e.g., 150 meters to 20 meters) and pedestrians with sizes within the range of 100 meters to 50 meters are collected at different time periods and weather conditions. Scene pictures such as falling rocks within cm are used as sample detection images.

[0072] Optionally, the pedestrians can choose to wear clothes of different colors, and the falling rocks can choose simulated stones of different shapes and different colors.

[0073] Optionally, the sample detection images in the test line scene are at least 11000.

[0074] Specifically, for the actual running line scene, pedestrian data appearing in different railway line monitoring scenes, such as construction personnel data at night, are collected under different monitoring angles and different lighting conditions.

[0075] Optionally, the sample detection images in the actual running line scene are at least 21000.

[0076] Further, the data set is divided into a training set, a test set and a validation set according to a preset ratio (for example, 7:2:1), and based on the training set and the validation set, an initial model constructed based on the improved yolo11 algorithm is pre-trained to obtain a small target detection model.

[0077] Further, based on the test set, the pre-trained small target detection model is tested for performance indicators.

[0078] Optionally, the performance indicator test includes recall rate indicator test and average precision indicator test.

[0079] To further illustrate the improvement effect of the small target detection model of the present embodiment, the accuracy rate and the recall rate indicators of the benchmark model yolo11n network model and the improved small target detection model of the present embodiment are tested respectively, and the indicators of the two models are compared as shown in Table 1.

[0080] Table 1

[0081] As can be seen from Table 1, the improved small target detection model of the present embodiment has improved accuracy rate and recall rate.

[0082] The present application also provides a track small target detection device. Please refer to Figure 6 , Figure 6 is a structural schematic diagram of the track small target detection device provided by the present application. In the present embodiment, the track small target detection device comprises an acquisition module 610 and a small target detection module 620.

[0083] The acquisition module 610 is used for acquiring a to-be-detected image.

[0084] The to-be-detected image is obtained after photographing a track area containing small targets, and the small targets include pedestrians and / or falling rocks.

[0085] The small target detection module 620 is configured to input the to-be-detected image into a pre-trained small target detection model to obtain a small target detection result output by the small target detection model.

[0086] The small target detection model is obtained by pre-training an initial model based on an improved yolo11 algorithm based on a sample detection image and a sample small target detection result corresponding to the sample detection image. The small target detection model is configured to perform feature extraction on the to-be-detected image to generate image features, perform feature alignment and feature fusion on the image features to generate fused features, and perform small target detection based on the fused features.

[0087] In some embodiments, the small target detection model comprises a backbone network, a neck network and a head network connected in sequence; the backbone network is configured to perform feature extraction on the to-be-detected image to generate image features; the neck network is configured to perform feature alignment and feature fusion on the image features to generate fused features; and the head network is configured to perform small target detection based on the fused features to generate a small target detection result.

[0088] In some embodiments, the backbone network comprises a first network layer, a second network layer, a third network layer, a fourth network layer and a fifth network layer connected in sequence, and the image features comprise a first image feature, a second image feature, a third image feature and a fourth image feature; the first network layer is configured to perform feature extraction on the to-be-detected image to generate an initial image feature; the second network layer is configured to perform feature extraction on the initial image feature to generate the first image feature; the third network layer is configured to perform feature extraction on the first image feature to generate the second image feature; the fourth network layer is configured to perform feature extraction on the second image feature to generate the third image feature; and the fifth network layer is configured to perform feature extraction on the third image feature to generate the fourth image feature.

[0089] In some embodiments, the neck network comprises an MLP module, a feature alignment network and a feature fusion network connected in sequence, the feature alignment network and the feature fusion network are deformable convolution networks, and the fusion features comprise first fusion features, second fusion features, third fusion features and fourth fusion features; wherein the MLP module is configured to perform channel screening processing on the first image features, the second image features, the third image features and the fourth image features respectively to generate first feature maps, second feature maps, third feature maps and fourth feature maps; the feature alignment network is configured to perform feature alignment on the first feature maps, the second feature maps, the third feature maps and the fourth feature maps to obtain first to-be-fused feature maps, second to-be-fused feature maps, third to-be-fused feature maps and fourth to-be-fused feature maps; and the feature fusion network is configured to perform feature fusion based on the first to-be-fused feature maps, the second to-be-fused feature maps, the third to-be-fused feature maps and the fourth to-be-fused feature maps to generate the first fusion features, the second fusion features, the third fusion features and the fourth fusion features respectively.

[0090] In some embodiments, the head network comprises a first prediction head, a second prediction head, a third prediction head and a fourth prediction head; wherein the second prediction head, the third prediction head and the fourth prediction head are configured to perform target detection based on the second fusion features, the third fusion features and the fourth fusion features to generate track target detection results; and the first prediction head is configured to perform small target detection based on the first fusion features to generate small target detection results; the small target detection results comprise small target position information and small target category information.

[0091] In some embodiments, the small target detection model is pre-trained based on a preset loss function, and the preset loss function is determined based on a positioning loss, a classification loss, a confidence loss and a self-supervised contrast loss.

[0092] In some embodiments, the small target detection module 620 is further configured to construct a data set; the data set comprises a plurality of sample detection images and sample small target detection results corresponding to each sample detection image collected in a test line scene and an actual running line scene; the data set is divided into a training set, a test set and a validation set; an initial model constructed based on the improved yolo11 algorithm is pre-trained based on the training set and the validation set to obtain a small target detection model; and the performance of the small target detection model is tested based on the test set.

[0093] The application further provides an electronic device. Figure 7 is a structural schematic diagram of the electronic device provided by the application, as Figure 7As shown, the electronic device can include a processor 710, a communications interface 720, a memory 730, and a communications bus 740, wherein the processor 710, the communications interface 720, and the memory 730 complete communications with each other through the communications bus 740. The processor 710 can invoke a logical instruction in the memory 730 to execute the track small target detection method.

[0094] In addition, the logical instruction in the memory 730 described above can be implemented in the form of a software function unit and sold or used as an independent product, and can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for making a computer device (which can be a personal computer, a server, or a network device, etc.) execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various program code storage media.

[0095] The present application also provides a non-transitory computer readable storage medium having a computer program stored thereon, wherein the computer program is executed by a processor to implement the track small target detection method provided by the above-mentioned methods.

[0096] The device embodiments described above are only schematic, wherein the units shown as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, i.e. they can be located in one place, or distributed on multiple network units. Part or all of the modules can be selected to achieve the purpose of the present embodiment according to actual needs. Those skilled in the art can understand and implement it without creative labor.

[0097] Those skilled in the art can clearly understand the technical solutions of the various embodiments from the above description of the embodiments, and the various embodiments can be implemented by means of software with the necessary general hardware platforms, and of course, can also be implemented by hardware. Based on such understanding, the above technical solutions, essentially or in other words, the part of the prior art that makes a contribution, can be embodied in the form of a software product, which can be stored in a computer readable storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, and the like, and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0098] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, rather than limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for some technical features therein; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for detecting small track targets, characterized in that: include: Acquire an image to be detected; the image to be detected is obtained by photographing a track area containing small targets, wherein the small targets include pedestrians and / or fallen rocks; Inputting the image to be detected into a pre-trained small target detection model to obtain a small target detection result output by the small target detection model; The small target detection model is obtained by pre-training an initial model constructed based on the improved YOLO11 algorithm based on the sample detection image and the sample small target detection results corresponding to the sample detection image; The small target detection model is used to extract features from the image to be detected, generate image features, perform feature alignment and feature fusion on the image features, generate fusion features, and perform small target detection based on the fusion features.

2. The method for detecting small track targets according to claim 1, wherein: The small target detection model includes a backbone network, a neck network and a head network connected in sequence; The backbone network is used to extract features from the image to be detected and generate the image features; The neck network is used to perform feature alignment and feature fusion on the image features to generate the fused features; The head network is used to perform small target detection based on the fusion features and generate the small target detection result.

3. The method for detecting small track targets according to claim 2, wherein: The backbone network includes a first network layer, a second network layer, a third network layer, a fourth network layer and a fifth network layer connected in sequence, and the image features include a first image feature, a second image feature, a third image feature and a fourth image feature; The first network layer is used to extract features from the image to be detected and generate initial image features; The second network layer is used to extract features from the initial image features to generate the first image features; The third network layer is used to extract the first image features to generate the second image features; The fourth network layer is used to extract the second image features to generate the third image features; The fifth network layer is used to perform feature extraction on the third image feature to generate the fourth image feature.

4. The method for detecting small track targets according to claim 3, wherein: The neck network includes an MLP module, a feature alignment network, and a feature fusion network connected in sequence, wherein the feature alignment network and the feature fusion network are deformable convolutional networks, and the fusion features include a first fusion feature, a second fusion feature, a third fusion feature, and a fourth fusion feature; The MLP module is configured to perform channel screening processing on the first image feature, the second image feature, the third image feature, and the fourth image feature, respectively, to generate a first feature map, a second feature map, a third feature map, and a fourth feature map; The feature alignment network is used to perform feature alignment on the first feature map, the second feature map, the third feature map, and the fourth feature map to obtain a first feature map to be fused, a second feature map to be fused, a third feature map to be fused, and a fourth feature map to be fused; The feature fusion network is used to perform feature fusion based on the first feature map to be fused, the second feature map to be fused, the third feature map to be fused, and the fourth feature map to be fused, to generate the first fused feature, the second fused feature, the third fused feature, and the fourth fused feature, respectively.

5. The method for detecting small track targets according to claim 4, wherein: The head network includes a first prediction head, a second prediction head, a third prediction head and a fourth prediction head; The second prediction head, the third prediction head, and the fourth prediction head are used to perform target detection based on the second fusion feature, the third fusion feature, and the fourth fusion feature to generate a track target detection result; The first prediction head is used to perform small target detection based on the first fusion feature to generate the small target detection result; the small target detection result includes small target position information and small target category information.

6. The method for detecting small track targets according to claim 1, wherein: The small target detection model is pre-trained based on a preset loss function, which is determined based on positioning loss, classification loss, confidence loss and self-supervised contrast loss.

7. The method for detecting small track targets according to claim 1, wherein: Before inputting the image to be detected into a pre-trained small target detection model and obtaining the small target detection result output by the small target detection model, the method further includes: Constructing a data set; the data set includes a plurality of sample detection images collected in the test line scene and the actual operation line scene and a sample small target detection result corresponding to each sample detection image; Dividing the dataset into a training set, a test set, and a validation set; Based on the training set and the validation set, pre-training the initial model constructed based on the improved YOLO11 algorithm to obtain the small target detection model; Based on the test set, a performance index test is performed on the small target detection model.

8. A small track target detection device, characterized in that: include: An acquisition module is configured to acquire an image to be detected; the image to be detected is obtained by photographing a track area containing small targets, such as pedestrians and / or fallen rocks; A small target detection module is used to input the image to be detected into a pre-trained small target detection model to obtain a small target detection result output by the small target detection model; The small target detection model is obtained by pre-training an initial model constructed based on the improved YOLO11 algorithm based on the sample detection image and the sample small target detection results corresponding to the sample detection image; The small target detection model is used to extract features from the image to be detected, generate image features, perform feature alignment and feature fusion on the image features, generate fusion features, and perform small target detection based on the fusion features.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the method for detecting small track targets as described in any one of claims 1 to 6 is implemented.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for detecting small track targets as described in any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Target detection method and system based on self-supervised contrast learning

    CN114549985A

  • Remote sensing image rotating small target detection method fusing cognitive features

    CN119229278A

  • Unmanned aerial vehicle aerial image small target detection method, system and device and medium

    CN119723374A

  • Unmanned aerial vehicle aerial photography small target detection method based on improved YOLOv11s algorithm

    CN120014244A