Improved YOLOv11-based laser detection method for target detection and identification

By improving the neck network of the YOLOv11 model and introducing a multi-scale attention module and a channel attention mechanism, the detection accuracy and speed issues of YOLO in small target detection and complex scenes are solved, and efficient laser detection is achieved.

CN120765951APending Publication Date: 2025-10-10西安中科立德红外科技有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510696477.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-10-10

AI Technical Summary

Technical Problem

The existing YOLO target detection method has poor accuracy in detecting small targets and low detection performance in complex scenes, which is prone to false detection and missed detection.

Method used

An improved YOLOv11 model is adopted. By building a self-built laser image dataset and introducing an improved multi-scale attention module (MSCA) in the neck network, combined with the channel attention mechanism (SE module) to enhance the feature extraction capability, the improved model includes a backbone network module, a neck network module and a detection head.

Benefits of technology

It improves the accuracy and speed of small target detection, adapts to detection capabilities in complex scenarios, reduces the computational burden, and is applicable to a wider range of mission scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120765951A_ABST
    Figure CN120765951A_ABST
Patent Text Reader

Abstract

The invention provides a laser detection method for target detection and identification based on improved YOLOv11, and the method comprises the following steps: 1, building a laser image data set, and carrying out the processing of the collected laser data set; 2, inputting the laser image data set obtained in the step 1 into an improved YOLOv11 model to obtain a laser feature image; the laser detection method for target detection and identification based on the improved YOLOv11 has the advantages of high detection rate and high detection speed. The problem of low performance in small target detection can be solved, an improved MSCA attention mechanism is added in a neck network, context information can be effectively extracted, richer details can be captured after improvement, calculation burden is reduced, and the method has higher expressive force in small target detection and is suitable for wider task scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of target detection, and in particular relates to a laser detection method for target detection and recognition based on improved YOLOv11. Background Art

[0002] Laser detection, originating from the widespread use of laser technology on the battlefield, is a countermeasure technology designed to target laser guidance and laser targeting systems, with the goal of protecting targets from enemy attack. Early battlefield countermeasures against enemy lasers focused on physical protection, such as applying laser-absorbing coatings or laser-reflective materials to the surface of the protected object to offset the laser's effects. However, traditional physical protection gradually became insufficient to meet battlefield demands. As the threat of lasers grew in the mid-20th century, countermeasures shifted from passive defense to active interference. In addition to transmitting laser signals to disrupt the enemy's laser targeting system, electronic jamming or signal shielding techniques can be used to disrupt its aiming or tracking functions, or camouflage nets and smoke bombs can be used to obscure the target. On the modern battlefield, laser detection methods have evolved to enable real-time monitoring of laser weapon use, integrating sensors, and rapidly analyzing laser signals through data analysis, enabling countermeasures to be initiated before a threat materializes. Currently, common laser detection systems utilize photoelectric detectors to receive laser signals, convert the optical signals into electrical signals, and then analyze various aspects of the incident laser light. This type of system can effectively reflect information such as laser wavelength and light source position, but due to detector and cost limitations, it cannot be deployed on a large scale on the battlefield.

[0003] With the development of artificial intelligence (AI), deep learning-based object detection methods have become widely used. Their principle involves constructing a dataset containing the objects to be detected, extracting features from the image using a neural network, generating candidate regions to classify the objects, and predicting bounding box coordinates. The detection results are then post-processed. Deep learning methods have seen rapid development in recent years due to their high accuracy, wide application scenarios, and real-time performance. Commonly used object detection methods are categorized as single-stage and two-stage. YOLO is a classic single-stage method. Its principle is to transform the object detection task into a single regression problem. The input image is divided into an S×S grid, with each grid responsible for predicting multiple bounding boxes and their corresponding confidence categories. The model processes the entire image at once, significantly improving detection speed. However, YOLO's accuracy is relatively poor when detecting small objects, and detection is also poor in complex scenes, potentially leading to false detections and missed detections. These issues result in YOLO's poor robustness in complex scenes, leaving room for improvement. Summary of the Invention

[0004] In view of the fact that the existing YOLO target detection method has relatively poor accuracy when detecting small targets and poor detection in complex scenes, which may result in false detection and missed detection, the present invention provides a laser detection method for target detection and recognition based on improved YOLOv11, comprising the following steps:

[0005] Step 1: Build your own laser image dataset and process the collected laser dataset;

[0006] Step 2: Input the laser image data set obtained in step 1 into the improved YOLOv11 model to obtain the laser feature image.

[0007] Furthermore, the image data collected in step 1: self-building the laser image data set is image data containing laser features by collecting visible light bands.

[0008] Furthermore, the improved YOLOv11 model includes a backbone network module, a neck network module, and a detection head, and an improved multi-scale attention module (MSCA) is provided in the neck network module.

[0009] Furthermore, the neck network module includes a first upsampling module, a first connection module, a first C3K2 module, a second upsampling module, a second connection module, a second C3K2 module, a first improved multi-scale attention module, a first convolution module, a third connection module, a third C3K2 module, a second improved multi-scale attention module, a second convolution module, a fourth connection module, a fourth C3K2 module, and a third improved multi-scale attention module, which are connected in sequence; the output end of the first C3K2 module is connected to the input end of the third C3K2 module, the input end of the fourth connection module is connected to the first input end of the backbone network module, the input end of the first upsampling module is connected to the first input end of the backbone network module, the input end of the first connection module is connected to the second input end of the backbone network module, the input end of the second connection module is connected to the third input end of the backbone network module, the output of the first improved multi-scale attention module is the first output end of the backbone network module, the output of the second improved multi-scale attention module is the second output end of the backbone network module, and the output of the third improved multi-scale attention module is the third output end of the backbone network module.

[0010] Furthermore, the first improved multi-scale attention module, the second improved multi-scale attention module, and the third improved multi-scale attention module have the same structure, and their calculation formulas are as follows:

[0011]

[0012] Where F represents the input feature, Att and Out are the attention map and output respectively, It is an element matrix multiplication operation, DW-Conv means depth convolution, Scale i , i∈{0, 1, 2, 3} represents the i-th branch, Scale0 is the identification connection;

[0013]

[0014] z c is the cth element statistic, u C is the cth element feature map, H represents the height of the image, and W represents the width of the image.

[0015] The advantages of the present invention are: It provides a laser detection method for target detection and recognition based on an improved YOLOv11, which has the advantages of high detection rate and fast detection speed. It can solve the problem of low performance in small target detection. The improved MSCA attention mechanism is added to the neck network, which can effectively extract contextual information. The improvement can also capture richer details while reducing the computational burden. It has stronger expressiveness in small target detection and is suitable for a wider range of task scenarios.

[0016] The present invention is described in detail below with reference to the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 Schematic diagram of laser detection using the laser detection method based on improved YOLOv11 target detection and recognition.

[0018] Figure 2 Flowchart of the laser detection method for target detection and recognition based on improved YOLOv11.

[0019] Figure 3 This is the framework diagram of the C3k2 attention module.

[0020] Figure 4 This is a framework diagram of the improved MSCA of the present invention. DETAILED DESCRIPTION

[0021] In order to further illustrate the technical means and effects adopted by the present invention to achieve the predetermined purpose, the specific implementation methods, structural features and effects of the present invention are described in detail below with reference to the accompanying drawings and examples.

[0022] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0023] All features disclosed in this specification, or all steps in the disclosed methods or processes, except mutually exclusive features and / or steps, can be combined in any manner.

[0024] Any feature disclosed in this specification (including any appended claims, abstract and drawings), unless otherwise stated, may be replaced by other equivalent or similar features. That is, unless otherwise stated, each feature is only an example of a series of equivalent or similar features.

[0025] Example 1

[0026] The existing YOLO target detection method has relatively poor accuracy when detecting small targets; the detection in complex scenes is also poor, which may result in false detection and missed detection.

[0027] This embodiment provides a Figures 2 to 4 The laser detection method for target detection and recognition based on improved YOLOv11 includes the following steps:

[0028] Step 1: Build your own laser image dataset and process the collected laser dataset;

[0029] Step 2: Input the laser image data set obtained in step 1 into the improved YOLOv11 model to obtain the laser feature image.

[0030] Furthermore, the image data collected in step 1: self-building the laser image data set is image data containing laser features by collecting visible light bands.

[0031] Given the limited widespread use of current target detection methods for laser detection, this example uses a self-built laser image dataset to train the YOLOv11 model. This step primarily involves collecting image data in the visible light band that contains laser characteristics. During the laser image data collection process, various battlefield environments are simulated and different background conditions are set to ensure the dataset's diversity and representativeness, thereby improving the model's detection accuracy in various complex environments.

[0032] Furthermore, the specific process for processing the collected laser dataset in step 1 is as follows: first, the image data is screened to remove images where the laser spot is not captured or where the spot is not obvious. After the screening is completed, image processing operations such as unified cropping, transformation, and resolution adjustment are performed. Some images are selected for rotation and noise addition to improve the model's generalization ability. The collected laser image data is then annotated using lamelling software to mark the laser spot position in the image, clarifying the characteristics and position of the laser under different environmental conditions, completing the process of building a self-built laser image dataset.

[0033] Furthermore, the improved YOLOv11 model includes a backbone network module, a neck network module, and a detection head. An improved multi-scale attention module (MSCA) is provided in the neck network module.

[0034] YOLOv11 is a newer model version in the YOLO series, developed by Ultralytics. It inherits YOLOv8 and is improved and optimized based on it. It adopts an improved backbone and neck architecture, such as Figure 2 . In the backbone network, compared with YOLOv8, the C2f attention module is replaced with the C3k2 attention module, which is the core improvement of YOLOv11. The input of the backbone network module of YOLOv11 is the original image data. Specifically, the input is a tensor with a shape of (B, C, H, W), where: B represents the batch size. C represents the number of channels of the image, usually 3 (RGB image). H and W represent the height and width of the image, respectively, usually 640x640 or other fixed sizes. The output of the backbone network module of YOLOv11 is multi-scale feature maps, which will be passed to the neck part of the network for further processing. Figure 3 C3k2 adds an optional parameter, c3k, which switches between c3k and bottleneck, providing flexibility to meet different task requirements. Two additional convolution operations enhance local feature extraction, improving feature resolution and representation in complex scenarios. In the backbone network, YOLOv11 adds C2PSA after the SPPF layer. This is an extension of C2f, incorporating a PSA module. This mechanism aims to extract global features using a multi-head attention mechanism to enhance feature extraction capabilities, but its performance for small object detection is less than satisfactory. The backbone network module of YOLOv11 utilizes a PAN architecture and the C3K2 module. The output is multi-scale feature maps after feature fusion and enhancement. These feature maps are passed to the network's detection head for final object detection. The head employs an anchor-free + decoupled-head architecture, with the regression head using normal convolution and the classification head using depthwise separable convolution (DWConv) to reduce computational overhead.

[0035] The improved YOLOv11 of the technical solution of this embodiment has the same backbone network module and detection head as YOLOv11.

[0036] Furthermore, the neck network module includes a first upsampling module, a first connection module, a first C3K2 module, a second upsampling module, a second connection module, a second C3K2 module, a first improved multi-scale attention module, a first convolution module, a third connection module, a third C3K2 module, a second improved multi-scale attention module, a second convolution module, a fourth connection module, a fourth C3K2 module, and a third improved multi-scale attention module, which are connected in sequence; the output end of the first C3K2 module is connected to the input end of the third C3K2 module, the input end of the fourth connection module is connected to the first input end of the backbone network module, the input end of the first upsampling module is connected to the first input end of the backbone network module, the input end of the first connection module is connected to the second input end of the backbone network module, the input end of the second connection module is connected to the third input end of the backbone network module, the output of the first improved multi-scale attention module is the first output end of the backbone network module, the output of the second improved multi-scale attention module is the second output end of the backbone network module, and the output of the third improved multi-scale attention module is the third output end of the backbone network module.

[0037] Furthermore, the first improved multi-scale attention module, the second improved multi-scale attention module, and the third improved multi-scale attention module have the same structure.

[0038] To address the low performance of YOLOv11 in small object detection, this embodiment proposes an improved Multi-Scale Convolutional Attention (MSCA) module. The existing MSCA module consists of three modules: deep convolution to aggregate local information, multi-branch deep strip convolution to capture multi-scale context, and 1×1 convolution to model the relationship between different channels.

[0039] The calculation formula is as follows:

[0040]

[0041] Among them, F represents the input feature, Att and Out are attention mapping and output respectively, It is an element matrix multiplication operation, DW-Conv means depth convolution, Scale i , i∈{0, 1, 2, 3} represents the i-th branch, and Scale0 is the identifier connection. By adding the MSCA module to the neck network of YOLOv11, contextual information can be effectively extracted. However, the attention calculation of multiple scales is fused through additive calculation, which does not consider the proportional relationship between the attention scales, which may lead to information loss or imbalance.

[0042] On the basis of the existing MSCA, Figure 4This patent introduces a channel attention mechanism (SE module). The channel attention mechanism (SE module, Squeeze-and-Excitation Module) is an attention mechanism used to enhance the feature representation capabilities of convolutional neural networks (CNNs). It explicitly models the dependencies between channels, enabling the network to adaptively enhance important features and suppress redundant information, thereby improving model performance. Through adaptive average pooling, a global description of each channel is obtained, and learning is performed through two 1×1 convolutions and ReLU activations. Finally, the weight of each channel is obtained through Simoid activation.

[0043]

[0044] z c is the cth element statistic, u C is the cth element feature map, where H represents the image height and W represents the image width. This improves the model's ability to adjust across different channels, enabling it to capture richer details. It also adds learnable weighted fusion parameters to dynamically adjust the weights of features at different scales, enabling the model to adaptively select the appropriate scale combination based on the specific input. The improved MSCA module avoids the redundant computations associated with directly stacking all scales, while reducing the computational burden. This provides stronger performance in small object detection and is suitable for a wider range of tasks.

[0045] In summary, the laser detection method for target detection and recognition based on improved YOLOv11 provided in this embodiment uses a YOLO-based target detection and recognition method for laser detection. This method has low deployment cost, large coverage, fast response time, is suitable for various complex background conditions, and is more suitable for battlefield environments; it has the advantages of high detection rate and fast detection speed; compared with the target detection and recognition algorithm based on deep learning, this method uses the latest YOLOv11 as the basis, which has a high detection rate and fast detection speed compared to previous versions. In order to solve the problem of low performance of YOLOv11 in small target detection, the improved MSCA attention mechanism is added to the neck network, which can effectively extract contextual information. At the same time, the improvement can capture richer details, while reducing the computational burden, and has stronger expressiveness in small target detection, suitable for a wider range of task scenarios.

[0046] Example 2

[0047] like Figure 1As shown, the laser detection method for target detection and recognition based on the improved YOLOv11 shown in Example 1 is used to perform laser detection on the established laser image dataset, and then the obtained laser feature image is input into the YOLO module, which is loaded into the detection module by the YOLO module for detection and recognition. When a laser is detected, the system pops up an alarm pop-up signal and transmits the alarm information back to the optical display module for observation by the user to remind the user to take timely defense.

[0048] This laser detection method, based on improved YOLOv11 object detection and recognition, is designed to be deployed on AR smart helmets. The helmet's image acquisition module transmits battlefield information to an intelligent computing unit, which is equipped with the laser detection system. The image acquisition module transmits information containing laser images to the intelligent computing unit for detection and recognition by YOLO. Upon detecting laser light, the system displays a pop-up alert and transmits the alert information back to the user's optical display module, prompting them to take timely defensive measures. If the user deems the threat low or resolved, they can choose to ignore the alert, and the pop-up window will automatically disappear after 2 seconds. However, the system will record the alert for future reference. The alert signal can be transmitted to the user's voice module to provide voice prompts to the entire team, or protective equipment can be connected. Upon receiving the alert signal, appropriate protective measures are automatically activated. After multiple rounds of training, the YOLO model is deployed in a detection program. The running detection program can effectively detect and identify laser light in various complex environments, and promptly display an alert window to provide users with early warnings.

[0049] This method uses a pop-up alert mechanism to alert the user when laser light is detected, prompting them to immediately take defensive or counter-attack measures. Other alert methods can also be supported, or appropriate protective measures can be automatically activated.

[0050] The above is a further detailed description of the present invention in conjunction with specific preferred embodiments, and the specific implementation of the present invention should not be considered to be limited to these descriptions. For those skilled in the art of the present invention, without departing from the concept of the present invention, several simple deductions or substitutions can be made, which should be considered to fall within the scope of protection of the present invention.

Claims

1. A laser detection method for target detection and recognition based on improved YOLOv11, characterized in that: The steps include: Step 1: Build your own laser image dataset and process the collected laser dataset; Step 2: Input the laser image data set obtained in step 1 into the improved YOLOv11 model to obtain the laser feature image.

2. The laser detection method for target detection and recognition based on improved YOLOv11 according to claim 1, characterized in that: The image data collected in step 1: self-building the laser image data set is image data containing laser features by collecting visible light bands.

3. The laser detection method for target detection and recognition based on improved YOLOv11 according to claim 1, characterized in that: The improved YOLOv11 model includes a backbone network module, a neck network module, and a detection head. The neck network module is provided with an improved multi-scale attention module (MSCA).

4. The laser detection method for target detection and recognition based on improved YOLOv11 according to claim 3, characterized in that: The neck network module includes a first upsampling module, a first connection module, a first C3K2 module, a second upsampling module, a second connection module, a second C3K2 module, a first improved multi-scale attention module, a first convolution module, a third connection module, a third C3K2 module, a second improved multi-scale attention module, a second convolution module, a fourth connection module, a fourth C3K2 module, and a third improved multi-scale attention module connected in sequence; The output end of the first C3K2 module is connected to the input end of the third C3K2 module, the input end of the fourth connection module is connected to the first input end of the backbone network module, the input end of the first upsampling module is connected to the first input end of the backbone network module, the input end of the first connection module is connected to the second input end of the backbone network module, the input end of the second connection module is connected to the third input end of the backbone network module, the output of the first improved multi-scale attention module is the first output end of the backbone network module, the output of the second improved multi-scale attention module is the second output end of the backbone network module, and the output of the third improved multi-scale attention module is the third output end of the backbone network module.

5. The laser detection method for target detection and recognition based on improved YOLOv11 according to claim 4, characterized in that: The first improved multi-scale attention module, the second improved multi-scale attention module, and the third improved multi-scale attention module have the same structure, and their calculation formulas are as follows: Where F represents the input feature, Att and Out are the attention map and output respectively, It is an element matrix multiplication operation, DW-Conv means depth convolution, Scale i , i∈{0, 1, 2, 3} represents the i-th branch, Scale0 is the identification connection; z c is the cth element statistic, u C is the cth element feature map, H represents the height of the image, and W represents the width of the image.