A laser jamming scene recognition method based on deep learning
By combining the improved ResNet50 model with deep learning methods such as CBAM and external attention mechanism, the problem of laser interference recognition in complex scenes of optoelectronic imaging systems is solved, and high-precision laser interference recognition and suppression decisions are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- NAT UNIV OF DEFENSE TECH
- Filing Date
- 2025-05-14
- Publication Date
- 2026-05-12
AI Technical Summary
In existing technologies, photoelectric imaging systems have weak model generalization ability and poor scene adaptability when processing high-dimensional laser interference features, making it difficult to effectively identify laser interference features in complex scenes.
A deep learning-based laser interference scene recognition method is adopted. By constructing an improved ResNet50 model, combining the CBAM attention mechanism and the external attention mechanism, and using DCNv2 for efficient convolution operations, the method is trained on laser interference scene images collected by the optoelectronic imaging system to identify whether the optoelectronic imaging system is interfered with by laser.
It achieves high-precision laser interference identification in complex scenarios, reduces dependence on optical hardware parameters, provides accurate decision-making basis for laser interference suppression, and ensures optimized allocation of system resources.
Smart Images

Figure CN120495842B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of laser interference recognition, and in particular to a laser interference scene recognition method based on deep learning. Background Technology
[0002] As a core component of modern detection systems, optoelectronic imaging systems are widely used in fields such as medical imaging, video media, security management, high-resolution target reconnaissance and identification, optoelectronic precision guidance, fire control and targeting, flight assistance, and autonomous driving. However, direct light from natural light sources such as the sun, artificial light sources such as neon lights and streetlights, and reflected light from objects can all cause serious optoelectronic interference to the image information generated by optoelectronic imaging systems.
[0003] Currently, scholars have conducted research on methods for identifying and eliminating incoherent light (sunlight or lamplight) interference scenes, mainly using traditional image processing methods such as Gaussian filtering and histogram equalization. However, these methods rely excessively on optical hardware parameters, resulting in feature representation capabilities that are limited by prior knowledge. They are difficult to capture laser interference features in complex scenes without prior knowledge, and suffer from systemic defects such as weak model generalization ability and poor scene adaptability, especially when dealing with high-dimensional interference features. Summary of the Invention
[0004] The purpose of this application is to provide a laser interference scene recognition method based on deep learning, which can intelligently, effectively and accurately identify laser interference scenes.
[0005] To achieve the above objectives, this application provides the following solution:
[0006] This application provides a deep learning-based method for laser interference scene recognition, including:
[0007] Acquire images of laser interference scenes captured by the photoelectric imaging system;
[0008] The laser interference scene image is input into a preset laser interference scene recognition model to obtain the corresponding laser interference state; the laser interference state is either that the photoelectric imaging system is interfered with by laser or that the photoelectric imaging system is not interfered with by laser.
[0009] The preset laser interference scene recognition model is the optimal model obtained by training and comparing different deep learning networks using a laser interference scene sample set.
[0010] According to the specific embodiments provided in this application, this application has the following technical effects: This application adopts a preset laser interference scene recognition model, which is obtained by training and comparing different deep learning networks using a laser interference scene sample set. By applying deep learning networks to laser interference scene recognition, it fills a technical gap in this field and effectively addresses the dependence on optical hardware parameters in existing technologies. Furthermore, the deep learning-based intelligent recognition method used in this application for laser interference scene recognition provides accurate decision-making basis for subsequent laser interference suppression algorithms, thereby avoiding ineffective suppression operations and ensuring optimized allocation of system resources in critical scenarios. Attached Figure Description
[0011] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0012] Figure 1 This is a flowchart illustrating a laser interference scene recognition method based on deep learning, provided as an embodiment of this application.
[0013] Figure 2 This is a schematic diagram illustrating the process of determining a preset laser interference scene recognition model.
[0014] Figure 3 This is a schematic diagram illustrating the interference of lasers of different power on target objects.
[0015] Figure 4 This is a schematic diagram of the ResNet50 structure.
[0016] Figure 5 This is a schematic diagram of a laser interference scene recognition model.
[0017] Figure 6 This is another structural diagram of a laser interference scene recognition model.
[0018] Figure 7 This is a schematic diagram of a DCNv2 convolutional subunit.
[0019] Figure 8 This is a schematic diagram of the CBAM attention subunit.
[0020] Figure 9 This is a schematic diagram of the external attention subunit.
[0021] Figure 10 This represents the loss-accuracy curves of the deep learning model on the training and validation sets. Detailed Implementation
[0022] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0023] This application can identify laser interference scenarios and effectively determine whether the scene content of the optical imaging system is affected by laser interference, so as to facilitate the system to implement different suppression measures for laser interference scenarios and has wide application adaptability.
[0024] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0025] In one exemplary embodiment, such as Figure 1 As shown, a laser interference scene recognition method based on deep learning is provided, including the following steps 201 to 202.
[0026] Step 201: Obtain the laser interference scene image collected by the photoelectric imaging system.
[0027] Step 202: Input the laser interference scene image into a preset laser interference scene recognition model to obtain the corresponding laser interference state; the laser interference state is that the photoelectric imaging system is interfered with by the laser or the photoelectric imaging system is not interfered with by the laser; wherein, the preset laser interference scene recognition model is the optimal model obtained by training and comparing different deep learning networks using a laser interference scene sample set.
[0028] Before performing steps 201 and 202 above, it is first necessary to determine a preset laser interference scene recognition model, such as... Figure 2 As shown, in an application example, the process of determining the preset laser interference scene recognition model includes the following three parts.
[0029] (a) Creating a laser interference image dataset.
[0030] The laser interference image dataset is a laser interference scene sample set. The data in this dataset mainly consists of scenes where target objects (such as cars) are interfered with by lasers. The steps include the following (11)-(14).
[0031] (11) Multiple laser interference sample images were acquired using an optoelectronic imaging system; different laser interference sample images were subjected to laser interference of different power.
[0032] In a specific practical application, a photoelectric imaging system, a laser, a computer, a target vehicle, and a Soleb beam expander can be used to acquire sample images. The specific parameters of the above instruments and equipment are shown in Table 1 below. Among them, the photoelectric imaging system includes two beam splitters, four 25mm focal length lenses, three apertures, and three image plane detectors.
[0033] Based on the aforementioned instruments and equipment, the sample image acquisition process includes: aligning the optical imaging system with the target object (i.e., the target vehicle), with the distance between the target vehicle and the image detector of the photoelectric imaging system approximately 3 meters; the laser beam emitted from the laser being directed to the target vehicle via a Sorebo beam expander; fixing the laser position relative to the image detector; and acquiring images of laser interference at different power levels by adjusting the laser power, allowing the target scene light spot to continuously expand, such as... Figure 3 As shown, by using computer-controlled real-time image acquisition, 4800 laser interference images (corresponding to the laser interference state where the photoelectric imaging system is interfered with by laser) and 4800 laser-free images (corresponding to the laser interference state where the photoelectric imaging system is not interfered with by laser) can be acquired.
[0034] Table 1
[0035] Instruments and equipment Specific parameters Photoelectric imaging system Detector resolution: 1440*1080 laser Wavelength: 532nm computer Model: Legion Y7000PIRH8 Target car Size: 30cm*20cm*20cm Sorebo beam expander GBE10-A
[0036] (12) Determine the laser interference state corresponding to each laser interference sample image.
[0037] (13) Perform preprocessing operations on all the laser interference sample images; the preprocessing operations include standardization, normalization and data augmentation.
[0038] The laser interference sample images obtained through step (11) above have a resolution of 1440*1080. To simplify model training and evaluation in subsequent steps, the collected images are standardized by adjusting all images to a uniform size of 224×224 pixels and performing normalization to eliminate the influence of different resolutions and ensure data consistency. To improve the model's generalization ability and reduce overfitting, data augmentation techniques include random flipping, random cropping, and center cropping.
[0039] (14) Any laser interference sample image after preprocessing and its corresponding laser interference state constitute a sample; multiple samples constitute a laser interference scene sample set.
[0040] (II) Construction of laser interference scene recognition model.
[0041] Prior to model training, this application constructed several deep learning networks, including improved ResNet50, ResNet50, VGG16, and AlexNet.
[0042] Among them, ResNet50 effectively overcomes the gradient vanishing problem in deep learning classification models. The ResNet50 architecture is divided into four main parts, each called a stage, and each performs a different function. In the first stage, ResNet50 uses a 7×7 convolutional kernel combined with a 3×3 max pooling operation, which effectively reduces the size of the input image and provides a more compact data representation for subsequent processing. In the second stage, through a series of innovative residual structure designs, namely residual blocks such as Conv2, Conv3, Conv4, and Conv5, the model can deeply explore the high-level features of the image. Then, these highly abstract features are fed into the fully connected layer in the third stage to complete the final classification task.
[0043] ResNet primarily addresses the gradient problem through residual networks, using cross-layer connections. This structure allows the input to skip certain layers and be added to the processed output, forming the concept of a "residual." Assuming the input image is x and the output is H(x), the output after convolution is a non-linear function F(x), and the final output is H(x) = F(x) + x. This can be transformed into finding the residual function F(x) = H(x) - x, which is easier to optimize than F(x) = H(x). The structure of ResNet50 is as follows... Figure 4 As shown, the ResNet residual network introduces residual structures, stacking multiple residual units to construct networks of different depths and complexities to adapt to task requirements, achieving significant performance improvements in various image-related tasks.
[0044] VGG16 and AlexNet, among others, use existing architectures, which will not be elaborated upon here.
[0045] In one exemplary application, the preset laser interference scene recognition model is an improved ResNet50, such as... Figure 5 and Figure 6 As shown, it includes an input module, a convolutional pooling module, a ResNet residual module, and a pooling classification output module arranged in sequence.
[0046] (21) The input module is used to receive the laser interference scene image; specifically, a 224×224 pixel laser interference image is used as the input to the network.
[0047] (22) The convolutional pooling module includes a first convolutional layer (conv), a batch normalization layer, a ReLU activation function layer and a max pooling layer arranged in sequence; specifically, the batch normalization layer is used to speed up the convergence speed, the ReLU activation function is used to increase nonlinearity, and the max pooling layer is used for sampling.
[0048] (23) The ResNet residual module includes multiple residual sub-modules (ResNet50 Bottleneck), and the number of residual units contained in different residual sub-modules is different; the residual units include convolution processing, DCNv2 (Deformable Convolutional Networks version 2) convolution processing, CBAM attention mechanism and external attention mechanism; specifically, this application sets up four residual sub-modules, which contain 3, 4, 6 and 3 residual units respectively.
[0049] The residual unit includes a second convolutional layer, a DCNv2 convolutional sub-unit, a third convolutional layer, a CBAM attention sub-unit, and an external attention sub-unit arranged sequentially; the input end of the second convolutional layer serves as the input end of the residual unit; the input end of the second convolutional layer is connected to the output end of the external attention sub-unit through an addition operation, and the connection end serves as the output end of the residual unit.
[0050] Specifically, for each residual unit, the feature map input to the residual unit first passes through a 1×1 second convolutional layer to adjust the number of channels; then through a 3×3 DCNv2 convolutional sub-unit for multi-scale feature extraction and spatial transformation; next, it passes through a 1×1 third convolutional layer to further adjust the number of channels or fuse features; attention mechanisms in both channels and spatial dimensions are enhanced through a CBAM attention sub-unit; and the external attention mechanism of the external attention sub-unit captures potential correlations between different samples. Finally, all processed features are fused through an addition operation to form the final output feature map. The structure of the residual units in the second, third, and fourth residual sub-modules is the same as above.
[0051] In an exemplary embodiment, the DCNv2 convolutional subunit includes an offset generation unit and a sampling convolution unit; the offset generation unit is used to: for a received feature map, calculate the sampling point offset of the convolution kernel in the horizontal and vertical directions of the feature map through a preset convolution operation; the sampling convolution unit is used to: perform bilinear interpolation based on the sampling point offset and the feature map to determine the sampling points of the convolution kernel on the feature map; and perform a convolution operation on the feature map according to the sampling points of the convolution kernel.
[0052] Specifically, such as Figure 7As shown, DCNv2 is an improved version based on deformable convolutional networks. It introduces deformable convolutional layers into traditional convolutional neural networks to improve the network's adaptability to changes in the shape of objects in images. The offset generation unit accurately calculates the sampling point offsets of the convolution kernel in the horizontal (x-axis) and vertical (y-axis) directions on the received feature map through specific convolution operations. The sampling convolution unit uses the above offset information (i.e., sampling point offsets) and the input feature map to perform bilinear interpolation to accurately locate the sampling points of the convolution kernel, thereby completing efficient and flexible convolution processing. Finally, the convolution operation is performed using these sampling points.
[0053] In another exemplary embodiment, the CBAM attention subunit includes a channel attention unit and a spatial attention unit; the feature map received by the channel attention unit and the feature map output by the channel attention unit are multiplied and then input to the spatial attention unit; the feature map received by the spatial attention unit and the feature map output by the spatial attention unit are multiplied and then input to the external attention subunit.
[0054] like Figure 8 As shown, the channel attention unit includes a first global max pooling layer, a second global average pooling layer, a shared fully connected layer, a second global max pooling layer, a third global average pooling layer, and a first sigmoid function layer. The inputs of the first global max pooling layer and the second global average pooling layer are both connected to the output of the third convolutional layer. The outputs of the first global max pooling layer and the second global average pooling layer are both connected to the input of the shared fully connected layer. The output of the shared fully connected layer is connected to the inputs of the second global max pooling layer and the third global average pooling layer, respectively. After performing an addition operation, the outputs of the second global max pooling layer and the third global average pooling layer are connected to the first sigmoid function layer. The first sigmoid function layer is used to activate and output the feature map. The spatial attention unit includes, in sequence, an average pooling layer, a max pooling layer, a fourth convolutional layer, and a second sigmoid function layer.
[0055] Specifically, in the channel attention unit, for the received input feature map, global max pooling and global average pooling are first performed on each channel of the feature map to calculate the maximum and average eigenvalues for each channel, resulting in two vectors containing the number of channels, representing the global maximum and average features for each channel, respectively. These global maximum and average features are then input into a shared fully connected layer. The fully connected layer learns the attention weights for each channel in the feature map, and the output is then subjected to global max pooling and global average pooling again. The global maximum and average eigenvectors are then intersected to obtain the final attention weight vector. Finally, the output feature map is activated by the sigmoid function.
[0056] In the spatial attention layer, for the received input feature map, average pooling and max pooling operations are performed along the channel dimension. The features obtained after average pooling and max pooling are concatenated along the channel dimension, and then passed through a convolutional layer to obtain a feature map with different scale contextual information. Finally, the output feature map is activated by the sigmoid function.
[0057] In another exemplary embodiment, such as Figure 9 As shown, the ExternalAttention mechanism in the ExternalAttention subunit enhances network capabilities by using two external memory matrices, MemoryUnit, independent of the input feature map, as the key and value. Specifically, the feature map received by the ExternalAttention subunit is mapped to a Query matrix Q, a key matrix k (corresponding to one external memory matrix), and a value matrix v (corresponding to another external memory matrix) to model the similarity between the i-th pixel and the j-th row of the feature map. The external memory matrices MemoryUnit are learnable and variable in size. Furthermore, MemoryUnit can model the relationships between different samples in the entire dataset as training progresses, significantly reducing computational and memory costs. Here, Linear refers to a linear function, and Matrix multiplication refers to matrix multiplication.
[0058] (24) The pooling classification output module is used to output the laser interference status; the pooling classification output module includes a first global average pooling layer, a fully connected layer (FC), a softmax classification layer and an output layer set in sequence; specifically, the feature map is reduced to 1×1 by global average pooling, classified by a fully connected layer, and the final classification result is output by softmax.
[0059] (III) Model training and model index evaluation.
[0060] In an exemplary application instance, the process of determining the preset laser interference scene recognition model includes the following steps (31)-(36).
[0061] (31) Obtain a laser interference scene sample set; each sample in the laser interference scene sample set includes a laser interference sample image and the corresponding laser interference state; this step can directly call the data results obtained in (a) above. That is, obtain the laser interference sample image and the non-laser interference sample image obtained by the image plane detector (such as a camera) in the photoelectric imaging system, and label them with the corresponding laser interference state.
[0062] (32) The laser interference scene sample set is divided into a training set, a validation set and a test set; specifically, the 9,600 collected image data are divided into a training set, a validation set and a test set according to the category, with a division ratio of 8:1:1.
[0063] (33) The training set is used to train different deep learning networks to obtain multiple optimized deep learning models.
[0064] Specifically, the hardware environment used in this application consisted of an Intel Core i5 processor, 32GB of RAM, and an NVIDIA GeForce RTX 3090 24GB graphics card. The programming languages used were Python 3.9, PyTorch version 2.2.0, and CUDA version 12.1. During model training, the parameters of each model were adjusted to determine the optimal learning rate. All models used the SGD optimizer with a momentum of 0.9 and a learning rate of 0.01. The environment parameters used for model training and testing are shown in Table 2. All network models were trained for 100 epochs to ensure that the models could fully learn and optimize their performance.
[0065] Table 2
[0066] Experimental System Windows 10 CPU 3.00GHzIntel(R)Xeon(R)Gold6248R GPU NVIDIA GeForce RTX 3090 Memory 32GB Development Environment Python 3.9 Deep learning framework PyTorch 2.2.0
[0067] (34) Using the validation set, each of the optimized deep learning models is validated and adjusted to obtain multiple corresponding final deep learning models; such as Figure 10 As shown, the accuracy of the trained model on the validation set stabilizes at 45 epochs, and on the training set it stabilizes at 50 epochs. Simultaneously, the model converges rapidly on the loss curve. Based on these results, the proposed method effectively accelerates the model's convergence process, improves its stability and generalization ability on the validation set, and lays a solid foundation for subsequent model applications.
[0068] (35) Using the test set, each of the final deep learning models is tested to obtain multiple test results. Specifically, to measure the performance of the models, accuracy, recall, and precision are used for evaluation, and their calculation formulas are as follows:
[0069]
[0070] Where TP represents the number of samples correctly predicted as positive, TN represents the number of samples correctly predicted as negative, FP represents the number of samples incorrectly predicted as positive, and FN represents the number of samples incorrectly predicted as negative.
[0071] (36) Compare multiple test results and mark the best-performing final deep learning model as the preset laser interference scene recognition model.
[0072] Table 3 below shows the training results of three models—the improved ResNet50 (this application), ResNet50, VGG16, and AlexNet—under the same parameter settings, including accuracy, recall, and precision. It can be seen that the improved ResNet50 model performs best in accuracy, recall, and precision, with scores of 0.9913, 0.9889, and 0.9761, respectively. This indicates that the improved ResNet50 model has higher accuracy and stability in recognizing laser interference scenes. In contrast, the performance of the ResNet50, VGG16, and AlexNet models is relatively poor, but it still verifies the application potential of the improved deep learning model in the field of laser interference scene recognition to some extent.
[0073] Table 3
[0074]
[0075]
[0076] To verify the effectiveness and generalization of the improved model, under the same experimental conditions and parameter settings, the proposed improvement strategy was integrated into the ResNet50 model, and three sets were tested to compare the model's accuracy, recall, and precision, as shown in Table 4. √ indicates that the module was added to the model, and × indicates that the module was not added. The experimental results show that adding each module improved the model's metrics, demonstrating the effectiveness of the improved model. Comparatively, the improvement of the external attention mechanism (EA) had the greatest impact on the model's performance. Therefore, the improved ResNet50 in this application fully combines the powerful feature extraction capabilities of the CBAM attention mechanism and the external attention mechanism (EA), and utilizes DCNv2 for efficient convolution operations, achieving high-precision classification and recognition of laser interference scenes.
[0077] Table 4
[0078] Model CBAM EA DCNNv2 accuracy Recall rate accuracy 1 × × × 0.9408 0.9321 0.9414 2 √ × × 0.9515 0.9456 0.9448 3 √ √ × 0.9748 0.9669 0.9694 4 √ √ √ 0.9913 0.9889 0.9761
[0079] In practical applications, for laser interference scene images that are not accurately identified, the diversity of the dataset can be increased by collecting and labeling more images of the scene. Then, data augmentation techniques (such as image cropping and image rotation) can be used to simulate interference under different conditions. Next, these new data are integrated into the training set, and the model is retrained to enhance its adaptability and generalization ability to complex scenes. Finally, iterative optimization and continuous monitoring are used to ensure the robustness and accuracy of the model when facing new problems.
[0080] In summary, the deep learning-based laser interference scene recognition implemented in this application is an improved recognition model that uses the ResNet network model as the base network and combines it with CBAM attention mechanisms and external attention mechanisms. First, images of different laser interference scenes are collected, along with images of scenes without laser interference. The collected images are then used to train different classification and recognition algorithms, and the optimal algorithm is determined. The trained algorithm model is then used to verify its performance. For scenes with incorrect classification, supplementary datasets are added, and the model is retrained and optimized to increase the accuracy and generalization of the model's classification and recognition. A residual learning architecture is introduced into the above algorithm model to maintain the stability of deep network training while constructing a multi-scale feature fusion mechanism, effectively improving the representation ability of complex interference patterns.
[0081] This application proposes the feasibility of applying deep learning classification and recognition algorithms to the field of laser interference. Secondly, it uses deep learning classification algorithms to identify whether an optoelectronic imaging system is interfered with by lasers. Several common classification and recognition algorithms are compared, and the algorithms with better performance are improved. Finally, a laser interference suppression scene recognition method based on a deep learning model is proposed, which can effectively identify whether an optoelectronic imaging system is interfered with by lasers, thus preparing for subsequent laser interference suppression.
[0082] In addition, although there are other alternatives (such as improvements to traditional image processing methods, other deep learning networks such as other variants of convolutional neural networks (CNNs) such as VGG, GoogLeNet, etc., and combinations of traditional image processing methods and deep learning methods) that can be attempted for the recognition of laser interference scenes, the laser interference scene recognition method based on the improved ResNet50 proposed in this application has significant advantages in terms of feature extraction capability, deep network training stability, and adaptability to complex scenes. It can more effectively achieve high-precision classification and recognition of laser interference scenes, thereby better achieving the purpose of the invention.
[0083] Based on the same inventive concept, embodiments of this application also provide a laser interference scene recognition system based on deep learning, comprising:
[0084] The data acquisition subsystem is used to acquire images of laser interference scenes collected by the photoelectric imaging system.
[0085] A laser interference recognition subsystem is used to: input the laser interference scene image into a preset laser interference scene recognition model to obtain the corresponding laser interference state; the laser interference state is that the photoelectric imaging system is interfered with by laser or the photoelectric imaging system is not interfered with by laser; wherein, the preset laser interference scene recognition model is the optimal model obtained by training and comparing different deep learning networks using a laser interference scene sample set.
[0086] Compared with the prior art, this application also has the following advantages:
[0087] (1) This application does not rely on prior knowledge, but directly utilizes deep learning theory to extract features from images output by photoelectric imaging systems for laser interference scene recognition, filling a technological gap in this field and demonstrating significant innovation and practicality. Furthermore, the training network for deep learning includes, but is not limited to, residual networks.
[0088] (2) This application utilizes an improved ResNet to achieve high-precision classification and recognition of laser interference scenarios. This application innovatively adopts a ResNet as the core architecture and improves it for recognizing laser interference scenarios. The ResNet, combined with CABM and external attention mechanisms, possesses powerful feature extraction capabilities, effectively learning complex features in laser interference scenarios. Simultaneously, it incorporates DCNv2 for efficient convolutional operations, and the stability of its deep network training ensures good model performance in complex data environments. By utilizing the improved ResNet50, key information in laser interference scenarios can be automatically extracted without excessive reliance on manual feature extraction, thus better adapting to complex and ever-changing interference scenarios.
[0089] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0090] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A laser interference scene recognition method based on deep learning, characterized in that, The method includes: Acquire images of laser interference scenes captured by the photoelectric imaging system; The laser interference scene image is input into a preset laser interference scene recognition model to obtain the corresponding laser interference state; the laser interference state is either that the photoelectric imaging system is interfered with by laser or that the photoelectric imaging system is not interfered with by laser. The preset laser interference scene recognition model is the optimal model obtained by training and comparing different deep learning networks using a laser interference scene sample set. The preset laser interference scene recognition model includes an input module, a convolutional pooling module, a ResNet residual module, and a pooling classification output module arranged sequentially. The ResNet residual module includes multiple residual sub-modules, and the number of residual units contained in different residual sub-modules varies. The residual unit includes a second convolutional layer, a DCNv2 convolutional sub-unit, a third convolutional layer, a CBAM attention sub-unit, and an external attention sub-unit arranged sequentially. The input end of the second convolutional layer serves as the input end of the residual unit. The input end of the second convolutional layer and the output end of the external attention sub-unit are connected through an addition operation, and the connection end serves as the output end of the residual unit. The DCNv2 convolutional subunit includes an offset generation unit and a sampling convolution unit; the CBAM attention subunit includes a channel attention unit and a spatial attention unit; the external attention mechanism in the external attention subunit is to enhance network capabilities by using two external memory matrices (Memory Units) independent of the input feature map as key and value.
2. The laser interference scene recognition method based on deep learning according to claim 1, characterized in that, The input module is used to receive the laser interference scene image; The convolutional pooling module includes a first convolutional layer, a batch normalization layer, a ReLU activation function layer, and a max pooling layer arranged sequentially. The residual unit includes convolution processing, DCNv2 convolution processing, CBAM attention mechanism, and external attention mechanism; The pooling classification output module is used to output the laser interference status; the pooling classification output module includes a first global average pooling layer, a fully connected layer, a Softmax classification layer and an output layer arranged in sequence.
3. The laser interference scene recognition method based on deep learning according to claim 2, characterized in that, The offset generation unit is used to: calculate the sampling point offset of the convolution kernel in the horizontal and vertical directions of the feature map for the received feature map through a preset convolution operation; The sampling convolution unit is used to: perform bilinear interpolation based on the sampling point offset and the feature map to determine the sampling points of the convolution kernel on the feature map; and perform a convolution operation on the feature map according to the sampling points of the convolution kernel.
4. The laser interference scene recognition method based on deep learning according to claim 2, characterized in that, The feature map received by the channel attention unit and the feature map output by the channel attention unit are multiplied and then input into the spatial attention unit. The feature map received by the spatial attention unit and the feature map output by the spatial attention unit are multiplied and then input to the external attention subunit.
5. The laser interference scene recognition method based on deep learning according to claim 4, characterized in that, The channel attention region includes a first global max pooling layer, a second global average pooling layer, a shared fully connected layer, a second global max pooling layer, a third global average pooling layer, and a first Sigmoid function layer; The input terminals of the first global max pooling layer and the second global average pooling layer are both connected to the output terminal of the third convolutional layer; The outputs of the first global max pooling layer and the second global average pooling layer are both connected to the input of the shared fully connected layer. The output of the shared fully connected layer is connected to the input of the second global max pooling layer and the input of the third global average pooling layer, respectively. After performing an addition operation on the outputs of the second global max pooling layer and the third global average pooling layer, they are connected to the first sigmoid function layer; the first sigmoid function layer is used to activate and output the feature map. The spatial attention region includes an average pooling layer, a max pooling layer, a fourth convolutional layer, and a second sigmoid function layer arranged sequentially.
6. The laser interference scene recognition method based on deep learning according to claim 1, characterized in that, The process of determining the preset laser interference scene recognition model includes: Obtain a laser interference scene sample set; each sample in the laser interference scene sample set includes a laser interference sample image and the corresponding laser interference state; The laser interference scenario sample set is divided into a training set, a validation set, and a test set. The training set is used to train different deep learning networks to obtain multiple optimized deep learning models. Using the validation set, each of the optimized deep learning models is validated and adjusted to obtain multiple final deep learning models. Using the test set, each of the final deep learning models is tested to obtain multiple corresponding test results; By comparing multiple test results, the best-performing final deep learning model is designated as the preset laser interference scene recognition model.
7. The laser interference scene recognition method based on deep learning according to claim 1, characterized in that, The process of acquiring the laser interference scenario sample set includes: Multiple laser interference sample images were acquired using an optoelectronic imaging system; different laser interference sample images were subjected to laser interference of different power. Determine the laser interference state corresponding to each laser interference sample image; All the laser interference sample images are preprocessed; the preprocessing operations include standardization, normalization and data augmentation. Any preprocessed laser interference sample image and its corresponding laser interference state constitute a sample; multiple samples constitute a laser interference scene sample set.