Small target detection method and device
By using cat-eye effect and parallel feature extraction of convolutional branch object detection technology, the problem of high false alarm rate and difficult real-time processing of traditional technology in complex backgrounds is solved, and more efficient target recognition and position prediction are achieved.
Patent Information
- Application Number
- CN202510316085.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-03-18
AI Technical Summary
Traditional optoelectronic target detection technology is susceptible to noise interference in complex backgrounds, resulting in a high false alarm rate. Convolutional neural networks lack the feature extraction ability of low signal-to-noise ratio echo images, making it difficult to meet the real-time processing needs.
By stimulating the cat-eye effect of the target optical system, the target echo signals of multiple bands are collected and preprocessed to generate a reconstruction feature map of multiple bands. Then, these feature maps are input into the object detection model with parallel feature extraction convolutional branches, and the attention mechanism is used to fuse the output to generate the target information feature map. At the same time, use the position prediction model to generate future position coordinates.
It improves the target recognition capability, reduces the false alarm rate, enhances the target detection performance in complex backgrounds, and meets the real-time processing requirements.
Smart Images

Figure CN119850929B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of target detection, and in particular to a small target detection method and device. Background Art
[0002] Traditional optoelectronic target detection technology mainly relies on passive imaging (such as infrared and visible light cameras), but it is easy to cause false detection or missed detection of targets with changes in ambient light or complex background noise. The cat's eye effect originates from the reflection characteristics of the lens of optical instruments. Its echo signal intensity is much higher than the background environment and is suitable for long-distance detection of targets.
[0003] Traditional target detection methods rely on threshold segmentation or template matching, but they are easily affected by noise in complex backgrounds, resulting in a high false alarm rate. The existing convolutional neural network (CNN) has insufficient feature extraction capabilities for low signal-to-noise ratio echo images and a large number of model parameters, making it difficult to meet real-time processing requirements.
[0004] Therefore, it is necessary to improve image processing algorithms to enhance target recognition capabilities. Summary of the invention
[0005] The purpose of the present invention is to provide a small target detection method and device, which can solve at least one of the above-mentioned technical problems. The specific solution is as follows:
[0006] According to a specific embodiment disclosed in the present invention, the first aspect of the present invention discloses a small target detection method based on the cat's eye effect of an optical system, comprising:
[0007] The target echo signals generated in multiple bands due to the cat's eye effect are collected and preprocessed to obtain the reconstructed feature maps of multiple bands;
[0008] The reconstructed feature maps of the multiple bands are respectively input into a target detection model having a corresponding number of parallel feature extraction convolution branches, and the attention weights are assigned to each of the feature extraction convolution branches and then fused and output to obtain the information feature map of the target; the expression of the weight of each of the feature extraction convolution branches is:
[0009] (1)
[0010] in, For the band, is the activation function;
[0011] Expressing the i The band is The reconstructed feature map Apply the average pooling operation to obtain a feature vector;
[0012] Indicates the band x The weight matrix of the attention module;
[0013] Represents a bias vector with the same dimensions as the output feature;
[0014] The positioning information feature map of the target is input into the position prediction model to generate future position coordinates.
[0015] Preferably, collecting the echo signal of the target generated by the cat's eye effect and preprocessing it includes:
[0016] A multi-band laser is used to synchronously emit detection lasers of different wavelengths through a beam splitter prism, and the reflection characteristics of the cat's eye effect are used to capture the echo signals of multiple bands of the target optical window to form point spot images of different bands.
[0017] The multi-point light spot image is subjected to noise reduction processing to generate a reconstructed feature map of multiple bands.
[0018] Preferably, the step of performing noise reduction processing on the point light spot image to generate a reconstructed feature map of multiple bands includes:
[0019] Establishing the echo intensity model of the point spot image , performing wavelet transform and inverse wavelet transform on the echo intensity model to obtain the transformed echo intensity model ;
[0020] Perform principal component analysis on the transformed echo intensity model to obtain the reconstructed feature graphs of the multiple bands .
[0021] Preferably, principal component analysis is performed on the transformed echo intensity model to obtain the reconstructed characteristic graphs of the multiple bands, including:
[0022] Calculate the transformed echo intensity model The mean And centralize, the expression is: ;
[0023] Calculate the echo intensity model after centering The covariance matrix of
[0024] Perform eigenvalue decomposition on the covariance matrix and retain the previous k The eigenvector corresponding to the largest eigenvalue ,
[0025] According to the feature vector and the centralized echo intensity model , obtain the reconstructed feature maps of the multiple bands , the expression is: .
[0026] Preferably, the target detection network is based on the YOLOv5s architecture, including:
[0027] Multiple parallel feature extraction convolution branches use Ghost convolution to replace the original ordinary convolution to extract the features of the reconstructed feature map of each band respectively , the expression is:
[0028] (2)
[0029] in, X is the reconstructed feature map of the multiple bands The matrix formed;
[0030] and are two sets of convolution kernels, represents a linear transformation;
[0031] The cross-band attention fusion module is used to assign weights to each of the feature extraction convolution branches and generate feature maps after weighted fusion. : ,in, n is the number of bands.
[0032] Preferably, the target detection network is based on the YOLOv5s architecture and further includes: an efficient channel attention mechanism ECA-Net, in which Perform channel attention on it to generate a weighted feature map , the expression is:
[0033] (3)
[0034] in, is the channel attention weight coefficient;
[0035] Represents a one-dimensional convolution operation.
[0036] Preferably, the position prediction model includes: a multi-layer long short-term memory network and a Kalman filter. The long short-term memory network is used to predict the target position according to the positioning information feature map of the target, and the Kalman filter is used to optimize the prediction result of the long short-term memory network to generate the future position coordinates of the target.
[0037] Preferably, the input feature vector of the long short-term memory network includes: position, speed, echo intensity change rate and Doppler shift related to the target; the feature vector is normalized in the interval [0, 1];
[0038] The LSTM network has 256 or 512 hidden units.
[0039] Preferably, the process of updating the predicted state by the Kalman filter is expressed as:
[0040] (4)
[0041] The process of updating the Kalman filter estimation covariance matrix is expressed as:
[0042] (5)
[0043] in, represents the Kalman gain, Indicates the measured position of the target provided by the new echo signal;
[0044] H represents the observation matrix; represents the updated estimated covariance matrix, I represents the identity matrix;
[0045] Indicates the target location for the update.
[0046] According to a specific embodiment disclosed in the present invention, a second aspect of the present invention discloses a small target detection device, comprising:
[0047] A signal acquisition unit collects target echo signals generated in multiple bands due to the cat's eye effect and performs preprocessing to obtain reconstructed feature maps of multiple bands;
[0048] The feature extraction unit inputs the reconstructed feature maps of the multiple bands into a target detection model having a corresponding number of parallel feature extraction convolution branches, assigns attention weights to each of the feature extraction convolution branches, and then fuses the outputs to obtain an information feature map of the target; the expression of the weight of each of the feature extraction convolution branches is:
[0049] (1)
[0050] in, For the band, is the activation function;
[0051] Indicates i The band is The reconstructed feature map Apply the average pooling operation to obtain a feature vector;
[0052] Indicates the specific band x The weight matrix of the attention module;
[0053] Represents a bias vector with the same dimensions as the output feature;
[0054] The prediction unit inputs the positioning information feature map of the target into the position prediction model to generate future position coordinates.
[0055] Compared with the prior art, the above solution disclosed in the present invention has at least the following beneficial effects:
[0056] The present invention realizes active detection of the target by stimulating the cat's eye effect of the target optical system, improves the echo image processing efficiency by improving the YOLOv5s deep learning network, and reduces the target missed detection rate by combining the position prediction model. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] The drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the disclosure of the present invention, and together with the specification, are used to explain the principles disclosed in the present invention. Obviously, the drawings described below are only some embodiments disclosed in the present invention, and for ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work. In the drawings:
[0058] Figure 1 A detection flow chart of a small target detection method of the present invention;
[0059] Figure 2 It is a schematic diagram of algorithm improvement of an embodiment of the present invention;
[0060] Figure 3 Schematic diagram of a small target detection device according to an embodiment of the present invention;
[0061] Figure 4 It is a schematic diagram of the structure of the electronic device provided in this embodiment. DETAILED DESCRIPTION
[0062] In order to make the purpose, technical scheme and advantages disclosed in the present invention clearer, the present invention will be further described in detail below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments disclosed in the present invention, rather than all the embodiments. Based on the embodiments disclosed in the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection disclosed in the present invention.
[0063] The terms used in the embodiments disclosed in the present invention are only for the purpose of describing specific embodiments, and are not intended to limit the disclosure of the present invention. The singular forms "a", "said" and "the" used in the embodiments disclosed in the present invention and the appended claims are also intended to include plural forms, unless the context clearly indicates other meanings, and "multiple" generally includes at least two.
[0064] It should be understood that the term "and / or" used in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship.
[0065] It should be understood that although the terms first, second, third, etc. may be used to describe in the disclosed embodiments of the present invention, these descriptions should not be limited to these terms. These terms are only used to distinguish the descriptions. For example, without departing from the scope of the disclosed embodiments of the present invention, the first may also be referred to as the second, and similarly, the second may also be referred to as the first.
[0066] As used herein, the words "if" and "if" may be interpreted as "at the time of" or "when" or "in response to determining" or "in response to detecting", depending on the context. Similarly, the phrases "if it is determined" or "if (stated condition or event) is detected" may be interpreted as "when it is determined" or "in response to determining" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)", depending on the context.
[0067] It should also be noted that the term "includes", "comprising" or any other variation thereof is intended to cover non-exclusive inclusion, so that a commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprising a ..." do not exclude the existence of other identical elements in the commodity or device including the elements.
[0068] Optoelectronic imaging equipment is generally composed of optical lenses and photoelectric sensors. When laser is used to scan and detect the space where the target is located, once the laser is irradiated to the target photoelectric device, a reflected echo will be formed returning along the original path. This effect is called the cat's eye effect.
[0069] Active detection based on the cat's eye effect means that when the detection laser enters the optical window of the target, the cat's eye effect of the optical system shows that the echo beam contains a variety of information about the target. After passing through the high-speed acquisition system, the signal is sent to the control processing system for processing. By demodulating and inverting the echo signal, the type of detector of the target's optoelectronic equipment, as well as its working band, modulation method and other characteristics can be identified, so as to quickly locate it and blind or attack it.
[0070] The following is combined with Figure 1-2 Detailed description of the alternative embodiments of the present disclosure.
[0071] Example 1
[0072] like Figure 1 As shown, the present invention provides a small target detection method, comprising the following steps:
[0073] Step S100: collecting target echo signals generated in multiple bands due to the cat's eye effect and performing preprocessing to obtain reconstructed feature maps of the multiple bands.
[0074] After the target is captured and roughly tracked, the illumination laser emits a high-repetition-rate illumination laser to illuminate the target. The laser within the wavelength band of the target optical system is reflected by the target detector, forming an echo signal in the opposite direction of the incident light. The cat's eye echo signal generated by the target optical system is received by an imaging detector with high frame rate, high spatial resolution and narrow spectrum to form a point spot image.
[0075] Furthermore, step S100 includes the following steps:
[0076] Step S101, using a multi-band laser to synchronously emit light beams of different wavelengths through a beam splitter prism, using the reflection characteristics of the cat's eye effect to capture echo signals of multiple bands of the target optical window, forming point light spot images of different bands; performing noise reduction processing on the multi-point light spot image to generate a reconstructed feature map of multiple bands.
[0077] Among them, a three-band fiber laser can be used, and a beam splitter prism is used to emit a detection laser with three wavelengths. After the detection laser enters the optical window of the target to be measured, an echo signal with three bands is generated due to the cat's eye effect. The reflected echo signal is received by the imaging detector and forms a point spot image with three bands.
[0078] Specifically, the laser uses a three-band fiber laser (wavelength 1550nm, 1064nm, 850nm) with adjustable output power (10mW~500mW). The pulse modulation signal is generated by FPGA to drive the laser to emit a synchronous modulated beam. A collimating beam expander (beam expansion ratio 10:1) is used to form a detection light curtain with a coverage range of ≥100m×100m.
[0079] A cooled EMCCD camera with a quantum efficiency of >90%@850-1550nm and a frame rate of ≥100fps was used to collect echo image information. Morlet wavelet transform was performed on the echo signal to extract the time-frequency feature matrix (size 64×64). PCA was used for dimensionality reduction (95% energy principal component was retained).
[0080] Step S102, performing noise reduction processing on the multi-point spot image to generate a reconstructed feature map of multiple bands, specifically comprising the following steps:
[0081] Step S1021: Establishing the echo intensity model of the spot image , the expression is:
[0082] (4)
[0083] in, Represented in space coordinates The wavelength is The strength of the echo signal;
[0084] represents the emission power of the laser, represents the efficiency of the optical system used for detection;
[0085] represents the target reflectivity, R Indicates the target distance; represents the atmospheric attenuation coefficient;
[0086] Indicates the target in space coordinates Surface geometric characteristics.
[0087] Echo Intensity Model Perform wavelet transform, extract time-frequency features, pre-process echo data acquisition information, and extract wavelet coefficients after wavelet transform , the expression is:
[0088] (5)
[0089] in, Indicates the echo signal of a certain band The sampling sequence in the time dimension,
[0090] ;
[0091] represents the wavelet basis function, a is the scale factor, b is the translation factor, and * is the matrix operation.
[0092] right Threshold processing is performed to remove noise, and the distance in the echo intensity model is R and atmospheric attenuation Dynamically adjust threshold , to enhance the retention of long-distance weak signals.
[0093] in, ; is the local noise standard deviation, It is adjusted as a dynamic factor combining target characteristics and detection distance.
[0094] Get the processed coefficients , and then perform inverse wavelet transform to obtain the transformed echo intensity model ;
[0095] (6)
[0096] in, is the normalization constant of the wavelet function.
[0097] Step S103: Perform principal component analysis on the transformed echo intensity model to obtain the reconstructed feature graphs of the multiple bands. . By performing principal component analysis (PCA) on multi-band data to suppress noise, the number of bits of data can be reduced, and the efficiency and performance of the algorithm can be improved.
[0098] The specific steps include:
[0099] Step S1031: Calculate the transformed echo intensity model The mean .
[0100] Step S1032: echo intensity model Centralization, the expression is: .
[0101] Step S1033: Calculate the centralized echo intensity model The covariance matrix of is:
[0102]
[0103] in, N is the number of data points, The characteristic contribution of high repetition rate laser bands is enhanced according to band weighting.
[0104] Step S1034: performing eigenvalue decomposition on the covariance matrix. ;in, is the eigenvector matrix, is the characteristic diagonal matrix.
[0105] Before Retention k The eigenvector corresponding to the largest eigenvalue , obtain the reconstructed feature maps of the multiple bands , the expression is: .
[0106] Specifically, when processing cat-eye echo data, the variance contribution rate of each principal component is calculated, and then the cumulative variance contribution rate is calculated, and the threshold is set to 95%. If the cumulative variance explanation rate of the first three principal components reaches 95%, K is 3. At this time, the three principal components can retain 95% of the information of the original echo data, which is enough to reflect the main characteristics of the data, while removing some noise and redundant information.
[0107] Furthermore, the reconstructed feature maps of multiple bands are obtained The information is input into the target detection model for processing to generate a target information feature map containing the target optical system visual information and position information.
[0108] When high repetition rate laser is emitted, different bands are affected differently by atmospheric attenuation and target reflection characteristics. Among the received cat-eye echo signals, the near-infrared band (such as 1.06 μm ) has relatively small transmission loss in the atmosphere, and the optical components of most optoelectronic devices have strong cat's eye reflection in this band; short-wave infrared band (such as 1.55 μm ) has a good reflection effect on targets of certain materials and is relatively less affected by background radiation; mid-infrared bands (such as 3-5 μm ) can effectively detect targets with thermal characteristics and can be used to distinguish the temperature difference between the target and the background. Therefore, the image characteristics of these three bands are mainly referred to in the image processing process.
[0109] Step S200: reconstruct the characteristic graphs of the multiple bands The information feature maps of the target are obtained by respectively inputting the corresponding parallel feature extraction convolution branches into the target detection model, assigning attention weights to each of the feature extraction convolution branches and then merging the outputs.
[0110] The target detection network in this embodiment is based on the YOLOv5s architecture, and the feature map is reconstructed Echo images in different bands Input the three parallel feature extraction convolution branches in the YOLOv5s head network respectively.
[0111] Compared with the original YOLOv5s architecture, this embodiment uses Ghost convolution to replace the traditional convolution layer to reduce the number of network parameters and the amount of calculation. The Ghost module extracts the features of the reconstructed feature maps of the three bands through linear transformation. , the expression is:
[0112] (2)
[0113] in, ; and are two sets of convolution kernels, Represents a linear transformation.
[0114] Specifically, considering the target size and feature size in the cat's eye echo image, for the near-infrared band (1.06 μ m ), since it is mainly used to detect long-distance targets, the size of the generated image of the target is relatively small, and a smaller convolution kernel size is sufficient. A 3×3 convolution kernel size is selected, and the number of convolution kernels in the two groups is set to be , ; For the short-wave infrared and mid-infrared bands, the convolution kernel size can be adjusted to 5×5 according to the detailed characteristics of the target and the image resolution, and the number of convolution kernel groups can be appropriately increased. , To better capture the target features.
[0115] Furthermore, a cross-band attention fusion module is added to learn the correlation between features of different bands and assign weights to each band. By weighted fusion of features of different bands, important features are enhanced and unimportant features are suppressed. The weight matrix is initialized according to the sensitivity of different bands to echo signals. For the near-infrared band, since it is more sensitive to long-distance cat's eye targets, a larger weight is given when initializing the weight matrix. For the short-wave infrared band and the medium-wave infrared band, the weights are reasonably assigned according to their material characteristics and thermal feature detection capabilities of the target.
[0116] Specifically, the cross-band attention fusion module uses the features output by Ghost convolution As input, calculate the attention weight of each band :
[0117] (1)
[0118] in, For the band, is the Sigmoid activation function;
[0119] Indicates iReconstructed feature map of bands Apply the average pooling operation to obtain a feature vector;
[0120] Indicates the specific band x The weight matrix of the attention module;
[0121] Represents a bias vector with the same dimensions as the output feature;
[0122] In this embodiment, the specific band x It is the near infrared band, short wave infrared band or mid infrared band.
[0123] In this embodiment, the background noise level of different bands is taken into consideration, and the bias vector value is appropriately increased for the band with larger noise, such as the mid-infrared band. (The vector dimension should be consistent with the output dimension). The other two bands can refer to the smaller bias vector such as .
[0124] Finally, the feature map is generated after weighted fusion : .
[0125] The feature image after the cross-band attention fusion module contains the dominant features of each band, but the processed feature image usually retains certain band-related information in the channel dimension, and the features originally extracted from different bands may be distributed in different channels.
[0126] Therefore, the target detection network based on the YOLOv5s architecture in this embodiment also includes: an efficient channel attention mechanism ECA-Net. By learning the dependencies between channels, adaptively enhancing target-related features, suppressing background noise, and improving the detection accuracy of low signal-to-noise ratio echo images, the ECA module can significantly improve the stability and generalization ability of target detection in complex backgrounds.
[0127] Specifically, in the feature map Perform channel attention weight coefficient on ;
[0128]
[0129] in, Represents a one-dimensional convolution operation, which is used to process the feature vector after global average pooling and calculate the size of the one-dimensional convolution kernel according to the number of channels of the feature map;
[0130] GAP Indicates that average pooling generates a feature map after weighted fusion Convert to ;
[0131] One-dimensional convolution can reduce the number of parameters while capturing the dependencies between channels. The sigmoid function is selected to compress the convolution operation to between 0 and 1, generate weight coefficients, and consider the characteristics of signals in different bands.
[0132] Finally, the weighted feature map is generated , the expression is:
[0133] (3)
[0134] in, Represents a weighted multiplication and summation operation.
[0135] The weighted feature map Continue to input into the head network to obtain the final target information feature map. At this time, the target information feature map contains the target's visual information and position information.
[0136] In other embodiments, it also includes Figure 2 The steps for training the target detection network are shown.
[0137] During the training process, synthetic datasets (including complex scenes such as fog, night, and occlusion) were used to evaluate the model performance and random noise injection (signal-to-noise ratio between 5 and 20 dB) and motion blur (blur size of 15×15) were used. The model was trained on a relevant dataset containing targets and then fine-tuned. The cosine annealing strategy was used to adjust the learning rate, with an initial learning rate of 1×10 -4 .
[0138] Step S300: input the positioning information feature map of the target into a position prediction model to generate future position coordinates.
[0139] Specifically, the position prediction model includes: a multi-layer long short-term memory network and a Kalman filter. The long short-term memory network is used to predict the target position according to the positioning information feature map of the target, and the Kalman filter is used to optimize the prediction result of the long short-term memory network to generate the future position coordinates of the target.
[0140] Furthermore, the use of LSTM networks for image processing and target localization is more focused on processing single images or static scenes than other methods. At the same time, LSTM networks are more flexible in dynamically processing target scenes, predicting future states, and handling multi-target problems.
[0141] In this embodiment, due to the particularity of the detected target, the target's motion pattern is mostly a complex trajectory, which requires LSTM to have a strong learning ability, appropriately increase the number of hidden units, and improve the model fitting ability. Compared with the general target detection scenario using 128 hidden units, this embodiment adds 256 or 512 hidden units. And use 2-3 layers of long short-term memory network to let the model learn more complex time series features, observe the model convergence and prediction accuracy on the training set and verification set, and use the prediction accuracy and mean absolute error for evaluation. The accuracy on the training set and verification set gradually increases and stabilizes after reaching a certain level. It can be determined that the model tends to converge. If the learning ability is insufficient, consider increasing the number of layers.
[0142] Furthermore, the input feature vector of the LSTM network is Based on the cat's eye echo image, data related to the position, speed, echo intensity and frequency of the target are added. Assuming that the echo intensity increases rapidly in consecutive time steps, it may mean that the target is approaching, so the echo intensity change rate is calculated. and Doppler shift Also used as input feature value.
[0143] Then the input feature vector of the long short-term memory network can be expressed as: .
[0144] Since the numerical ranges of different eigenvalues may vary greatly, in order to avoid excessive impact on the training of long short-term memory networks, the input features need to be normalized and normalized to the interval [0, 1] according to the actual measurement range.
[0145] This embodiment predicts the target position through the long short-term memory network, uses the obtained data results as prior information, and uses the Kalman filter combined with the new cat-eye echo measurement data to update the state. Especially when the target state changes suddenly, the Kalman filter can use the new echo measurement information to quickly correct the prediction results of the long short-term memory network, making the tracking more accurate.
[0146] Specifically, assume that the new echo information provides the measured position of the target , using the Kalman gain To update the prediction status, there is the following expression:
[0147] (4).
[0148] When the target mutates, the process noise covariance matrix can be increased Q The diagonal elements of the Kalman filter enable the Kalman filter to respond quickly to changes in the target state. The Kalman gain The need is to increase accordingly, that is, the proportion of new measured values relative to the estimated values becomes larger, so that the prediction results can be adjusted quickly to adapt to the target mutation.
[0149] Finally update the estimated covariance matrix for:
[0150] (5)
[0151] in, represents the Kalman gain, Indicates the measured position of the target provided by the new echo signal;
[0152] H represents the observation matrix; represents the updated estimated covariance matrix, I represents the identity matrix;
[0153] represents the updated target location, Indicates the target position in the previous echo signal.
[0154] When predicting future position coordinates, the embodiment of the present invention first converts the image coordinate system to the polar coordinate system for ease of processing and prediction. Input the feature vectors of the five historical frames into the long short-term memory network X ,
[0155] ;
[0156] Generate feature vectors that predict future positions and velocities .
[0157] The improved prediction model in this embodiment has low computational complexity and can meet the real-time tracking requirements of ≥30FPS. The average accuracy of the overall image processing unit can be improved by about 5%-10%, and real-time processing can be achieved on the device, with a single frame processing time of ≤15 ms .
[0158] Example 2
[0159] The present invention also provides an apparatus embodiment that is consistent with the above embodiment, which is used to implement the method steps described in the above embodiment. The explanation based on the same name meaning is the same as the above embodiment, and has the same technical effect as the above embodiment, which will not be repeated here.
[0160] like Figure 3 As shown, the present invention discloses a small target detection device, comprising:
[0161] The signal acquisition unit 302 acquires target echo signals generated in multiple bands due to the cat's eye effect and performs preprocessing to obtain reconstructed feature maps of the multiple bands;
[0162] The feature extraction unit 304 inputs the reconstructed feature maps of the multiple bands into a target detection model having a corresponding number of parallel feature extraction convolution branches, assigns attention weights to each of the feature extraction convolution branches, and then fuses the outputs to obtain an information feature map of the target; the expression of the weight of each of the feature extraction convolution branches is:
[0163] (1)
[0164] in, For the band, is the Sigmoid activation function;
[0165] Indicates i The band is The reconstructed feature map Apply the average pooling operation to obtain a feature vector;
[0166] Indicates the specific band x The weight matrix of the attention module;
[0167] Represents a bias vector with the same dimension as the output feature;
[0168] The prediction unit 306 inputs the positioning information feature map of the target into the position prediction model to generate future position coordinates.
[0169] Example 3
[0170] like Figure 4 As shown, this embodiment provides an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method steps described in the above embodiment.
[0171] Example 4
[0172] The disclosed embodiment of the present invention provides a non-volatile computer storage medium, wherein the computer storage medium stores computer executable instructions, and the computer executable instructions can execute the method steps described in the above embodiment.
[0173] Example 5
[0174] Reference below Figure 4, which shows a schematic diagram of the structure of an electronic device suitable for implementing the disclosed embodiment of the present invention. The terminal device in the disclosed embodiment of the present invention may include but is not limited to mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 4 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments disclosed in the present invention.
[0175] like Figure 4 As shown, the electronic device may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 401, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 402 or a program loaded from a storage device 408 to a random access memory (RAM) 403. In RAM 403, various programs and data required for the operation of the electronic device are also stored. The processing device 401, ROM 402, and RAM 403 are connected to each other via a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.
[0176] Typically, the following devices may be connected to the I / O interface 405: an input device 406 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 407 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 408 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 409. The communication device 409 may allow the electronic device to communicate with other devices wirelessly or by wire to exchange data. Although Figure 4 An electronic device having various devices is shown, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead.
[0177] In particular, according to the embodiments disclosed in the present invention, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the embodiments disclosed in the present invention include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 409, or installed from the storage device 408, or installed from the ROM 402. When the computer program is executed by the processing device 401, the above-mentioned functions defined in the method of the embodiment disclosed in the present invention are executed.
[0178] It should be noted that the computer-readable medium disclosed in the present invention may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device. In the present invention, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable program code. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Computer readable signal media may also be any computer readable medium other than computer readable storage media, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer readable medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0179] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0180] Computer program code for performing the operations disclosed in the present invention may be written in one or more programming languages or a combination thereof, including object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0181] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to the various embodiments disclosed in the present invention. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some implementations as replacements, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0182] The units involved in the embodiments disclosed in the present invention may be implemented by software or hardware, wherein the name of a unit does not limit the unit itself in some cases.
Claims
1. A small target detection method, characterized in that: Cat's eye effect based on optical system, including: The target echo signals generated in multiple bands due to the cat's eye effect are collected and preprocessed to obtain the reconstructed feature maps of multiple bands; The reconstructed feature maps of the multiple bands are respectively input into a target detection model having a corresponding number of parallel feature extraction convolution branches, and the attention weights are assigned to each of the feature extraction convolution branches and then fused and output to obtain the information feature map of the target; the expression of the weight of each of the feature extraction convolution branches is: a λi =σ(ψ x ·GAP(F λi )+B x ) (1) Among them, λ is the band and σ is the activation function; GAP(F λi ) represents the reconstructed feature map F for the i-th band λ λi Apply the average pooling operation to obtain a feature vector; ψ x Represents the weight matrix of the attention module for band x; B x Represents a bias vector with the same dimension as the output feature; The information feature map of the target is input into the position prediction model to generate future position coordinates.
2. The method according to claim 1, characterized in that The collecting and preprocessing of the target echo signal generated by the cat's eye effect includes: A multi-band laser is used to synchronously emit detection lasers of different wavelengths through a beam splitter prism, and the reflection characteristics of the cat's eye effect are used to capture the echo signals of multiple bands of the target optical window to form point spot images of different bands. The point light spot image is subjected to noise reduction processing to generate a reconstructed feature map of multiple bands.
3. The method according to claim 2, characterized in that The step of performing noise reduction processing on the point light spot image to generate a reconstructed feature map of multiple bands includes: Establish the echo intensity model I of the point spot image λ (x, y), perform wavelet transform and inverse wavelet transform on the echo intensity model to obtain the transformed echo intensity model I λ ′(t); The transformed echo intensity model is subjected to principal component analysis to obtain the reconstructed characteristic graph I of the multiple bands. PCA (t).
4. The method according to claim 3, characterized in that The transformed echo intensity model is subjected to principal component analysis to obtain reconstruction feature maps of the multiple bands, including: Calculate the transformed echo intensity model I λ The mean μ of ′(t) is centered and expressed as: λ ″(t)=I λ ′(t)-μ; Calculate the echo intensity model after centralization I λ The covariance matrix of ″(t); Perform eigenvalue decomposition on the covariance matrix and retain the eigenvectors U corresponding to the first k largest eigenvalues k , According to the feature vector U k and the centralized echo intensity model I λ ″(t), obtain the reconstructed feature graph I of the multiple bands PCA (t), the expression is:
5. The method according to claim 1, characterized in that The target detection network is based on the YOLOv5s architecture and includes: Multiple parallel feature extraction convolution branches use Ghost convolution to replace the original ordinary convolution to extract the features F of the reconstructed feature map of each band respectively. λi , the expression is: F λi =X*K1+X*K2Φ(X) (2) Wherein, X is the reconstructed feature map I of the multiple bands PCA (t) the matrix formed; K1 and K2 are two sets of convolution kernels, and Φ represents linear transformation; The cross-band attention fusion module is used to assign weights to each of the feature extraction convolution branches and generate a feature map F′ after weighted fusion: Where n is the number of bands.
6. The method according to claim 5, characterized in that The target detection network is based on the YOLOv5s architecture and also includes an efficient channel attention mechanism ECA-Net, which performs channel attention on the feature map F′ to generate a weighted feature map F ECA , the expression is: F ECA =ψ c ⊙F′ (3) Among them, ψ c =σ(Conv 1D (GAP(F′))) is the channel attention weight coefficient; Conv 1D Represents a one-dimensional convolution operation.
7. The method according to claim 1, characterized in that The position prediction model includes: a multi-layer long short-term memory network and a Kalman filter. The long short-term memory network is used to predict the target position according to the positioning information feature map of the target, and the Kalman filter is used to optimize the prediction result of the long short-term memory network to generate the future position coordinates of the target.
8. The method according to claim 7, characterized in that The input feature vector of the long short-term memory network includes: position, speed, echo intensity change rate and Doppler shift related to the target; the feature vector is normalized in the interval [0, 1]; The LSTM network has 256 or 512 hidden units.
9. The method according to claim 7, characterized in that: The process of updating the predicted state by the Kalman filter is expressed as: The process of updating the Kalman filter estimation covariance matrix is expressed as: P t =(I-K t H)P t-1 (5) Among them, K t represents the Kalman gain, z t Indicates the measured position of the target provided by the new echo signal; H represents the observation matrix; P t represents the updated estimated covariance matrix, and I represents the identity matrix; Indicates the target location for the update.
10. A small target detection device, characterized in that: include: A signal acquisition unit collects target echo signals generated in multiple bands due to the cat's eye effect and performs preprocessing to obtain reconstructed feature maps of multiple bands; The feature extraction unit inputs the reconstructed feature maps of the multiple bands into a target detection model having a corresponding number of parallel feature extraction convolution branches, assigns attention weights to each of the feature extraction convolution branches, and then fuses the outputs to obtain an information feature map of the target; the expression of the weight of each of the feature extraction convolution branches is: a λi =σ(ψ x ·GAP(F λi )+B x ) (1) Among them, λ is the band and σ is the activation function; GAP(F λi ) represents the reconstructed feature map F for the i-th band λ λi Apply the average pooling operation to obtain a feature vector; ψ x Represents the weight matrix of the attention module for band x; B x Represents a bias vector with the same dimension as the output feature; The prediction unit inputs the information feature map of the target into the position prediction model to generate future position coordinates.
Citation Information
Patent Citations
Target extraction method based on cat eye effect
CN118135234A
Radar multi-frame image target detection and tracking integrated processing method
CN118962627A