Method and equipment for generating decoy jamming bomb discrimination model
By coupling data from the seeker's infrared focal plane array and inertial measurement unit during the terminal guidance phase of infrared homing, a decoy flare discrimination model is generated, solving the problem of difficulty in distinguishing between real targets and decoy flares, and achieving high-accuracy discrimination under complex conditions.
Patent Information
- Application Number
- CN202510915296.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-03
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2045-07-03
AI Technical Summary
Existing technologies struggle to effectively distinguish between real targets and decoy flares during the terminal guidance phase of infrared homing, resulting in a high false alarm rate, especially during rapid rolls or lateral maneuvers, and performance degrades under complex cloud cover or parallax conditions.
By coupling data based on the seeker's infrared focal plane array and inertial measurement unit, a preset network is used to generate a decoy jammer discrimination model, including the construction of spatiotemporal embedding vectors, a single-layer fully connected classification head, and an adversarial discriminator. Combined with extreme value and pseudo-label training, the discrimination accuracy is improved.
It effectively reduced the false alarm rate of decoy flares and improved the discrimination accuracy and anti-interference capability under complex conditions.
Smart Images

Figure CN120852307A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of infrared homing terminal guidance technology, and particularly relates to a method and device for generating a decoy chaff discrimination model. Background Technology
[0002] During the infrared homing terminal guidance phase, the missile relies on the infrared focal plane array within the seeker head to detect, track, and lock onto the target. In real tactical scenarios, the enemy often interferes with the missile's guidance system by releasing decoy flares (such as high-temperature magnesium flares, chaff flares, etc.), causing it to deviate from the actual target and severely reducing guidance accuracy and strike effectiveness.
[0003] In related technologies, a combination of spectral difference and motion consistency is typically used to address the problem of distinguishing between true targets and decoys / flares during the terminal guidance phase. A typical process is as follows: First, all bright candidate points are extracted from a monocular infrared sequence using multi-threshold segmentation and morphological filtering. Then, an energy spectrum curve and a short-time velocity vector are constructed for each candidate point. These two types of features are mapped onto a two-dimensional plane, and Bayesian classification or support vector machines are used for initial screening. In a parallel branch, Kalman spectroscopy is used to fit the candidate points to continuous frame trajectories. Targets whose trajectory direction remains consistent with the line of sight over a long period are output as a high-confidence set, and a true / false label is given after a rule-based decision.
[0004] While the aforementioned scheme simultaneously introduces spectral difference and motion consistency, these two typically operate within independent pathways, and radiated energy and attitude dynamics are not fully coupled on the same feature representation. When the target enters a rapid roll or lateral maneuver, the correlation between the thermal field morphology and the velocity field is weakened, and the classifier can only rely on thresholds and rule trees for backend correction. Consequently, the distinguishability between decoys and real targets at the single-frame scale decreases, and the false alarm rate still increases significantly under complex cloud cover or parallax conditions. Summary of the Invention
[0005] The embodiments of this application provide a method and apparatus for generating a decoy chaff discrimination model, which can at least to a certain extent accurately discriminate decoy chaff and reduce the false alarm rate of decoy chaff.
[0006] Other features and advantages of this application will become apparent from the following detailed description, or may be learned in part from practice of this application.
[0007] According to a first aspect of the embodiments of this application, a method for generating a decoy chaff discrimination model is provided, comprising:
[0008] Based on the original image output by the infrared focal plane array inside the seeker and the attitude data of the projectile output by the inertial measurement unit, a time-consistent radiation sequence and attitude sequence are obtained.
[0009] The first encoder in the preset network is used to couple the radiation sequence and the attitude sequence to obtain the spatiotemporal embedding vector;
[0010] Adjust the single-layer fully connected classification head in the preset network according to the spatiotemporal embedding vector;
[0011] An initial discrimination model is constructed based on the first encoder and the adjusted single-layer fully connected classification head;
[0012] Based on the limiting values of the physical parameters of the decoy flares, the decoy sequence is obtained, whereby the decoy sequence is used to characterize the radiation sequence and attitude sequence of the decoy flares.
[0013] Based on the initial discrimination model and the decoy sequence, an intermediate discrimination model is obtained;
[0014] The intermediate discrimination model is trained using decoy sequences, time-consistent radiation sequences, and attitude sequences to obtain the target discrimination model.
[0015] In some embodiments, the infrared focal plane array within the seeker includes an infrared detector, and the inertial measurement unit includes a three-axis gyroscope. The interrupt triggering time of the infrared detector upon completion of integration readout is the same as the interrupt triggering time of the three-axis gyroscope upon completion of sampling. Based on the original image output by the infrared focal plane array within the seeker and the attitude data of the projectile output by the inertial measurement unit, a time-consistent radiation sequence and attitude sequence are obtained, including:
[0016] The gain matrix and bias matrix of the infrared focal plane array inside the seeker are corrected based on the focal plane temperature of the infrared focal plane array inside the seeker.
[0017] The linear radiation matrix is determined based on the original image, the corrected gain matrix, and the corrected bias matrix.
[0018] The current attitude data is rotated and projected to obtain the displacement field corresponding to the current attitude data.
[0019] The target radiation matrix corresponding to the current attitude data is determined using the displacement field, linear radiation matrix, and quadratic B-spline interpolation algorithm.
[0020] Based on the current attitude data and the target radiation matrix corresponding to the current attitude data, a time-consistent radiation sequence and attitude sequence are obtained.
[0021] In some embodiments, the radiation sequence and the attitude sequence are coupled using a first encoder in a preset network to obtain a spatiotemporal embedding vector, including:
[0022] Based on the radiation sequence and attitude sequence, multiple minimum prediction units are determined;
[0023] After mapping the attitude data in each smallest prediction unit to a planar vector field, it is concatenated with the corresponding target radiation matrix by channel dimension to obtain multiple multi-channel tensors.
[0024] The initial network is trained using multi-channel tensors to obtain the preset network. The initial network includes an initial encoder, a radiative decoding head, and an optical flow decoding head. The preset network includes a first encoder, a radiative decoding head, and an optical flow decoding head.
[0025] The multi-channel tensor is input into the preset network to obtain the first feature vector output by the first encoder, and the first feature vector is determined as the spatiotemporal embedding vector.
[0026] In some embodiments, an initial network is trained using multi-channel tensors to obtain a preset network, including:
[0027] The multi-channel tensor is input into the initial network to obtain the second feature vector output by the initial encoder, the radiation matrix prediction value output by the radiation decoder head, and the optical flow prediction value output by the optical flow decoder head;
[0028] Based on the second eigenvector corresponding to the multichannel tensor of the current smallest prediction unit, determine the positive cluster center of the same group;
[0029] Based on the second eigenvector corresponding to the multichannel tensor of any minimum prediction unit other than the current minimum prediction unit, determine the cross-group negative clustering center.
[0030] Substitute the predicted values of the radiation matrix and optical flow corresponding to the multi-channel tensor of the current smallest prediction unit, the positive cluster centers in the same group, and the negative cluster centers across groups into the loss function of the initial network. If the value of the loss function is lower than a preset threshold, the convergence of the loss function is determined, and the preset network is obtained.
[0031] In some embodiments, adjusting the single-layer fully connected classification head in a preset network according to the spatiotemporal embedding vector includes:
[0032] Clustering and filtering of the spatiotemporal embedding vectors yields the first spatiotemporal embedding vector, and the label corresponding to the first spatiotemporal embedding vector is obtained. The label includes a first label for representing the real target and a second label for representing the decoy chaff.
[0033] With the parameters of the first encoder unchanged, the single-layer fully connected classification head in the preset network is adjusted according to the first spatiotemporal embedding vector and the label corresponding to the first spatiotemporal embedding vector.
[0034] The vectors other than the first spatiotemporal embedding vector in the spatiotemporal embedding vector are used as the second spatiotemporal embedding vector, and the original image corresponding to the second spatiotemporal embedding vector is subjected to weak augmentation to obtain a set of pseudo-label candidates.
[0035] Weak augmented images and their corresponding pseudo-labels are selected from the candidate set of pseudo-labels based on a preset decreasing threshold sequence.
[0036] A weakly augmented image is subjected to strong augmentation processing to obtain a strongly augmented image;
[0037] With the parameters of the first encoder unchanged, the single-layer fully connected classification head in the preset network is adjusted according to the pseudo-labels corresponding to the weakly augmented image and the strongly augmented image.
[0038] In some embodiments, a decoy sequence is obtained based on the limiting values of the physical parameters of the decoy flares, including:
[0039] A parameter generator is constructed based on the limit values of the physical parameters of the decoy flares;
[0040] Based on the parameters output by the parameter generator, the decoy sequence is obtained.
[0041] In some embodiments, an intermediate discrimination model is obtained based on an initial discrimination model and a decoy sequence, including:
[0042] Unfreeze the penultimate layer of the first encoder in the initial discriminant model and the single fully connected classification head in the initial discriminant model to construct an adversarial discriminator;
[0043] Based on the adversarial discriminator, determine the confusion probability of the decoy sequence;
[0044] The adversarial discriminator is trained based on the confusion probability. After the first preset condition is met, the adversarial discriminator is trained again using the decoy sequence, the time-consistent radiation sequence and the attitude sequence. After the second preset condition is met, an intermediate discriminant model is constructed based on the first encoder and the trained adversarial discriminator.
[0045] In some embodiments, the intermediate discrimination model is trained using a decoy sequence, a time-consistent radiation sequence, and an attitude sequence to obtain a target discrimination model, including:
[0046] The decoy sequence is inserted into the radiation sequence and attitude sequence that are in the same time sequence according to the time axis to obtain the synthetic sequence;
[0047] The second encoder is obtained by adjusting the last two layers of the first encoder in the intermediate discriminant model based on the synthesized sequence.
[0048] Adjust the single-layer fully connected classification head of the adversarial discriminator in the second encoder and intermediate discriminant model according to the original image corresponding to the first spatiotemporal embedding vector;
[0049] Based on the adjusted second encoder and the adjusted adversarial discriminator, a target discrimination model is constructed.
[0050] In some embodiments, the method for generating the decoy chaff discrimination model further includes:
[0051] The three-dimensional convolutional kernel of the target discrimination model is decomposed into an intra-frame two-dimensional convolutional kernel and an inter-frame one-dimensional convolutional kernel;
[0052] Based on the contribution of the output channel of the target discrimination model to the backpropagation gradient corresponding to the hard example sequence, the output channel is pruned, where the hard example sequence is the decoy sequence with medium confidence.
[0053] Inference delay and anti-decoy accuracy of the target discrimination model were tested.
[0054] The target discrimination model is deployed when the inference delay time is less than the preset time and the anti-decoy accuracy is greater than the preset accuracy.
[0055] According to a second aspect of the embodiments of this application, a device for generating a decoy chaff discrimination model is provided, including a processor and a memory. The memory stores computer program instructions that can be executed by the processor. When the processor executes the computer program instructions, it implements the steps of the method as described in any of the first aspects above.
[0056] In this application, a time-consistent radiation sequence and attitude sequence are obtained based on the original image output by the infrared focal plane array inside the seeker and the attitude data of the projectile output by the inertial measurement unit. The radiation sequence and attitude sequence are coupled using a first encoder in a preset network to obtain a spatiotemporal embedding vector. The single-layer fully connected classification head in the preset network is adjusted according to the spatiotemporal embedding vector. An initial discrimination model is constructed based on the first encoder and the adjusted single-layer fully connected classification head. A decoy sequence is obtained based on the limit values of the physical parameters of the decoy flare. An intermediate discrimination model is obtained based on the initial discrimination model and the decoy sequence. The intermediate discrimination model is trained using the decoy sequence, the time-consistent radiation sequence, and the attitude sequence to obtain a target discrimination model, which can reduce the false alarm rate of the decoy flare.
[0057] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description
[0058] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application. It is obvious that the drawings described below are merely some embodiments of this application, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort. In the drawings:
[0059] Figure 1A flowchart illustrating a method for generating a decoy chaff discrimination model according to some embodiments of this application is shown;
[0060] Figure 2 A block diagram of an apparatus for generating a decoy chaff discrimination model according to some embodiments of this application is shown;
[0061] Figure 3 A schematic diagram of the structure of a device for generating a decoy chaff discrimination model according to some embodiments of this application is shown. Detailed Implementation
[0062] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0063] Furthermore, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. Numerous specific details are provided in the following description to give a thorough understanding of embodiments of this application. However, those skilled in the art will recognize that the technical solutions of this application can be practiced without one or more of the specific details, or other methods, components, apparatuses, steps, etc., can be employed. In other instances, well-known methods, apparatuses, implementations, or operations are not shown or described in detail to avoid obscuring various aspects of this application.
[0064] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.
[0065] The flowcharts shown in the accompanying drawings are merely illustrative and do not necessarily include all content and operations / steps, nor do they necessarily have to be performed in the described order. For example, some operations / steps can be broken down, while others can be combined or partially combined; therefore, the actual execution order may change depending on the specific circumstances.
[0066] Figure 1 A flowchart illustrating a method for generating a decoy chaff discrimination model according to some embodiments of this application is shown. Figure 1 As shown, a method for generating a decoy chaff discrimination model is provided, which may include the following steps 101 to 107.
[0067] In step 101, a time-consistent radiation sequence and attitude sequence are obtained based on the original image output by the infrared focal plane array inside the seeker and the attitude data of the projectile output by the inertial measurement unit.
[0068] The seeker's internal infrared focal plane array may include multiple infrared detectors for outputting raw images; the inertial measurement unit may include a three-axis gyroscope and a digital signal processor (DSP); the attitude data may be attitude quaternions. After the three-axis angular velocity is sampled by the three-axis gyroscope, the attitude quaternion can be obtained by the trapezoidal integration of the three-axis angular velocity by the DSP.
[0069] During implementation, a single crystal clock can be configured within the seeker head, and all peripheral interrupts are driven by this clock. An interrupt is triggered each time the infrared detector completes integration and readout. And record time T IR (t), and an interrupt is triggered when the three-axis gyroscope completes one sampling. And record time T IMU (k). Write these two types of interrupts into the same timestamp register so that they satisfy T. IR (t)=T IMU (k)=τ c In this way, the raw image and pose data can share a global clock.
[0070] In some embodiments, τ can be configured on the data bus side. c The sequence is subjected to monotonicity detection; if τ appears... c (t+1)-τ c (t)≠1 / F IR (F IR If the imaging frame rate is high or an arbitrary out-of-order event is detected, a placeholder frame D can be immediately inserted at the gap location. gap Simultaneously generate mask M gap =1.
[0071] In some embodiments, the gain matrix and bias matrix of the infrared focal plane array inside the seeker can be corrected based on the focal plane temperature of the infrared focal plane array inside the seeker; a linear radiation matrix is determined based on the original image, the corrected gain matrix, and the corrected bias matrix; the current attitude data is rotated and projected to obtain the displacement field corresponding to the current attitude data; the target radiation matrix corresponding to the current attitude data is determined using the displacement field, the linear radiation matrix, and a quadratic B-spline interpolation algorithm; and a time-consistent radiation sequence and attitude sequence are obtained based on the current attitude data and the target radiation matrix corresponding to the current attitude data.
[0072] It is understandable that infrared focal plane arrays undergo multi-temperature calibration using a blackbody source at the factory to obtain the gain matrix G0 and bias matrix O0. During online operation, due to fluctuations in operating temperature, the pixel response exhibits linear drift. The focal plane temperature Θ can be acquired simultaneously with each frame of data readout. t And utilize gain drift ΔG(Θ) t ) and bias drift ΔO(Θ t Incremental corrections are applied to the gain matrix G0 and the bias matrix O0 respectively. Then, the linear radiation matrix is calculated based on the original digital counting matrix of the original image, the corrected gain matrix, the corrected bias matrix, and the following formula:
[0073] R t =[G0+ΔG(Θ)] t )]°D t +[O0+ΔO(Θ t )];
[0074] in, It is a linear radiation matrix. This is the original numerical counting matrix, where ° represents element-wise multiplication, and H×W is the focal plane size. Through linear mapping, any radiation source within the scene can be represented within the thermal drift range (Θ). min ,Θ max Maintaining a constant grayscale scale within the range provides a radiation conservation premise for comparison pre-training.
[0075] Focal plane temperature Θ t The closed-loop temperature control module can update at 100Hz. The device generating the decoy / flare discrimination model can periodically read the focal plane temperature Θ. t And calculate the gain drift ΔG(Θ) t ) = A G Θ t +B G Bias drift ΔO(Θ) t ) = A O Θ t +B O Matrix A G B G A O B O All data are derived from least squares regression results of experimental data from the full-temperature range of the hot plate test. EEPROM calibration coefficients are checked daily to prevent aging.
[0076] The inertial measurement unit outputs the current attitude data Q. t Then, it can be compared with the pose data Q of the original image from the previous frame. t-1 Constructing relative rotation Let the horizontal field of view of the image be f ov The focal plane size is H×W. Using the projection operator Π(·), Q is...Δt Mapped onto the imaging plane, a sub-pixel displacement field (u) is obtained. t ,v t ), where Π(Q) Δt ,f ov Image point drift is directly calculated using the three-dimensional rotation-perspective projection equation (H,W).
[0077] The above displacement field is applied to the linear radiation matrix R. t Sub-pixel level resampling is achieved using a quadratic B-spline interpolation kernel to obtain the target radiation matrix. in This is a subpixel-level rearrangement operator. In terms of hardware implementation, a field-programmable gate array (FPGA) can interpolate each pixel scan line in a streaming pipeline manner, realizing subpixel-level resampling of the thermal field within the FPGA pipeline and eliminating optical axis jitter.
[0078] In the implementation, the radiation sequence and attitude sequence can be output using a sliding window. In some examples, the length N of the sliding window can be defined. s =16 frames, buffer queue The jitter-reduced target radiation matrix S is written sequentially using a double-ended ring structure. t With attitude data Q t .when At that time, generate It is then pushed to the thermal-motion coupling self-supervised pre-training module. At the same time, it provides two information streams, radiation sequence and attitude sequence, so that the encoder can perform cross-modal comparative learning on the same alignment reference during subsequent training, realizing high-order thermal-motion coupling.
[0079] Through the above processing, it can be ensured that the infrared radiation matrix and the motion vector field are strictly aligned at each time index, and the jitter error is compressed to the sub-pixel scale, providing a stable input for subsequent cross-modal feature coupling. This is different from the loosely coupled mode of existing technologies that process the two data separately and then synchronize them.
[0080] In step 102, the radiation sequence and attitude sequence are coupled using the first encoder in the preset network to obtain the spatiotemporal embedding vector.
[0081] The preset network may include a first encoder E φ Radiation decoding head And optical flow decoder Training from the initial network (including the initial encoder ε) φ Radiation decoding head And optical flow decoder It is obtained that the encoders in both networks can output feature vectors, the radiation decoder can output radiation matrix predictions, and the optical flow decoder can output optical flow predictions. The radiation decoder and the optical flow decoder can share all backbone parameters. Step 102 can be completed in the thermal-motion coupled self-supervised pre-training module.
[0082] In some embodiments, step 102 may include the following sub-steps:
[0083] Step 1021: Determine multiple minimum prediction units based on the radiation sequence and attitude sequence;
[0084] Step 1022: After mapping the attitude data in each minimum prediction unit to the planar vector field, concatenate it with the corresponding target radiation matrix by channel dimension to obtain multiple multi-channel tensors;
[0085] Step 1023: Train the initial network using multi-channel tensors to obtain the preset network. The initial network includes an initial encoder, a radiative decoding head, and an optical flow decoding head. The preset network includes a first encoder, a radiative decoding head, and an optical flow decoding head.
[0086] Step 1024: Input the multi-channel tensor into the preset network to obtain the first feature vector output by the first encoder, and determine the first feature vector as the spatiotemporal embedding vector.
[0087] In step 1021, it is possible to retrieve from the queue The radiation-attitude sequence corresponding to k consecutive original images is sampled using frame order indexing {(S t-k+1 Q t-k+1 ),…,(S t Q t )}, and reserve an empty slot for the next frame. Combined into the smallest prediction unit
[0088] In step 1022, the pose data Q corresponding to each frame of the original image can be processed. τ Calculate the instantaneous angular velocity Ω τ It is then transformed into a planar vector field through perspective mapping Π(·). Then, the corresponding target radiation matrix S is assembled according to the channel dimension. τ With planar vector field Obtain a two-dimensional multichannel tensor It achieves pixel-level coupling of thermal and motion information, so that the multi-channel tensor contains both the spatial distribution characteristics of the thermal field (i.e., heat) and the motion vector information that evolves over time (i.e., motion).
[0089] In step 1023, a multi-channel tensor can be input into the initial network to obtain the second feature vector output by the initial encoder, the radiation matrix prediction value output by the radiation decoder, and the optical flow prediction value output by the optical flow decoder. Based on the second feature vector corresponding to the multi-channel tensor of the current smallest prediction unit, the positive cluster center of the same group is determined. Based on the second feature vector corresponding to the multi-channel tensor of any smallest prediction unit other than the current smallest prediction unit, the negative cluster center of the cross-group is determined. The radiation matrix prediction value and optical flow prediction value corresponding to the multi-channel tensor of the current smallest prediction unit, the positive cluster center of the same group, and the negative cluster center of the cross-group are substituted into the loss function of the initial network. If the value of the loss function is lower than a preset threshold, the loss function is determined to be converged, and the preset network is obtained.
[0090] The initial encoder is identical to the first encoder except for its parameters.
[0091] It is understandable that the input multi-channel tensor X t-k+1:t Then, the predicted value of the radiation matrix is generated based on the radiation decoding head. The initial encoder is guided by L1 distance to learn the time-transfer characteristics of heat distribution along the inertial path. The goal is to enable the initial network to not only memorize static textures, but also to perceive the direction and rate of evolution of radiation intensity driven by dynamics.
[0092] The current optical flow prediction value can be generated based on the optical flow decoder. True value of optical flow F t It can be generated online using a self-supervised estimator based on three-frame difference, without the need for external annotation.
[0093] The dual decoders share encoder parameters, forcing the initial encoder ε φ Simultaneously characterizing heat and motion within the same latent space avoids the shortcomings of traditional self-supervised methods that only capture static textures for small targets while ignoring dynamic information.
[0094] The second feature vector output by the initial encoder for each frame within the same minimum prediction unit By averaging within groups, we can obtain the positive cluster centers within the same group. Then, by randomly selecting frames within the smallest prediction unit of another time period and averaging them using the second feature vector output by the initial encoder, cross-group negative cluster centers are obtained. By comparing the loss, z is reduced. τ and Euclidean distance and magnification of z τ and The distance further suppresses background static redundancy and highlights the maneuver-radiation coupling characteristics.
[0095] The loss function of the initial network can be set as:
[0096]
[0097] in, S is the predicted value of the radiation matrix; τ+1 The true value of the radiation matrix; This is the predicted optical flow value; F τ This represents the true value of optical flow; z τ This is the second feature vector; It serves as the positive cluster center of the same group; α represents the negative cluster centers across groups; α, β, γ are weighting coefficients; m is the contrast loss safety margin; [·] + This indicates the ReLU operation.
[0098] By minimizing The encoder can obtain thermal-motion coupling representation capability without manual annotation, and provides a high-fidelity feature foundation for subsequent semi-supervised discriminative training with very few annotations.
[0099] In each iteration cycle, the errors of the radiation matrix prediction value and the optical flow prediction value can be statistically analyzed. The weight coefficients of the high-loss samples with the highest error ρ are refreshed and the training list is filled back. The hard mining mechanism accelerates the model's adaptation to high-acceleration maneuvers or rapidly heated decoy scenarios, and improves the encoder's memory persistence for extreme terminal guidance conditions.
[0100] In step 1024, when the value of the above loss function is lower than a set threshold ∈ stop At this point, the initial network has been trained. The encoder parameters φ are frozen, and the multi-channel tensor is input into the preset network. The first encoder E... φ It will output the spatiotemporal embedding vector z τ =E φ (X τ This is used for subsequent semi-supervised discriminative training with minimal annotations. The spatiotemporal embedding vector simultaneously encodes the radiation field morphology, instantaneous motion vectors, and local contrastive semantics within the same vector space, forming a solid foundation for minimally labeled learning.
[0101] In step 103, the single-layer fully connected classification head in the preset network is adjusted according to the spatiotemporal embedding vector.
[0102] In step 104, an initial discrimination model is constructed based on the first encoder and the adjusted single-layer fully connected classification head.
[0103] In the implementation process, representative vectors can be extracted from the spatiotemporal embedding vectors, and binary artificial labels can be injected into the original images corresponding to the vectors. The labels include a first label representing the real target and a second label representing the decoy flares. With the encoder parameters frozen, a single-layer fully connected classification head is added to the preset network, and the single-layer fully connected classification head is adjusted using the labels. Steps 103 and 104 can be completed in a minimally labeled semi-supervised discriminative training module, which generates the initial discriminative model with minimal annotations.
[0104] In some embodiments, step 103 may include the following sub-steps:
[0105] Step 1031: Cluster the spatiotemporal embedding vectors to obtain the first spatiotemporal embedding vector, and obtain the label corresponding to the first spatiotemporal embedding vector. The label includes a first label for representing the real target and a second label for representing the decoy chaff.
[0106] Step 1032: With the parameters of the first encoder unchanged, adjust the single-layer fully connected classification head in the preset network according to the first spatiotemporal embedding vector and the label corresponding to the first spatiotemporal embedding vector.
[0107] Step 1033: Take the vectors in the spatiotemporal embedding vector other than the first spatiotemporal embedding vector as the second spatiotemporal embedding vector, and perform weak augmentation processing on the original image corresponding to the second spatiotemporal embedding vector to obtain a set of pseudo-label candidates.
[0108] Step 1034: Select weak augmented images and their corresponding pseudo-labels from the pseudo-label candidate set according to a preset decreasing threshold sequence; perform strong augmentation processing on the weak augmented images to obtain strong augmented images.
[0109] Step 1035: With the parameters of the first encoder unchanged, adjust the single-layer fully connected classification head in the preset network according to the pseudo-labels corresponding to the weakly augmented image and the strongly augmented image.
[0110] It should be noted that, to compensate for the scarcity of annotations in real tactical scenarios, the industry generally adopts a strategy of large-scale synthetic sequence pre-training combined with a small amount of field data transfer. However, synthetic sequences have significant domain differences from real-world conditions in terms of spectral distribution, wake morphology, and ballistic noise, making it difficult for the model to maintain the discrimination boundary during the training period within the actual combat weather window. Furthermore, limitations in computing power and safety gaps lead to the deep pruning and solidification of the backbone network, allowing only online fine-tuning of the terminal classification layer. This local update lacks the ability to adapt to sudden changes in the radiation surface and motion coupling of the chaff.
[0111] In step 1031, the entire spatiotemporal embedding vector sequence can be targeted. Calculate the pairwise Euclidean distance matrix for all spatiotemporal embedding vectors, then perform K-Medoids clustering, setting the number of cluster centers to be... This achieves global coverage across three dimensions: distance, weather type, and maneuvering attitude.
[0112] Lock the original image frame index corresponding to each cluster center The field technicians only reviewed these raw image frames and assigned them binary labels. τ ∈{1,0}, where 1 represents the real target and 0 represents the decoy chaff.
[0113] In step 1032, the first encoder is frozen and a single-layer fully connected classification head is trained. In this sub-step, all parameters φ of the first encoder are frozen, and a single-layer fully connected classification head is added to the preset network. Its input dimension is the same as its embedding dimension, and its output dimension is 2.
[0114] The spatiotemporal embedding vector e for all labeled original image frames i With the corresponding label y i Construct a weighted cross-entropy objective:
[0115]
[0116] Among them, J init N is the initial loss function; L The sample size is manually labeled; Based on category frequency Adaptive weights; These represent the predicted probabilities of the classification head for the real target class and the decoy / chaff class, respectively.
[0117] By minimizing J init The first version of the decision hyperplane is given under sparse label conditions, laying the baseline for subsequent confidence assessment.
[0118] In step 1033, pseudo-labels are generated through weak augmentation. In this sub-step, the weights of the single-layer fully connected classification head are kept fixed, and two weak augmentation processes, brightness jitter and fine-grained rotation, are applied to the unlabeled original image frame to obtain the weakly augmented image A. w (·), then calculate the output probability. The highest output probability Record the confidence level and the triplet Included in the pseudo-label candidate set U1, where
[0119] Since weak augmentation does not destroy the geometric properties of the scene, the confidence level can truly reflect the reliability of the classification head's decision on the original features.
[0120] In step 1034, a decreasing threshold sequence {θ} is set for the pseudo-label candidate set U1. (r)}, and in each iteration, based on the current high confidence ratio σ (r) Adjusting the threshold θ (r+1) =max(θ) floor ,θ (r) -λσ (r) ), where θ floor This is the lower bound that cannot be lowered; λ is the attenuation factor. Only when c... τ ≥θ (r+1) Furthermore, the weakly augmented image x is only applied when the ratio of the two types of samples remains in the range [ρ, 1-ρ]. τ With pseudo-tags Add to training buffer pool In this way, we can control the influx of mislabeled categories, stabilize category balance, and prevent the encoder from forgetting scarce categories.
[0121] In step 1035, the training buffer pool is... Each frame of the image is additionally subjected to occlusion, random cropping, and temporal disordering to form a strongly augmented image A. s (·), ensuring that the predicted distribution of the same frame image remains consistent across both weak and strong views. Kullback–Leibler divergence is added to the loss as a regularization term, enhancing the model's robustness to occlusion, flicker, and sudden viewpoint changes based on significant deformation, and preventing weight shifts driven by pseudo-label noise through consistency constraints.
[0122] In step 1036, after training one round using the pseudo-labels corresponding to the strongly augmented and weakly augmented images, the false positive rate ε is calculated on the validation set. (r) If it reaches ε (r) >ε (r-1) Immediately roll back the classifier weights to the last snapshot, and simultaneously increase the threshold to θ. (r+1) =θ (r) +2λ.
[0123] During implementation, the last layer of the first encoder can also be unfrozen and fine-tuned in conjunction with it. When two consecutive rounds satisfy |ε (r) -ε (r-1) |<δ stop (where δ) stop When the preset stop threshold is reached, the parameter φ of the last layer of the first encoder is unlocked. tail Parameters of a single-layer fully connected classification head A low learning rate is used, along with WeightDecay fine-tuning, until the validation loss shows no further decrease. Finally, the parameters are frozen to obtain the initial discriminant model. Where φ * , This represents the globally optimal parameters. The initial discrimination model combines thermal-motion priors with sparse supervision adaptation capabilities, providing a stable and highly confident base for discriminating between real targets and decoys in subsequent adversarial inference stages.
[0124] By performing minimal binary annotations on representative original images, freezing the encoder, training a single-layer fully connected classification head, and then progressively expanding the annotation domain with weak augmented pseudo-labels and strong augmented consistency regularization, the limitation of difficulty in large-scale annotation in real tactical scenarios is overcome.
[0125] In step 105, the decoy sequence is obtained based on the limit values of the physical parameters of the decoy flare, wherein the decoy sequence is used to characterize the radiation sequence and attitude sequence of the decoy flare.
[0126] In step 106, an intermediate discrimination model is obtained based on the initial discrimination model and the decoy sequence.
[0127] Steps 105 and 106 can be completed in the adversarial hard example incremental training module, which uses the decoy sequence to generate an adversarially enhanced intermediate discriminant model.
[0128] The physical parameters of the decoy flares may include the initial temperature T0 and the temperature decay constant k. T Effective area of the burning surface A f Injection initial velocity v e , charge mass m f Tail evaporation rate η b Particle distribution scale σ s Spectral band ratio emissivity ε, etc.
[0129] In the implementation process, the infrared radiation curve, tail flame spectrum, and trajectory of the magnesium chaff can be obtained from the test range, and the relevant data of the above eight physical parameters can be statistically analyzed. Then, based on the statistically obtained limit values, upper and lower bound vectors can be constructed.
[0130] L≤p=[T0,k T A f ,v e ,m f ,η b ,σ s ,ε] T ≤U
[0131] Where p is the vector of physical parameters to be generated. The inequality ensures that subsequent samples maintain tactical realism and do not violate the law of conservation of heat.
[0132] In some embodiments, step 105 may include the following sub-steps:
[0133] Step 1051: Construct a parameter generator based on the limit values of the physical parameters of the decoy flares;
[0134] Step 1052: Obtain the decoy sequence based on the parameters output by the parameter generator.
[0135] Define parameter generator Where d z Denotes the Gaussian latent code dimension, 6-dimensional prior vector. The distance d between the current true target and the seeker t radial velocity and attitude Euler angles (α) t ,β t ,γ t ), Remaining flight time ξ t composition.
[0136] Each round samples from a normal distribution. through Output And by combining the upper and lower bound vectors, p is forcibly bound within the physical boundary.
[0137] In step 1052, the parameters p output by the parameter generator and the attitude prior can be fed into a fully differentiable infrared renderer. Render output decoy sequence The resolution and timestamps of the decoy sequence are perfectly aligned with the real data stream.
[0138] The infrared renderer may include: a radiative transfer submodule based on energy conservation, which will transfer T0,k T A f ,m f ε is transformed into a time-grayscale surface; the ballistic submodule based on Newtonian aerodynamics utilizes v e ,σ s Calculate the three-dimensional trajectory; a wake submodule based on the volatile decay model, with η b Modulation attenuation radius.
[0139] In some embodiments, step 106 may include the following sub-steps:
[0140] Step 1061: Unfreeze the penultimate layer of the first encoder in the initial discriminant model and the single-layer fully connected classification head in the initial discriminant model to construct an adversarial discriminator;
[0141] Step 1062: Determine the confusion probability of the decoy sequence based on the adversarial discriminator;
[0142] Step 1063: Train the adversarial discriminator based on the confusion probability. After the first preset condition is met, continue to train the adversarial discriminator using the decoy sequence, the time-consistent radiation sequence, and the attitude sequence. After the second preset condition is met, construct an intermediate discriminant model based on the first encoder and the trained adversarial discriminator.
[0143] In step 1061, the decoy sequence can be input into the first encoder of the initial discrimination model to obtain the spatiotemporal embedding vector. Unfreeze the parameters φ of the penultimate layer of the first encoder in the initial discriminant model tail The parameters of the single-layer fully connected classifier head in the initial discriminant model Combining adversarial discriminants
[0144] In step 1062, for each spatiotemporal embedding vector from step 1061, the label y=1 of the true target is taken, and the confusion probability is calculated. The closer the confusion probability is to 1, the more the decoy flare resembles the real target.
[0145] In step 1063, the parameter generator and adversarial discriminator are optimized using a game theory approach. This involves joint optimization of the adversarial objectives: maximizing confusion in the parameter generator and minimizing confusion in the adversarial discriminator while maintaining physical plausibility. This is achieved through the following joint training with the adversarial objectives:
[0146]
[0147] Where, λ phys The first term represents the physical plausibility penalty coefficient; the second term measures the degree to which the model is misled by the decoy; the third term constrains p to return to the legal interval during gradient backpropagation, preventing extreme parameters that violate thermodynamic laws. The game iterates until… The distance from the true target cluster boundary is reduced to a preset threshold δ adv (i.e., the first preset condition).
[0148] It is understandable that the temporally consistent radiation and attitude sequences are real sequences, and the decoy sequences can be mixed with the real sequences at a 1:1 ratio to continuously train the thawed parameter set. This continues until the false alarm rate on the validation set decreases and stabilizes (i.e., the second precondition). This process expands the decision boundary towards the safe side, reducing the false alarm risk caused by highly similar decoy chaff, and ultimately outputs an adversarially enhanced intermediate discriminant model.
[0149] In step 107, the intermediate discrimination model is trained using the decoy sequence, the time-consistent radiation sequence, and the attitude sequence to obtain the target discrimination model.
[0150] Understandably, the decoy sequence can be inserted into the real sequence to obtain a synthetic sequence, which can then be used to train the intermediate discrimination model again, ultimately yielding the target discrimination model. After obtaining the target discrimination model, it can be used to distinguish decoy flares during the terminal guidance phase of infrared homing. Step 107 can be completed in the difficult example insertion and re-pre-training module.
[0151] In some embodiments, step 107 may include the following sub-steps:
[0152] Step 1071: Insert the decoy sequence into the radiation sequence and attitude sequence that are in the same time sequence according to the time axis to obtain the synthetic sequence;
[0153] Step 1072: Adjust the last two layers of the first encoder in the intermediate discriminant model according to the synthesized sequence to obtain the second encoder;
[0154] Step 1073: Adjust the single-layer fully connected classification head of the adversarial discriminator in the second encoder and intermediate discriminant model according to the original image corresponding to the first spatiotemporal embedding vector;
[0155] Step 1074: Construct a target discrimination model based on the adjusted second encoder and the adjusted adversarial discriminator.
[0156] In step 1071, a time index set T of the real sequence can be constructed. real ={t1,t2,…,t M}, where t i = i·Δt, where Δt is a fixed frame period; Read the decoy sequence Search for the closest one in sequence Original empty frame index If the position is already occupied, search forward along the timeline until an unfilled slot is found; insert the difficult-to-insert frame at the desired location. Then, the synthetic sequence is regenerated. Where y t Label the target with 0-1 tags (1 represents the real target, 0 represents the decoy / flare); record the interpolation label matrix M∈{0,1} M , used for explicit grouping of real target-decoy frames.
[0157] In step 1072, a partial thawing thermal-motor contrast retraining is performed. In this sub-step, the first L-2 layers of the first encoder are frozen, and only the parameter sets of the last two layers are unlocked. exist The thermal-motion dual-contrast pre-task is executed in parallel, specifically: the temperature surface prediction sub-task uses the target radiation matrix S. t The supervised network captures the temporal patterns of radiation; the ballistic vector prediction subtask uses attitude data Q.t This encourages the network to retain prior knowledge of motion dynamics; an additional two-positive-two-negative hard-example separation regularization is added: positive-negative contrast pairs are constructed for real target-decoy frames at the same time, allowing the network to actively widen the distance between them in the local embedding space. The overall multi-target loss is:
[0158]
[0159] in: For the predicted output of re-pre-training; These are the embedding vectors of the same frame before and after retraining; Represents the set of indexes of the actual target frames. Represents the set of decoy frame indices; Embed the centroid of the real target; α, β, γ, κ are hyperparameters.
[0160] The first two items maintain thermo-kinetic consistency, the third item protects old memories, and the fourth item deliberately punishes difficult examples to get closer to the real goal.
[0161] The average embedding distance increment ΔD of the decoy can be calculated in real time using the following formula, and an exponential moving average can be applied to ΔD to eliminate random oscillations:
[0162]
[0163] When ΔD≥δ gap (where δ) gap (for the preset separation threshold) and If a monotonically decreasing inflection point is observed in the most recent K mini-batches, immediately freeze all trainable weights to obtain the second encoder.
[0164] The separation effect between the decoy flares and the real targets is quantified in real time. Training stops immediately when the separation degree reaches a threshold to avoid overfitting and unnecessary iteration.
[0165] In step 1073, the set of original images (or representative frames) corresponding to the first spatiotemporal embedding vector is obtained. Using sets Weak-strong augmentation consistency training is performed, specifically: weak augmentation (brightness jitter, fine-angle rotation) ensures the stability of the base distribution; strong augmentation (occlusion, random pruning, temporal disorder) forces the single-layer fully connected classifier head of the adversarial discriminator in the intermediate discriminant model to maintain consistent predictions in scenarios with large deformations; only the two-layer classifier head is optimized. We employ weighted cross-entropy plus KL consistency regularization to force consistent prediction distributions under strong and weak views without modifying the encoder body; we monitor the false alarm rate ε on the validation set. val If no significant decrease is observed in 20 consecutive small batches, the fine-tuning process ends.
[0166] After fine-tuning, the remaining difficult example sequences can be scanned and progressively stacked using the new model. Full scan of the data pool: Record the confidence level corresponding to all decoy sequences. c t Samples ∈ (0.45, 0.55) are grouped into a set of difficult examples as a sequence of difficult examples. right Reverse lookup to generate parameters And store in the parameter pool Will Add to the training buffer to provide a sharper sample base for the next iteration. Repeat steps 1072 and 1073, determining the end of training using the following two criteria: a decrease in the proportion of remaining hard example sequences. (where M is the total number of samples in a single full scan, ε) stop For two consecutive rounds (with a preset stop false positive rate); and the validation set false positive rate is stable: (where η) flat To verify the flatness threshold of the false alarm rate (the value of the verification set), both conditions must be met, i.e., the encoder version must be fixed. With the final classification head
[0167] In step 1074, according to the adjusted encoder Construct a target discrimination model using the adjusted adversarial discriminator.
[0168] The target discrimination model can be validated using a closed-loop terminal guidance simulation chain. Specifically, this can be achieved by: [The target discrimination model is then...] A six-DOF terminal guidance simulation chain was embedded, maintaining synchronization between the embedded-decision module and the guidance loop at a frequency of 200Hz; a strong disturbance scenario matrix Ω was constructed, including multiple ballistic incident angles, rapid attitude rolls, and dual thermal-cloud cover. sim Simulations were run scene by scene, and the latency τ was captured in real time. cap ε, the false capture rate of the bait sim 1. Miss distance Δr; compare and evaluate with the control model, if ε sim ≤0.4ε baseline And τ cap If the increase is less than 3%, the technical gain is deemed to meet the engineering indicators; output the test report and software version label, and complete the entire chain delivery.
[0169] This application employs a multi-channel tensor-based dual-task self-supervised framework based on thermal field-motion field channel splicing. The same encoder simultaneously undertakes future radiation prediction and current optical flow estimation, with two decoders sharing all backbone parameters. During training, the frame-level feature vectors output by the encoder are averaged within a sliding window to form positive cluster centers within the same group. These are then compared with randomly selected negative cluster centers across groups to form contrast pairs. A joint loss function allows radiation and dynamics to be jointly represented within a unified latent space. After self-supervision, only a small number of binary annotations are applied to representative frames. The encoder is frozen, and a single-layer fully connected classification head is trained. The annotation domain is then progressively expanded using weak augmented pseudo-labels and strong augmented consistency regularization. Furthermore, a decoy sequence is generated through a fully differentiable physical rendering pipeline, participating in adversarial games alongside the real sequence. After convergence, the decoy sequence is inserted back into the timeline of the real sequence, the tail layer parameters are locally unfrozen, and a two-positive, two-negative hard-example separation regularization is added to continuously increase the distance between the real target and the decoy in the embedding space. The integrated training chain achieves high-order coupling of radiation and motion, reinforcement of adversarial examples, and sparse supervision adaptation at the encoder level. It has significant structural differences from the existing methods that separate spectral thresholds, motion filtering, and empirical rules without mutual feedback. It effectively improves the accuracy of decoy chaff detection and reduces the false alarm rate.
[0170] After obtaining the target model, the method for generating the decoy chaff discrimination model may further include step 108, which involves compressing and deploying the target discrimination model. In some embodiments, step 108 may include the following sub-steps:
[0171] Step 1081: Decompose the three-dimensional convolution kernel of the target discrimination model into an intra-frame two-dimensional convolution kernel and an inter-frame one-dimensional convolution kernel;
[0172] Step 1082: Based on the contribution of the output channel of the target discrimination model to the backpropagation gradient corresponding to the difficult example sequence, the output channel is pruned, wherein the difficult example sequence is a decoy sequence with moderate confidence.
[0173] Step 1083: Detect the inference delay time and anti-decoy accuracy of the target discrimination model;
[0174] Step 1084: Deploy the target discrimination model if the inference delay time is less than the preset time and the anti-decoy accuracy is greater than the preset accuracy.
[0175] In step 1081, a hierarchical time consumption analysis is performed on the target discrimination model. In this sub-step, the hardware counters of the FPGA-DSP collaborative environment can be collected to record the measured time consumption vector l = [l1, l2, ..., l] of each layer of the target discrimination model. J ] TWhere J is the number of network layers. The temporal grouping and reconstruction of the 3D convolutional kernels of the object discrimination model is performed, while keeping the weight values unchanged, to convert the original 3D convolutional kernels... (where C) in C represents the number of channels in the input feature map. out The number of output channels (also the number of convolutional kernels in this layer, where k is the temporal kernel length) is decomposed into intra-frame depth 2D convolutional kernels. with inter-frame one-dimensional convolution kernel And independently set up thermal field passages With light flow path The output feature Y at the same time index t after transformation t Written as:
[0176]
[0177] in, For the current frame input of channel c, The cached features are concatenated from the path. * indicates two-dimensional convolution, and ⊙ indicates element-wise multiplication and accumulation.
[0178] By decomposing the original 3D convolution kernel memory access into two accesses—spatial local and temporal sparse—the off-chip bandwidth requirement per frame is reduced, while the parallel operator libraries of FPGA and DSP are reused.
[0179] In step 1082, the gradient-sensitive channel contribution is evaluated. In this sub-step, the backpropagation gradient corresponding to the hard example sequence is evaluated. Calculate the contribution of each output channel c across the entire training set:
[0180]
[0181] in, Let Ω be the spatial index of the feature map, representing the sample set. Let be the activation value of the d-th sample at position p in output channel c.
[0182] According to Γ c Sort by size from smallest to largest, and output channels with low contribution are included in the pruning candidate set. Ensure that high-contribution output channels are fully preserved and avoid accidental deletion of thermo-optical modes that are crucial for decoy chaff detection.
[0183] Candidate sets can be processed in 10% increments. Perform layer-by-layer pruning: Remove a subset of candidate channels in each round. The FPGA logic resource table is statically recompiled, allowing for immediate shrinking of the on-chip SRAM and convolutional array; only the parameters of the last two layers are fine-tuned three times in a small batch until the anti-decoy accuracy is verified through the validation set. Keep
[0184] Repeat substeps 1081-1082 to measure the end-to-end inference delay τ in real time. e2e If τ e2e ≤4ms and (η min If the percentage is fixed at 97%, then stop the iteration and solidify the lightweight model M. lite If continued cutting leads to Roll back to the previous version's parameters and structure to ensure real-time performance and robustness.
[0185] In step 1083, the lightweight model M can be... lite The convolution kernel, bias, and normalization scale are mapped to 8-bit integers and then mapped to the multiplier array of the FPGA-DSP collaborative platform in a fixed-point format (Q2.6); a three-stage pipeline buffer is constructed simultaneously: input buffer → depthwise convolution → group convolution.
[0186] It's worth noting that, in terms of model deployment, the industry standard practice is to split the model's 3D convolutional kernel into spatial 2D convolution and temporal recursive updates, supplemented by 8-bit fixed-point quantization, on-chip dual-end caching, and Ping-Pong scheduling to reduce external bandwidth overhead. This approach weakens the spatial-temporal correlation representation across frames. When high-speed cloud veils pass by or background radiation fluctuates rapidly, the recursive state becomes delayed, further amplifying the risk of terminal lockout. Overall, current technologies still struggle to consistently reduce the false alarm rate of decoys in the terminal guidance phase to a more stringent numerical target under the hard constraint of millisecond-level inference latency.
[0187] After performing hierarchical time consumption analysis on the target discrimination model, the cross-frame 3D convolution kernel is decomposed into an intra-frame 2D convolution kernel and a one-dimensional convolution kernel for inter-frame grouping. The contribution of the backpropagation gradient corresponding to the hard example sequence is statistically analyzed in the entire training set, and low contribution channels are pruned in an incremental manner. Then, the convolution kernel and bias are quantized into 8-bit fixed-point parameters and mapped to the three-stage pipeline of the FPGA-DSP collaborative platform to achieve end-to-end millisecond-level inference.
[0188] After the target model is deployed, the method for generating the decoy chaff discrimination model may further include step 109, which involves on-orbit micro-updating and maintaining the target discrimination model. In some embodiments, step 109 may include the following sub-steps:
[0189] Step 1091: Extract the ten most recent original images and their corresponding optical flow fields in chronological order during missile flight;
[0190] Step 1092: Calculate the drift intensity of the current window based on the extracted original image and the corresponding optical flow field;
[0191] Step 1093: Construct a weighted consistency loss function based on the drift intensity, and apply the weighted consistency loss function only to the statistical layer;
[0192] Step 1094: Perform a finite number of micro-updates only on the mean and variance of each layer of BatchNorm;
[0193] Step 1095: Adjust the classification threshold by sliding according to the F1 score curve of the window;
[0194] Step 1096: If the target is continuously verified to be hit and there are no new false alarms, then the update is fixed; otherwise, the version is rolled back to the previous version and the update record is transmitted back to the ground station through the telemetry channel.
[0195] In step 1091, the ten most recent frames of original image I can be extracted from the seeker's circular buffer in chronological order. t With the corresponding optical flow field To ensure that the image and optical flow are aligned at the same sampling timestamp, a complete thermal-motion co-observation is prepared for this round of micro-updates.
[0196] In step 1092, the ten original image-optical flow field pairs can be... Injecting lightweight model M lite Then, hot branches are extracted and embedded respectively. With motion branch embedding (b is the pixel index). The residual difference is calculated pixel by pixel and accumulated over time to obtain the drift intensity of the current window:
[0197]
[0198] Where Ω represents the set of pixel indices for a single frame. A larger D value indicates a more significant drift in the center band of the atmospheric transmission window, leading to thermal-motion coupling mismatch.
[0199] In step 1093, the weight ω = min(κD,1) (κ is a proportionality constant) is calculated based on the drift intensity D, and the weighted consistency loss function is constructed as follows:
[0200]
[0201] Only It is applied to the statistical layer to avoid disrupting the already compressed convolutional kernel weights.
[0202] In step 1094, all convolutional kernels and biases are frozen, and only the mean vector μ of each BatchNorm layer of the target discrimination model is processed. l With variance vector Perform three rounds of micro-updates: using a learning rate η bn (<10 -3 Gradient descent correction Immediately after each update, re-normalize and reactivate to ensure the new distribution does not exhibit explosive peaks; constraints Avoid excessive drift.
[0203] In step 1095, the set of positive and negative class output scores for ten frames after the micro-update is collected. Calculate the F1 score curve F1(θ) based on the separation threshold θ. Slide the discrimination threshold of the fully connected layer to the θ corresponding to the peak F1 score. ★ This offsets the global bias introduced by changes in cloud thickness and humidity.
[0204] In step 1096, in the following two frames I t′ Run the new threshold θ ★ If the target is hit continuously without any new false alarms, then the status is fixed. Otherwise, immediately roll back to the previous version of the parameters to avoid accidental updates affecting the safety of the guidance closed loop.
[0205] After each three successful solidifications, the corresponding timestamps and drift metrics {t, D} are combined to form a trajectory sequence. The data is written to the encrypted telemetry channel and transmitted back to the ground station for subsequent offline batch simulation reproduction, and the model repository is updated synchronously so that subsequent firmware versions inherit the latest statistical priors.
[0206] The above scheme enables the construction of a real-time discrimination model that simultaneously characterizes the coupling relationship between thermal radiation distribution and maneuver dynamics during the terminal guidance phase of infrared homing, under conditions of minimal annotation and within a low-power millisecond-level inference framework limited by the flight platform. This model can accurately distinguish between real targets and decoy flares even when facing multi-spectral, complex atmospheric windows and high-intensity magnesium bomb jamming scenarios, and keeps the false alarm rate stable within a predetermined threshold throughout the entire process.
[0207] The following describes an embodiment of the apparatus described in this application, which can be used to execute the method for generating the decoy / flare discrimination model in the above embodiments of this application. For details not disclosed in the apparatus embodiments of this application, please refer to the embodiments of the method for generating the decoy / flare discrimination model described above in this application.
[0208] See Figure 2 The diagram shows a block diagram of the device for generating the decoy chaff discrimination model in an embodiment of this application.
[0209] like Figure 2As shown, the decoy chaff discrimination model generation device of this application embodiment includes: a sequence acquisition module 201, a sequence coupling module 202, an initial model construction module 203, an intermediate model construction module 204, and a target model construction module 205. The sequence acquisition module 201 is used to obtain a time-consistent radiation sequence and attitude sequence based on the original image output by the infrared focal plane array inside the seeker and the attitude data of the projectile output by the inertial measurement unit. The sequence coupling module 202 is used to couple the radiation sequence and attitude sequence using a first encoder in a preset network to obtain a spatiotemporal embedding vector. The initial model construction module 203 is used to generate a target model based on the spatiotemporal embedding vector. The input vector is adjusted in the preset network to form a single-layer fully connected classifier head; the initial model building module 203 is also used to build an initial discrimination model based on the first encoder and the adjusted single-layer fully connected classifier head; the intermediate model building module 204 is used to obtain the decoy sequence based on the limit values of the physical parameters of the decoy flares, wherein the decoy sequence is used to characterize the radiation sequence and attitude sequence of the decoy flares; the intermediate model building module 204 is also used to obtain an intermediate discrimination model based on the initial discrimination model and the decoy sequence; the target model building module 205 is used to train the intermediate discrimination model using the decoy sequence, the time-consistent radiation sequence and attitude sequence to obtain the target discrimination model.
[0210] Based on the same inventive concept, embodiments of this application also provide a device for generating a decoy chaff discrimination model, see reference. Figure 3 The diagram shows a schematic of the structure of the device for generating the decoy chaff discrimination model in an embodiment of this application. The device for generating the decoy chaff discrimination model includes one or more memories 304, one or more processors 302, and at least one computer program (computer program instructions) stored in the memory 304 and executable on the processor 302. When the processor 1202 executes the computer program, it implements the method described above.
[0211] Among them, Figure 3 In this document, a bus architecture (represented by bus 300) is used. Bus 300 may include any number of interconnected buses and bridges, linking various circuits including one or more processors represented by processor 302 and memory represented by memory 304. Bus 300 may also link various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and therefore will not be described further herein. Bus interface 305 provides an interface between bus 300 and receiver 301 and transmitter 303. Receiver 301 and transmitter 303 may be the same element, i.e., a transceiver, providing a unit for communicating with various other devices over a transmission medium. Processor 302 is responsible for managing bus 300 and general processing, while memory 304 can be used to store data used by processor 302 during operation.
[0212] Based on the same inventive concept, embodiments of this application provide a computer-readable storage medium storing computer program instructions, which, when executed by a processor, cause the processor to perform the steps of the method described above.
[0213] Based on the same inventive concept, embodiments of this application provide a computer program product, including a computer program, which, when executed by a processor, causes the processor to perform the steps of the method described above.
[0214] The functions described herein may be implemented in hardware, software executed by a processor, firmware, or any combination thereof. If implemented in software executed by a processor, the functions may be stored as one or more instructions or codes on or transmitted via a computer-readable medium. Other examples and embodiments are within the scope and spirit of this application and the appended claims. For example, due to the nature of software, the functions described above may be implemented using software executed by a processor, hardware, firmware, hardwired, or any combination thereof. Furthermore, the functional units may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit.
[0215] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units can be a logical functional division, and in actual implementation, there may be other division methods. For instance, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.
[0216] The units described as separate components may or may not be physically separate. Similarly, the components of the control device may or may not be physical units; they may be located in one place or distributed across multiple units. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.
[0217] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing computer program instructions, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.
[0218] The above description is merely an embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. A method for generating a decoy chaff discrimination model, characterized in that, include: Based on the original image output by the infrared focal plane array inside the seeker and the attitude data of the projectile output by the inertial measurement unit, a time-consistent radiation sequence and attitude sequence are obtained. The radiation sequence and the attitude sequence are coupled using the first encoder in the preset network to obtain a spatiotemporal embedding vector; Adjust the single-layer fully connected classification head in the preset network according to the spatiotemporal embedding vector; An initial discrimination model is constructed based on the first encoder and the adjusted single-layer fully connected classification head; Based on the limit values of the physical parameters of the decoy flares, a decoy sequence is obtained, wherein the decoy sequence is used to characterize the radiation sequence and attitude sequence of the decoy flares. Based on the initial discrimination model and the decoy sequence, an intermediate discrimination model is obtained; The intermediate discrimination model is trained using the decoy sequence, the time-consistent radiation sequence, and the attitude sequence to obtain the target discrimination model.
2. The method for generating the decoy chaff discrimination model according to claim 1, characterized in that, The infrared focal plane array within the seeker head includes an infrared detector, and the inertial measurement unit includes a three-axis gyroscope. The interrupt triggering time of the infrared detector upon completion of integration readout is the same as the interrupt triggering time of the three-axis gyroscope upon completion of sampling. Based on the original image output by the infrared focal plane array within the seeker head and the attitude data of the projectile output by the inertial measurement unit, a time-consistent radiation sequence and attitude sequence are obtained, including: The gain matrix and bias matrix of the infrared focal plane array inside the seeker are corrected based on the focal plane temperature of the infrared focal plane array inside the seeker. The linear radiation matrix is determined based on the original image, the corrected gain matrix, and the corrected bias matrix; The current attitude data is rotated and projected to obtain the displacement field corresponding to the current attitude data. The target radiation matrix corresponding to the current attitude data is determined using the displacement field, the linear radiation matrix, and a quadratic B-spline interpolation algorithm. Based on the current attitude data and the target radiation matrix corresponding to the current attitude data, a time-consistent radiation sequence and attitude sequence are obtained.
3. The method for generating the decoy chaff discrimination model according to claim 2, characterized in that, The step of coupling the radiation sequence and the attitude sequence using a first encoder in a preset network to obtain a spatiotemporal embedding vector includes: Based on the radiation sequence and attitude sequence, multiple minimum prediction units are determined; After mapping the attitude data in each of the minimum prediction units to a planar vector field, the data is concatenated with the corresponding target radiation matrix by channel dimension to obtain multiple multi-channel tensors. The initial network is trained using the multi-channel tensor to obtain a preset network, wherein the initial network includes an initial encoder, a radiative decoding head, and an optical flow decoding head, and the preset network includes the first encoder, the radiative decoding head, and the optical flow decoding head; The multi-channel tensor is input into the preset network to obtain the first feature vector output by the first encoder, and the first feature vector is determined as the spatiotemporal embedding vector.
4. The method for generating the decoy chaff discrimination model according to claim 3, characterized in that, The step of training an initial network using the multi-channel tensor to obtain a preset network includes: The multi-channel tensor is input into the initial network to obtain the second feature vector output by the initial encoder, the radiation matrix prediction value output by the radiation decoder, and the optical flow prediction value output by the optical flow decoder. Based on the second eigenvector corresponding to the multichannel tensor of the current smallest prediction unit, determine the positive cluster center of the same group; Based on the second feature vector corresponding to the multichannel tensor of any of the plurality of minimum prediction units other than the current minimum prediction unit, determine the cross-group negative clustering center; Substitute the predicted radiation matrix and optical flow values corresponding to the multi-channel tensor of the current smallest prediction unit, the positive cluster centers of the same group, and the negative cluster centers of the cross-group into the loss function of the initial network. If the value of the loss function is lower than a preset threshold, the loss function is determined to have converged, and the preset network is obtained.
5. The method for generating the decoy chaff discrimination model according to claim 4, characterized in that, The step of adjusting the single-layer fully connected classification head in the preset network according to the spatiotemporal embedding vector includes: Clustering and filtering are performed on the spatiotemporal embedding vectors to obtain a first spatiotemporal embedding vector, and the label corresponding to the first spatiotemporal embedding vector is obtained, wherein the label includes a first label for characterizing the real target and a second label for characterizing the decoy flares. With the parameters of the first encoder unchanged, the single-layer fully connected classification head in the preset network is adjusted according to the first spatiotemporal embedding vector and the label corresponding to the first spatiotemporal embedding vector. The vectors other than the first spatiotemporal embedding vector in the spatiotemporal embedding vector are used as the second spatiotemporal embedding vector, and the original image corresponding to the second spatiotemporal embedding vector is subjected to weak augmentation processing to obtain a set of pseudo-label candidates. Weak augmented images and their corresponding pseudo-labels are selected from the candidate set of pseudo-labels based on a preset decreasing threshold sequence. The weakly augmented image is subjected to strong augmentation processing to obtain a strongly augmented image; With the parameters of the first encoder unchanged, the single-layer fully connected classification head in the preset network is adjusted according to the pseudo-labels corresponding to the weakly augmented image and the strongly augmented image.
6. The method for generating the decoy chaff discrimination model according to claim 5, characterized in that, The process of obtaining the decoy sequence based on the extreme values of the physical parameters of the decoy flares includes: A parameter generator is constructed based on the limit values of the physical parameters of the decoy flares; The decoy sequence is obtained based on the parameters output by the parameter generator.
7. The method for generating the decoy chaff discrimination model according to claim 6, characterized in that, The step of obtaining an intermediate discrimination model based on the initial discrimination model and the decoy sequence includes: Unfreeze the penultimate layer of the first encoder in the initial discriminant model and the single-layer fully connected classification head in the initial discriminant model to construct an adversarial discriminator; The confusion probability of the decoy sequence is determined based on the adversarial discriminator. The adversarial discriminator is trained based on the confusion probability. After the first preset condition is met, the adversarial discriminator is further trained using the decoy sequence, the time-consistent radiation sequence, and the attitude sequence. After the second preset condition is met, the intermediate discriminant model is constructed based on the first encoder and the trained adversarial discriminator.
8. The method for generating the decoy chaff discrimination model according to claim 7, characterized in that, The step of training the intermediate discrimination model using the decoy sequence, the time-consistent radiation sequence, and the attitude sequence to obtain the target discrimination model includes: The decoy sequence is inserted into the time-consistent radiation sequence and attitude sequence along the time axis to obtain a synthetic sequence; The last two layers of the first encoder in the intermediate discriminant model are adjusted according to the synthesized sequence to obtain the second encoder; Adjust the single-layer fully connected classification head of the second encoder and the adversarial discriminator in the intermediate discriminant model according to the original image corresponding to the first spatiotemporal embedding vector; The target discrimination model is constructed based on the adjusted second encoder and the adjusted adversarial discriminator.
9. The method for generating the decoy chaff discrimination model according to claim 1, characterized in that, Also includes: The three-dimensional convolutional kernel of the target discrimination model is decomposed into an intra-frame two-dimensional convolutional kernel and an inter-frame one-dimensional convolutional kernel; Based on the contribution of the output channel of the target discrimination model to the backpropagation gradient corresponding to the difficult example sequence, the output channel is pruned, wherein the difficult example sequence is a decoy sequence with moderate confidence. The inference delay time and anti-decoy accuracy of the target discrimination model are tested. The target discrimination model is deployed when the inference delay time is less than a preset time and the anti-decoy accuracy is greater than a preset accuracy.
10. A device for generating a decoy chaff discrimination model, comprising a processor and a memory, characterized in that, The memory stores computer program instructions that can be executed by the processor, and when the processor executes the computer program instructions, it implements the steps of the method as described in any one of claims 1 to 9.
Citation Information
Patent Citations
Foil type infrared surface source bait dynamic diffusion characteristic test acquisition method
CN114046696A
Weak and small target detection method and device capable of resisting infrared decoy interference
CN114359264A
Interference control method and system based on surface source infrared decoy image
CN118913023A
Model correction method and device, electronic equipment and storage medium
CN119004959A
IR decoy, and method for adjusting infrared radiation amount from the same
JP2013148271A
Cited By
Track panel joint damage state intelligent diagnosis method and system based on vehicle-mounted CCD camera
CN121582598A