Method and device for identifying target in infrared image, equipment and storage medium
By training a deep learning model that integrates visual and trajectory features in infrared images, and extracting temporal visual and trajectory features, the problem of identifying drones and birds in traditional methods is solved, and high-precision target recognition under complex conditions is achieved.
Patent Information
- Application Number
- CN202511402441.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-28
- Publication Date
- 2026-01-09
AI Technical Summary
Traditional identification methods based on single-frame infrared images are difficult to effectively distinguish between drones and birds, especially under long-distance imaging or low-resolution conditions, where spatiotemporal modeling methods have insufficient discrimination capabilities.
By training a deep learning recognition model that integrates visual and trajectory features, temporal visual features and trajectory temporal change features are extracted from multiple consecutive frames of infrared images to form a joint feature vector, which is then used for recognition and classification. Combined with camera parameters, the target's motion trajectory in physical space is inferred.
In the presence of nonlinear displacement, rapid attitude changes, and complex flight maneuvers, this system improves the accuracy and robustness of target recognition in infrared images, enhances the ability to distinguish between similar-looking targets, and improves recognition precision and system practicality.
Smart Images

Figure CN121305031A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing and target recognition technology, specifically to a method, apparatus, device, and storage medium for identifying targets in infrared images. Background Technology
[0002] Currently, in single-frame infrared images, drones and birds often appear extremely similar in outline shape and thermal radiation characteristics, especially under long-range imaging or low-resolution conditions. This makes it difficult for traditional single-frame image-based recognition methods to effectively distinguish between the two, severely limiting the practicality and accuracy of infrared target recognition systems.
[0003] To overcome this problem, 3D image recognition methods (such as 3D CNN and Transformer) have gradually become mainstream techniques in video analysis because they can simultaneously extract spatial features (structural and appearance information in images) and temporal features (motion patterns and behavioral information across frames). By jointly modeling spatiotemporal features, these methods have been widely applied in tasks such as video classification, action recognition, and behavior detection in the visible light domain. However, in the field of infrared image analysis, related research is still in its early stages, and its spatiotemporal modeling capabilities and adaptability to real-world scenarios need further exploration.
[0004] In related technologies, current spatiotemporal modeling methods based on continuous frame feature extraction can capture changes in the appearance of a target over time. However, in the presence of nonlinear displacement, drastic attitude changes, or complex flight maneuvers, the model's discriminative ability is significantly reduced, especially when distinguishing between drones and birds that have highly similar appearances.
[0005] Therefore, it is necessary to design a new method for identifying targets in infrared images to overcome the above problems. Summary of the Invention
[0006] This application provides a method, apparatus, device, and computer-readable storage medium for identifying targets in infrared images, which can solve the technical problem of low discrimination capability of spatiotemporal modeling methods in related technologies.
[0007] In a first aspect, embodiments of this application provide a method for identifying targets in infrared images, the method comprising: Based on the labeled target dataset, a deep learning recognition model that integrates visual and trajectory features is trained; Extracting temporal visual features of a target from multiple consecutive frames of infrared images; Based on the position changes of the detection boxes in multiple frames, the motion trajectory of the target in physical space is inferred, and the temporal change features of the trajectory are extracted. The temporal visual features of the target are fused with the temporal change features of the trajectory to form a joint feature vector, and the joint features of the target are identified and classified by a deep learning recognition model.
[0008] In conjunction with the first aspect, in one embodiment, extracting the temporal visual features of the target from consecutive multi-frame infrared images includes: The backbone network is used to extract features from each frame of infrared image to obtain multi-scale feature maps; The target bounding box is generated based on the multi-scale feature map, and the features of the corresponding target bounding box region are cropped on the multi-scale feature map to obtain a visual feature vector of fixed size. An association algorithm is used to track identical targets based on visual features and target motion information, generating target IDs; Temporal fusion of visual features from multiple time points within the same target ID generates temporal visual features of the target.
[0009] In conjunction with the first aspect, in one implementation, the step of inferring the target's motion trajectory in physical space based on the position changes of multi-frame detection boxes and extracting trajectory temporal change features includes: The target pixel displacement of the same target is converted into a spatial angle, and combined with the camera orientation to obtain the spatial orientation angle of the target; The spatial orientation angles of targets in multiple frames are used to construct a motion trajectory sequence, and the temporal variation features of the trajectory are extracted based on the motion trajectory sequence.
[0010] In conjunction with the first aspect, in one embodiment, the step of converting the target pixel displacement of the same target into a spatial angle and combining it with the camera orientation to obtain the spatial orientation angle of the target includes: The horizontal angle of the target in the real world is calculated based on the horizontal offset of the target's pixel relative to the center of the image, the camera's focal length, pixel size, and the current camera's azimuth angle. The vertical angle of the target in the real world is calculated based on the vertical offset of the target's pixel relative to the center of the image, the camera's focal length, pixel size, and the current camera's pitch angle.
[0011] In conjunction with the first aspect, in one implementation, the temporal visual features of the fused target and the temporal change features of its trajectory form a joint feature vector, and a deep learning recognition model is used to identify and classify the joint features of the target, including: The temporal visual features of the target are fused with the temporal change features of the trajectory to form a joint feature vector, and the temporal dependency is encoded through a sequence modeling network before recognition and classification.
[0012] In conjunction with the first aspect, in one implementation, training a deep learning recognition model that integrates visual and trajectory features based on a labeled target dataset includes: Using a labeled target dataset, temporal visual features extracted by the target detection backbone network are fused with trajectory features calculated based on continuous frame detection boxes to form a joint feature vector. The temporal dependencies are then encoded by a sequence modeling network and input into the classification layer for classification and recognition. During training, visual and trajectory features are normalized and aligned, and cross-entropy loss and trajectory constraint loss are jointly optimized.
[0013] In conjunction with the first aspect, in one implementation, before training the deep learning recognition model fusing visual and trajectory features based on the labeled target dataset, the method further includes: Acquire multiple frames of infrared image sequences while recording environmental information and camera parameter information. After acquisition, the multiple frames of infrared images are labeled, including target category, bounding box coordinates, and motion trajectory.
[0014] Secondly, embodiments of this application provide a target identification device in an infrared image, the target identification device in the infrared image comprising: Model training module: It is used to train a deep learning recognition model that integrates visual and trajectory features based on a labeled target dataset; A visual feature extraction module is used to extract temporal visual features of a target from multiple consecutive frames of infrared images; The trajectory feature extraction module is used to infer the target's motion trajectory in physical space based on the position changes of the detection boxes in multiple frames, and extract the trajectory temporal change features. The feature fusion and recognition module is used to fuse the temporal visual features of the target with the temporal change features of the trajectory to form a joint feature vector, and to identify and classify the joint features of the target through a deep learning recognition model.
[0015] Thirdly, embodiments of this application provide a target recognition device in an infrared image. The target recognition device in an infrared image includes a processor, a memory, and a target recognition program in an infrared image stored in the memory and executable by the processor. When the target recognition program in an infrared image is executed by the processor, it implements the steps of the target recognition method in an infrared image described above.
[0016] Fourthly, embodiments of this application provide a computer-readable storage medium storing a target identification program in an infrared image, wherein when the target identification program in the infrared image is executed by a processor, it implements the steps of the above-described target identification method in an infrared image.
[0017] The beneficial effects of the technical solutions provided in this application include: By introducing a target trajectory feature modeling mechanism and combining feature extraction from multiple consecutive frames of infrared images, the system integrates visual appearance features with physical space motion features, effectively enhancing the ability to distinguish between UAVs and similar-looking targets such as birds in terms of motion behavior. This improves the accuracy of identifying complex targets in infrared images and maintains high recognition accuracy and robustness even in the presence of nonlinear displacement, rapid attitude changes, and complex flight maneuvers. It solves the technical problem of low discrimination capability in spatiotemporal modeling methods in related technologies. Attached Figure Description
[0018] Figure 1 This is a flowchart illustrating an embodiment of the target identification method in infrared images according to this application; Figure 2 This is a schematic flowchart illustrating another embodiment of the target identification method in infrared images according to this application; Figure 3 This is a schematic diagram of the hardware structure of the target recognition device in the infrared image involved in the embodiments of this application. Detailed Implementation
[0019] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present application.
[0020] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.
[0021] In a first aspect, embodiments of this application provide a method for identifying targets in infrared images.
[0022] In one embodiment, reference is made to Figure 1 , Figure 1 This is a schematic flowchart of the first embodiment of the target identification method in infrared images according to this application. Figure 1 As shown, the methods for identifying targets in infrared images include: S1: Based on the labeled target dataset, train a deep learning recognition model that integrates visual and trajectory features.
[0023] S2: Extract the temporal visual features of the target from multiple consecutive infrared images.
[0024] S3: Based on the positional changes of the detection boxes in multiple frames, the motion trajectory of the target in physical space is inferred, and the temporal change features of the trajectory are extracted.
[0025] S4: The temporal visual features of the target are fused with the temporal change features of the trajectory to form a joint feature vector, and the joint features of the target are identified and classified by a deep learning recognition model.
[0026] The target recognition method in this embodiment is mainly used to identify drones and birds with extremely similar outlines. Of course, it can also be used to identify other moving objects with extremely similar outlines. In step S1, the labeled target dataset includes infrared images of drones and birds. The images can also be labeled with information such as target category, bounding box coordinates, and motion trajectory. By optimizing model parameters, the detection and classification accuracy of the model in complex environments is improved. In step S2, spatial and temporal features of drones and birds are extracted from multiple consecutive frames of infrared images to fully capture the appearance and dynamic changes of drones and birds, enhancing the ability to perceive details of drones and birds. In step S3, based on the position changes of the detection boxes in multiple frames, the motion trajectories of drones and birds in physical space are inferred, and motion behavior features are extracted to provide motion-level auxiliary information for distinguishing drones from birds. In step S4, the visual and trajectory features of drones and birds are fused together, and a deep learning recognition model is used to classify the joint features of drones and birds, achieving accurate classification and recognition of drones and birds.
[0027] This embodiment introduces a target trajectory feature modeling mechanism, combining feature extraction from multiple consecutive frames of infrared images with visual appearance features and physical space motion features. This effectively enhances the ability to distinguish between similar-looking targets such as drones and birds in terms of motion behavior, improving the accuracy of complex target recognition in infrared images. Even with nonlinear displacement, rapid attitude changes, and complex flight maneuvers, it maintains high recognition accuracy and robustness, solving the technical problem of low discrimination capability in spatiotemporal modeling methods in related technologies. Furthermore, this embodiment addresses the shortcomings of spatiotemporal feature modeling in the infrared field by combining mature spatiotemporal joint modeling ideas from the visible light domain with the characteristics of infrared video, compensating for the deficiencies of existing infrared target recognition in utilizing temporal information and trajectory modeling. Simultaneously, this embodiment improves the system's practicality and stability, ensuring stable and reliable recognition results even under long-distance imaging, low resolution, and high noise interference conditions, thus enhancing the practical deployment value of the infrared target recognition system.
[0028] Furthermore, in one embodiment, before training the deep learning recognition model fusing visual and trajectory features based on the labeled target dataset, the method may further include: acquiring a sequence of multiple infrared images, simultaneously recording environmental information and camera parameter information; after acquisition, labeling the multiple infrared images, including target category, bounding box coordinates, and motion trajectory. In this embodiment, before step S1, a data acquisition step is performed to obtain multiple consecutive frames of infrared video data to ensure the temporal continuity and integrity of the target. This ensures that the data covers different scenes, targets, and motion states, and preprocesses the acquired infrared images (denoising, contrast enhancement, and normalization).
[0029] Specifically, infrared cameras are set up in open areas to ensure targets (drones and birds) are within the field of view, and a calibration board is used for radiometric calibration to ensure data consistency. By simulating scenarios with different altitudes, speeds, and flight trajectories, and including complex backgrounds to increase data diversity, the system acquires multiple frames of infrared image sequences, while simultaneously recording environmental information (such as temperature, humidity, and lighting conditions) and camera parameter information (focal length, pixel size, servo azimuth angle, servo pitch angle, etc.). After acquisition, the video frames are labeled, including target category, bounding box coordinates, and motion trajectory, providing a high-quality and diverse data foundation for subsequent target detection, tracking, and recognition.
[0030] In some optional embodiments, training a deep learning recognition model that integrates visual and trajectory features based on a labeled target dataset may include: using the labeled target dataset, fusing temporal visual features extracted by the target detection backbone network with trajectory features calculated based on continuous frame detection boxes to form a joint feature vector, and then encoding the temporal dependencies through a sequence modeling network before inputting it into the classification layer for classification and recognition; wherein, during the training process, the visual and trajectory features are normalized and aligned, and cross-entropy loss and trajectory constraint loss are jointly optimized.
[0031] In this embodiment, using labeled datasets of drones and birds, temporal visual features extracted by the target detection backbone network are fused with trajectory features calculated based on continuous frame detection boxes to form a joint feature vector. This vector is then encoded using a sequence modeling network (such as Transformer or LSTM) to encode temporal dependencies before being input into the classification layer for drone and bird recognition. During training, visual and trajectory features are normalized and aligned, and cross-entropy loss and trajectory constraint loss are jointly optimized to ensure that the model can achieve high-precision, multimodal target recognition in scenarios with similar appearances or complex motion.
[0032] Furthermore, in one embodiment, in step S2, extracting the temporal visual features of the target from multiple consecutive infrared images may include: S21: Use the backbone network to extract features from each frame of infrared image to obtain multi-scale feature maps.
[0033] S22: Generate target bounding boxes based on multi-scale feature maps, and crop the features of the corresponding target bounding box regions on the multi-scale feature maps to obtain visual feature vectors of fixed size.
[0034] S23: Using an association algorithm, the same targets are tracked based on visual features and target motion information to generate target IDs.
[0035] S24: Perform temporal fusion of visual features from multiple time points within the same target ID to generate temporal visual features of the target.
[0036] In this embodiment, firstly, the backbone network of a general target detection model (such as YOLOv8 or DETR) is used to extract features from each frame of infrared image, obtaining a multi-scale feature map that fully reflects the spatial structure and texture information of the UAV and birds. Next, the detection head generates target bounding boxes based on the multi-scale feature maps and crops the features of the corresponding target box regions on the feature maps to obtain fixed-size visual feature vectors, which represent the visual features of the target in a single frame. Subsequently, in consecutive video frames, an association algorithm (such as Deep SORT) is used to track the same targets based on visual features and target motion information (position, velocity, etc.), generating target IDs. This step establishes pipeline associations for the same tracked target; this association mechanism ensures the temporal continuity of visual features and improves the stability of target tracking. Finally, the visual features from multiple moments within the same target ID are temporally fused to generate comprehensive temporal visual features that reflect both the target's appearance and its dynamic changes. This embodiment first detects the target region, tracks and establishes pipeline association, and then performs temporal recognition. This can effectively avoid the fact that the temporal recognition targets are not the same. There are no specific restrictions on the temporal recognition method (3DCNN, transformer, etc. are all acceptable).
[0037] In some embodiments, the step of inferring the target's motion trajectory in physical space based on the position changes of multi-frame detection boxes and extracting trajectory temporal change features may include: S31: Convert the target pixel displacement of the same target into a spatial angle, and combine it with the camera orientation to obtain the spatial orientation angle of the target.
[0038] S32: Construct a motion trajectory sequence from the spatial orientation angles of targets in multiple frames, and extract the temporal variation features of the trajectory based on the motion trajectory sequence.
[0039] In this embodiment, feature extraction first converts the pixel displacements of drones and birds in the same pipeline (i.e., those with the same ID) into spatial angles, and then combines this with the camera orientation to obtain the spatial orientation angles of the drones and birds respectively. Next, multiple frames of target spatial orientation angles are used to construct a motion trajectory sequence, which is then fed into a trajectory encoding module (such as 1D-CNN) to extract the temporal variation features of the trajectory. Because the servo of the photoelectric device continuously moves in tracking mode to keep the target in the center of the field of view, it is almost impossible to intuitively perceive the target's true motion trajectory (such as the bird's speed, acceleration, or vertical flight) from the video footage. Therefore, this embodiment first converts the pixel displacements of the same target in the image into spatial angles, and then combines this with camera attitude information to obtain the target's spatial orientation angles, thereby calculating the motion parameters in the real coordinate system. The reason for needing servo motion is that it enables long-term stable tracking of the target in tracking mode; if the servo remains stationary, the target will quickly fly out of the field of view, making continuous tracking impossible.
[0040] Based on the above technical solution, the step of converting the target pixel displacement of the same target into a spatial angle and combining it with the camera orientation to obtain the spatial orientation angle of the target includes: calculating the horizontal angle of the target in the real world based on the horizontal offset of the pixel relative to the center of the image, the focal length of the camera, the pixel size, and the azimuth angle of the current camera; and calculating the vertical angle of the target in the real world based on the vertical offset of the pixel relative to the center of the image, the focal length of the camera, the pixel size, and the pitch angle of the current camera.
[0041] In this embodiment, the formula for converting the horizontal pixel offset in the image into a horizontal angle in the real world is as follows: .
[0042] In the formula: It is the horizontal pixel offset of the target relative to the center of the image (e.g., the target). Coordinates - Image Center coordinate), It's the camera's focal length. It is the pixel size. The current azimuth angle of the camera, The formula for perpendicular angle is similar: .
[0043] In the formula: It is the vertical offset of the target pixel relative to the center of the image (e.g., the target pixel). Coordinates - Image Center coordinate), It's the camera's focal length. It represents the pixel size; the negative sign is due to the image size. The axis direction is opposite to the direction of view change (downward is positive in the image, and downward is negative in the view). The current camera's pitch angle. The formula above is used to convert the target's position in the image into a true geographic angle or viewing direction.
[0044] This embodiment calculates the motion trajectory by reversing the target's position in the real world based on its position in the image, thus avoiding interference from device motion during data acquisition.
[0045] Further, in one embodiment, the fusion of the temporal visual features of the target and the temporal change features of the trajectory forms a joint feature vector, and the joint features of the target are identified and classified using a deep learning recognition model. This can include: fusing the temporal visual features of the target and the temporal change features of the trajectory to form a joint feature vector, encoding the temporal dependencies through a sequence modeling network, and then performing recognition and classification. In this embodiment, visual features and trajectory features are fused to form a joint feature vector, and the temporal dependencies are encoded through a sequence modeling network. The vector is then input into a fully connected layer and a Softmax classifier to obtain the target category and confidence level.
[0046] This application introduces a continuous-frame target recognition and motion trajectory modeling mechanism to achieve high-precision classification and recognition of similar moving targets (such as drones and birds) in infrared video. Specifically, by fusing visual appearance features and physical space motion trajectory features, it effectively enhances the ability to distinguish highly similar targets (such as drones and birds), improving the accuracy of infrared video target recognition. Furthermore, by using the pixel coordinates of the target in the image combined with the camera imaging model to deduce its three-dimensional motion trajectory in the camera coordinate system, it can model the target's true motion path, compensating for the shortcomings of spatiotemporal joint modeling in the infrared field. It can stably output reliable recognition results under long-distance, low-resolution, or noisy conditions, enhancing the practical application value of the infrared target recognition system. Simultaneously, by using only monocular infrared video sequences for feature modeling without relying on depth maps or multimodal sensors, it enables lightweight and low-cost deployment, suitable for unattended long-distance monitoring and intelligent recognition scenarios.
[0047] Secondly, embodiments of this application also provide a target identification device in an infrared image.
[0048] In one embodiment, the target recognition device in infrared images includes: a model training module for training a deep learning recognition model that integrates visual and trajectory features based on a labeled target dataset; a visual feature extraction module for extracting temporal visual features of the target from multiple consecutive frames of infrared images; a trajectory feature extraction module for inferring the target's motion trajectory in physical space based on the position changes of detection boxes in multiple frames and extracting temporal trajectory change features; and a feature fusion and recognition module for fusing the target's temporal visual features and trajectory temporal change features to form a joint feature vector, and using the deep learning recognition model to identify and classify the joint features of the target.
[0049] Furthermore, in one embodiment, the visual feature extraction module is used to extract features from each frame of infrared image using a backbone network to obtain a multi-scale feature map; generate a target bounding box based on the multi-scale feature map, and crop the features of the corresponding target bounding box region on the multi-scale feature map to obtain a visual feature vector of fixed size; use an association algorithm to track the same target based on visual features and target motion information to generate a target ID; and perform temporal fusion of visual features at multiple times within the same target ID to generate temporal visual features of the target.
[0050] Furthermore, in one embodiment, the trajectory feature extraction module is used to convert the target pixel displacement of the same target into a spatial angle, and combine it with the camera orientation to obtain the spatial orientation angle of the target; construct a motion trajectory sequence from the spatial orientation angles of multiple frames of targets, and extract the trajectory temporal change features based on the motion trajectory sequence.
[0051] Furthermore, in one embodiment, the step of converting the target pixel displacement of the same target into a spatial angle and combining it with the camera orientation to obtain the spatial orientation angle of the target includes: calculating the horizontal angle of the target in the real world based on the horizontal offset of the target's pixel relative to the center of the image, the camera's focal length, pixel size, and the current camera's azimuth angle; and calculating the vertical angle of the target in the real world based on the vertical offset of the target's pixel relative to the center of the image, the camera's focal length, pixel size, and the current camera's pitch angle.
[0052] Furthermore, in one embodiment, the feature fusion and recognition module is used to fuse the temporal visual features of the target with the temporal change features of the trajectory to form a joint feature vector, and then encode the temporal dependency relationship through a sequence modeling network before recognition and classification.
[0053] Furthermore, in one embodiment, the model training module is used to utilize the labeled target dataset, fuse the temporal visual features extracted by the target detection backbone network with the trajectory features calculated based on the continuous frame detection boxes to form a joint feature vector, and encode the temporal dependencies through the sequence modeling network before inputting it into the classification layer for classification and recognition; wherein, during the training process, the visual and trajectory features are normalized and aligned, and the cross-entropy loss and trajectory constraint loss are jointly optimized.
[0054] Furthermore, in one embodiment, the target identification device in the infrared image further includes a data acquisition module, which is used to acquire a sequence of multiple infrared images, while recording environmental information and camera parameter information. After acquisition, the multiple infrared images are labeled, including target category, bounding box coordinates and motion trajectory.
[0055] The functions of each module in the target identification device in the infrared image described above correspond to the steps in the target identification method embodiment described above, and their functions and implementation processes will not be described in detail here.
[0056] This application innovatively analyzes the same target region in consecutive frames and, combined with a camera imaging model, inverts the two-dimensional positional changes of the target in the image into a three-dimensional motion trajectory in the camera coordinate system, thereby obtaining the target's true spatiotemporal behavioral characteristics. Furthermore, it fuses the extracted motion trajectory features with the visual spatiotemporal features of the image sequence to achieve joint modeling of the target's appearance and dynamic behavior, effectively improving the ability to distinguish similar-looking targets (such as drones and birds) in infrared images. Multiple experiments have verified that the method of this embodiment can significantly improve recognition accuracy.
[0057] In terms of applications, this application does not rely on depth maps or multi-view devices, and has the advantages of being lightweight and easy to deploy, making it suitable for intelligent identification tasks of small targets in complex environments. Its applications are wide-ranging, including UAV monitoring, air traffic management, military reconnaissance, and other fields.
[0058] Thirdly, embodiments of this application provide a target identification device in an infrared image. The target identification device in an infrared image can be a personal computer (PC), a laptop computer, a server, or other device with data processing capabilities.
[0059] Reference Figure 3 , Figure 3 This is a schematic diagram of the hardware structure of a target identification device in an infrared image according to an embodiment of this application. In this embodiment, the target identification device in an infrared image may include a processor, a memory, a communication interface, and a communication bus.
[0060] The communication bus can be of any type and is used to interconnect the processor, memory, and communication interface.
[0061] Communication interfaces include input / output (I / O) interfaces, physical interfaces, and logical interfaces used for interconnecting internal components of the infrared target recognition device, as well as interfaces used for interconnecting the infrared target recognition device with other devices (such as other computing devices or user equipment). Physical interfaces can be Ethernet interfaces, fiber optic interfaces, ATM interfaces, etc.; user equipment can be displays, keyboards, etc.
[0062] Memory can be various types of storage media, such as random access memory (RAM), read-only memory (ROM), non-volatile RAM (NVRAM), flash memory, optical storage, hard disk, programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), etc.
[0063] The processor can be a general-purpose processor, which can call the target recognition program in the infrared image stored in the memory and execute the target recognition method in the infrared image provided in the embodiments of this application. For example, the general-purpose processor can be a central processing unit (CPU). The method executed when the target recognition program in the infrared image is called can refer to the various embodiments of the target recognition method in the infrared image of this application, and will not be repeated here.
[0064] Those skilled in the art will understand that Figure 3 The hardware structure shown does not constitute a limitation of this application and may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0065] Fourthly, embodiments of this application also provide a readable storage medium.
[0066] The present application has a program for identifying targets in infrared images stored on a readable storage medium, wherein when the program for identifying targets in infrared images is executed by a processor, it implements the steps of the method for identifying targets in infrared images as described above.
[0067] The method implemented when the target identification program in the infrared image is executed can be referred to in the various embodiments of the target identification method in the infrared image of this application, and will not be repeated here.
[0068] It should be noted that the sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0069] The terms "comprising" and "having," and any variations thereof, in the specification, claims, and accompanying drawings of this application are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to such process, method, product, or apparatus. The terms "first," "second," and "third," etc., are used to distinguish different objects, etc., and do not indicate a sequence, nor do they limit "first," "second," and "third" to different types.
[0070] In the description of the embodiments of this application, terms such as "exemplary," "for example," or "for instance" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as "exemplary," "for example," or "for instance" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of terms such as "exemplary," "for example," or "for instance" is intended to present the relevant concepts in a concrete manner.
[0071] In the description of the embodiments of this application, unless otherwise stated, " / " means "or". For example, A / B can mean A or B. The "and / or" in the text is merely a description of the relationship between related objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist simultaneously, and B exists alone. In addition, in the description of the embodiments of this application, "multiple" means two or more.
[0072] In some processes described in the embodiments of this application, multiple operations or steps are included in a specific order. However, it should be understood that these operations or steps may not be executed in the order they appear in the embodiments of this application, or they may be executed in parallel. The sequence number of the operation is only used to distinguish different operations, and the sequence number itself does not represent any execution order. In addition, these processes may include more or fewer operations, and these operations or steps may be executed sequentially or in parallel, and these operations or steps may be combined.
[0073] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes several instructions to cause a terminal device to execute the methods described in the various embodiments of this application.
[0074] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.
Claims
1. A method for identifying targets in infrared images, characterized in that, The method for identifying targets in infrared images includes: Based on the labeled target dataset, a deep learning recognition model that integrates visual and trajectory features is trained; Extracting temporal visual features of a target from multiple consecutive frames of infrared images; Based on the position changes of the detection boxes in multiple frames, the motion trajectory of the target in physical space is inferred, and the temporal change features of the trajectory are extracted. The temporal visual features of the target are fused with the temporal change features of the trajectory to form a joint feature vector, and the joint features of the target are identified and classified by a deep learning recognition model.
2. The method for identifying targets in infrared images as described in claim 1, characterized in that, The extraction of temporal visual features of the target from consecutive multi-frame infrared images includes: The backbone network is used to extract features from each frame of infrared image to obtain multi-scale feature maps; The target bounding box is generated based on the multi-scale feature map, and the features of the corresponding target bounding box region are cropped on the multi-scale feature map to obtain a visual feature vector of fixed size. An association algorithm is used to track identical targets based on visual features and target motion information, generating target IDs; Temporal fusion of visual features from multiple time points within the same target ID generates temporal visual features of the target.
3. The method for identifying targets in infrared images as described in claim 1, characterized in that, The method of inferring the target's motion trajectory in physical space based on the position changes of multi-frame detection boxes and extracting trajectory temporal change features includes: The target pixel displacement of the same target is converted into a spatial angle, and combined with the camera orientation to obtain the spatial orientation angle of the target; The spatial orientation angles of targets in multiple frames are used to construct a motion trajectory sequence, and the temporal variation features of the trajectory are extracted based on the motion trajectory sequence.
4. The method for identifying targets in infrared images as described in claim 3, characterized in that, The step of converting the target pixel displacement of the same target into a spatial angle and combining it with the camera orientation to obtain the spatial orientation angle of the target includes: The horizontal angle of the target in the real world is calculated based on the horizontal offset of the target's pixel relative to the center of the image, the camera's focal length, pixel size, and the current camera's azimuth angle. The vertical angle of the target in the real world is calculated based on the vertical offset of the target's pixel relative to the center of the image, the camera's focal length, pixel size, and the current camera's pitch angle.
5. The method for identifying targets in infrared images as described in claim 1, characterized in that, The fused target's temporal visual features and trajectory temporal change features form a joint feature vector, and a deep learning recognition model is used to identify and classify the target's joint features, including: The temporal visual features of the target are fused with the temporal change features of the trajectory to form a joint feature vector, and the temporal dependency is encoded through a sequence modeling network before recognition and classification.
6. The method for identifying targets in infrared images as described in claim 1, characterized in that, The training of a deep learning recognition model that integrates visual and trajectory features based on the labeled target dataset includes: Using a labeled target dataset, temporal visual features extracted by the target detection backbone network are fused with trajectory features calculated based on continuous frame detection boxes to form a joint feature vector. The temporal dependencies are then encoded by a sequence modeling network and input into the classification layer for classification and recognition. During training, visual and trajectory features are normalized and aligned, and cross-entropy loss and trajectory constraint loss are jointly optimized.
7. The method for identifying targets in infrared images as described in claim 1, characterized in that, Before training the deep learning recognition model that fuses visual and trajectory features based on the labeled target dataset, the following steps are also included: Acquire multiple frames of infrared image sequences while recording environmental information and camera parameter information. After acquisition, the multiple frames of infrared images are labeled, including target category, bounding box coordinates, and motion trajectory.
8. A target identification device in an infrared image, characterized in that, The target identification device in the infrared image includes: Model training module: It is used to train a deep learning recognition model that integrates visual and trajectory features based on a labeled target dataset; A visual feature extraction module is used to extract temporal visual features of a target from multiple consecutive frames of infrared images; The trajectory feature extraction module is used to infer the target's motion trajectory in physical space based on the position changes of the detection boxes in multiple frames, and extract the trajectory temporal change features. The feature fusion and recognition module is used to fuse the temporal visual features of the target with the temporal change features of the trajectory to form a joint feature vector, and to identify and classify the joint features of the target through a deep learning recognition model.
9. A target identification device in an infrared image, characterized in that, The target identification device in the infrared image includes a processor, a memory, and a target identification program in the infrared image stored in the memory and executable by the processor, wherein when the target identification program in the infrared image is executed by the processor, it implements the steps of the target identification method in the infrared image as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a target identification program in an infrared image, wherein when the target identification program in the infrared image is executed by a processor, it implements the steps of the target identification method in an infrared image as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Method for distinguishing unmanned aerial vehicle and bird by fusing shape change characteristics
CN115063743A
Unmanned aerial vehicle identification method based on deep learning and target motion trail
CN117079196A
Unmanned aerial vehicle and bird monitoring method based on convolutional neural network
CN117876942A
Infrared image weak and small moving target detection method based on track discrimination network model
CN120125924A
Cited By
Heterogeneous feature fusion-based moving target trajectory distinguishing method, device and equipment
CN121598033A
Method, device and equipment for distinguishing moving target trajectory based on heterogeneous feature fusion
CN121598033B
Infrared target detection method and device based on dual-channel feature fusion
CN122024167A