Object identification method and system based on multiband infrared image, and medium

Through the multi-band infrared image recognition method, using a multi-band infrared camera to collect and preprocess images, combined with convolutional neural networks and weighted fusion strategies, the problem of insufficient accuracy of single-band infrared image recognition in complex environments is solved, and higher-precision object recognition is achieved.

CN120635592APending Publication Date: 2025-09-12HANGZHOU HUICUI INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510991991.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-18
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Traditional single-band infrared image recognition methods are difficult to meet the needs of object recognition in complex environments due to the lack of sufficient spectral information, resulting in insufficient recognition accuracy and robustness.

Method used

A multi-band infrared image recognition method is adopted. Images are collected by a multi-band infrared camera. After denoising, registration and normalization processing, convolutional neural network is used to extract image features, and object recognition is performed through weighted fusion strategy and classifier.

Benefits of technology

The accuracy and adaptability of object recognition are improved, and objects can be identified more accurately in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120635592A_ABST
    Figure CN120635592A_ABST
Patent Text Reader

Abstract

The invention provides an object recognition method and system based on a multi-band infrared image and a medium. The method comprises the steps that the multi-band infrared image of a target object is collected based on a multi-band infrared camera; the collected multi-band infrared image is preprocessed, the preprocessing comprises at least one of denoising, registration and normalization processing, and the preprocessed multi-band infrared image is obtained; extracting image features of the infrared image of each wave band based on a convolutional neural network; based on a weighted fusion strategy, carrying out fusion processing on the image features of the infrared images of all the wave bands to obtain fused image features; performing classification processing on the fused image features based on a classifier to obtain an object recognition result; by fusing the features of the multi-band infrared image, the precision and adaptability of object recognition are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of object recognition technology, and more specifically, to a method, system, and medium for object recognition based on multi-band infrared images. Background Art

[0002] Infrared imaging technology generates images by capturing infrared radiation emitted by objects and is widely used in military, security, and medical fields. Traditional single-band infrared image recognition methods are limited by the information contained in a single band, making them inadequate for object recognition in complex environments. Multi-band infrared imagery provides richer spectral information, helping to improve the accuracy and robustness of object recognition.

[0003] AI-based infrared object recognition technology is crucial for a variety of autonomous applications. Thermal infrared can be categorized as long-wave infrared (LWIR), medium-wave infrared (MWIR), and short-wave infrared (SWIR), covering the atmospheric window wavelengths of 8-14 microns, 3-5 microns, and 1-3 microns, respectively. These radiation bands have proven useful in a variety of imaging applications, particularly where signals from other bands (such as visible light, near-infrared, or short-infrared) are weak or unavailable. Long-wave infrared images capture the thermodynamic distribution of objects near room temperature, providing superior information, especially when the temperature of the imaged object differs from the ambient temperature.

[0004] Existing infrared (IR) cameras typically capture all radiation within a band and output a single-channel image whose intensity is the average of the radiation across the entire band. Similar to traditional visible light images (red, green, and blue), thermal imaging also captures the intensity of infrared sub-bands. However, three-channel visible light images typically offer better object recognition performance than their single-channel counterparts. Summary of the Invention

[0005] The purpose of the embodiments of the present application is to provide an object recognition method, system and medium based on multi-band infrared images, which improves the accuracy and adaptability of object recognition by fusing the features of multi-band infrared images.

[0006] The present application also provides an object recognition method based on multi-band infrared images, including: Collect multi-band infrared images of target objects based on a multi-band infrared camera; Preprocessing the collected multi-band infrared image, wherein the preprocessing includes at least one of the following: denoising, registration and normalization processing, to obtain a preprocessed multi-band infrared image; Extract image features of each band of infrared images based on convolutional neural network; Based on the weighted fusion strategy, the image features of each band of infrared images are fused to obtain the fused image features; The fused image features are classified based on the classifier to obtain the object recognition results.

[0007] Optionally, in the object recognition method based on multi-band infrared images described in the embodiment of the present application, collecting a multi-band infrared image of the target object based on a multi-band infrared camera specifically includes: Select a multi-band infrared camera and configure camera parameters, wherein the camera parameters include at least one of the following: operating band, gain, and integration time; Capture scene images based on camera parameters and analyze object positions within the scene images; Determine whether the object is in the center of the scene based on the object's position; If it is not in the center of the scene, adjust the camera acquisition angle; If it is at the center of the scene, the image background is set and a multi-band infrared image is collected in real time. The multi-band infrared image includes an infrared image of a first band, an infrared image of a second band and an infrared image of a third band. The wavelength of the first band is smaller than that of the second band, and the wavelength of the second band is smaller than that of the third band.

[0008] Optionally, in the object recognition method based on multi-band infrared images described in the embodiment of the present application, preprocessing the collected multi-band infrared images specifically includes: Set the neighborhood window size, obtain the pixel values ​​of all pixels in the neighborhood window, and calculate the average value of the pixel values ​​of all pixels to obtain the pixel average value; The pixel average value is used as the pixel value of the center point of the neighborhood window to obtain the denoised multi-band infrared image; Extract feature points from multi-band infrared images and find corresponding feature point pairs between images of different bands based on feature point matching algorithm; The transformation parameters are calculated based on the corresponding feature point pairs. The transformation parameters include translation parameters, rotation parameters and scaling parameters. Based on the transformation parameters, the multi-band infrared images are registered and transformed. The infrared image of each band after registration transformation is linearly normalized to obtain the preprocessed multi-band infrared image.

[0009] Optionally, in the object recognition method based on multi-band infrared images described in the embodiment of the present application, extracting image features of each band of infrared images based on a convolutional neural network specifically includes: Select the infrastructure, configure network parameters, and add target layers, including attention and pooling layers. The preprocessed multi-band infrared images are used as training sets to train the infrastructure and obtain a convolutional neural network model. Based on the convolutional neural network model, a layer that meets the preset conditions is selected as the feature extraction layer, and the infrared image of each band is input into the trained convolutional neural network model; The output is obtained based on the feature extraction layer to obtain a corresponding feature vector, which includes at least one of the following: texture features, shape features, and temperature distribution features of the multi-band infrared image.

[0010] Optionally, in the object recognition method based on multi-band infrared images described in the embodiment of the present application, the image features of each band of infrared images are fused based on a weighted fusion strategy, specifically including: Obtain the feature vector of the multi-band infrared image and analyze the dimension of the feature vector; Analyze the variance of each band’s eigenvector based on its dimension; Normalize the variance of each band’s feature vector and use it as the weight; The image features are weighted fused based on the weight of each band to obtain the fused image features.

[0011] Optionally, in the object recognition method based on multi-band infrared images described in the embodiment of the present application, classifying the fused image features based on a classifier to obtain an object recognition result specifically includes: Obtain the fused image features, associate the fused image features with the corresponding object category labels to obtain labeled data; Setting hyperparameters, training a classifier based on the labeled data, and obtaining a training result, wherein the classifier includes a support vector machine or a fully connected neural network; Adjust the hyperparameters based on the training results to obtain the trained classifier; Classify the fused image features based on the trained classifier to obtain the classification prediction results; The accuracy of the classification prediction results is evaluated based on the evaluation indicators to obtain the prediction accuracy; The classifier's hyperparameters are adjusted secondary based on the prediction accuracy.

[0012] In a second aspect, an embodiment of the present application provides an object recognition system based on multi-band infrared images, the system comprising: a memory and a processor, the memory comprising a program for a method for object recognition based on multi-band infrared images, the program for the method for object recognition based on multi-band infrared images, when executed by the processor, implementing the following steps: Collect multi-band infrared images of target objects based on a multi-band infrared camera; Preprocessing the collected multi-band infrared image, wherein the preprocessing includes at least one of the following: denoising, registration and normalization processing, to obtain a preprocessed multi-band infrared image; Extract image features of each band of infrared images based on convolutional neural network; Based on the weighted fusion strategy, the image features of each band of infrared images are fused to obtain the fused image features; The fused image features are classified based on the classifier to obtain the object recognition results.

[0013] Optionally, in the object recognition system based on multi-band infrared images described in the embodiment of the present application, collecting multi-band infrared images of the target object based on the multi-band infrared camera specifically includes: Select a multi-band infrared camera and configure camera parameters, wherein the camera parameters include at least one of the following: operating band, gain, and integration time; Capture scene images based on camera parameters and analyze object positions within the scene images; Determine whether the object is in the center of the scene based on the object's position; If it is not in the center of the scene, adjust the camera acquisition angle; If it is at the center of the scene, the image background is set and a multi-band infrared image is collected in real time. The multi-band infrared image includes an infrared image of a first band, an infrared image of a second band and an infrared image of a third band. The wavelength of the first band is smaller than that of the second band, and the wavelength of the second band is smaller than that of the third band.

[0014] Optionally, in the object recognition system based on multi-band infrared images described in the embodiment of the present application, preprocessing the collected multi-band infrared images specifically includes: Set the neighborhood window size, obtain the pixel values ​​of all pixels in the neighborhood window, and calculate the average value of the pixel values ​​of all pixels to obtain the pixel average value; The pixel average value is used as the pixel value of the center point of the neighborhood window to obtain the denoised multi-band infrared image; Extract feature points from multi-band infrared images and find corresponding feature point pairs between images of different bands based on feature point matching algorithm; The transformation parameters are calculated based on the corresponding feature point pairs. The transformation parameters include translation parameters, rotation parameters and scaling parameters. Based on the transformation parameters, the multi-band infrared images are registered and transformed. The infrared image of each band after registration transformation is linearly normalized to obtain the preprocessed multi-band infrared image.

[0015] In a third aspect, an embodiment of the present application further provides a computer-readable storage medium, which includes an object recognition method program based on multi-band infrared images. When the object recognition method program based on multi-band infrared images is executed by a processor, the steps of the object recognition method based on multi-band infrared images as described in any one of the above items are implemented.

[0016] As can be seen from the above, the embodiments of the present application provide an object recognition method, system and medium based on multi-band infrared images, which collect multi-band infrared images of target objects based on a multi-band infrared camera; pre-process the collected multi-band infrared images, wherein the pre-processing includes at least one of the following: denoising, alignment and normalization processing to obtain pre-processed multi-band infrared images; extract image features of each band infrared image based on a convolutional neural network; fuse the image features of each band infrared image based on a weighted fusion strategy to obtain fused image features; classify the fused image features based on a classifier to obtain object recognition results; and improve the accuracy and adaptability of object recognition by fusing the features of multi-band infrared images. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments of the present application. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without creative work.

[0018] Figure 1 A flowchart of an object recognition method based on multi-band infrared images provided in an embodiment of the present application; Figure 2 A multi-band infrared image acquisition flow chart of the object recognition method based on multi-band infrared images provided in an embodiment of the present application; Figure 3 This is a flow chart of a multi-band infrared image denoising, registration and normalization processing method for an object recognition method based on multi-band infrared images provided in an embodiment of the present application. DETAILED DESCRIPTION

[0019] The technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. The components of the embodiments of the present application generally described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the application for protection, but merely represents the selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without making creative work fall within the scope of protection of the present application.

[0020] It should be noted that similar reference numerals and letters represent similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings. At the same time, in the description of this application, the terms "first", "second", etc. are only used to distinguish the description and should not be understood as indicating or implying relative importance.

[0021] Please refer to Figure 1 , Figure 1 The flowchart of a method for object recognition based on multi-band infrared images in some embodiments of the present application is shown. The method for object recognition based on multi-band infrared images is used in a terminal device and includes the following steps: S101, collecting a multi-band infrared image of a target object using a multi-band infrared camera; S102, preprocessing the collected multi-band infrared image, wherein the preprocessing includes at least one of the following: denoising, registration, and normalization, to obtain a preprocessed multi-band infrared image; S103, extracting image features of each band of infrared images based on a convolutional neural network; S104, performing fusion processing on the image features of each band of infrared images based on a weighted fusion strategy to obtain fused image features; S105: Classify the fused image features based on the classifier to obtain an object recognition result.

[0022] It should be noted that automatic feature extraction using convolutional neural networks (CNNs), such as ResNet and DenseNet, can extract deep semantic features of images layer by layer, effectively capturing complex feature information of objects and showing good adaptability to different scenes and objects.

[0023] Image feature classification methods, including traditional classifiers such as support vector machines (SVMs), decision trees, and naive Bayesian methods, can be used for object recognition in multi-band infrared images. Based on the extracted and fused features, a classification model is trained to classify objects.

[0024] The formula for extracting image features is as follows: Fb=CNN(Ib); Among them, Ib represents the image of the b-th band, and Fb represents the image features extracted from the b-th band infrared image, which is a feature vector or feature map used for subsequent fusion and classification operations.

[0025] CNN, or Convolutional Neural Network, is a deep learning model used to automatically extract multi-level, discriminative features from input images.

[0026] Ib represents the infrared image of the bth band, that is, the original input image. Multiple bands represent images at different infrared wavelengths, such as shortwave, mediumwave, and longwave infrared.

[0027] The feature fusion calculation formula is as follows: Ffusion=∑b=1Bwb; “1” means starting from the first band; until the Bth band, the features extracted from all band images are weighted summed; Among them, Ffusion represents the fused image feature vector, wb is the weight of the b-th band, and B is the total number of bands.

[0028] y=Classifier(Ffusion), where Classifier represents the classifier model, such as support vector machine (SVM) or fully connected neural network, and y represents the recognition result.

[0029] Please refer to Figure 2 , Figure 2 This is a multi-band infrared image acquisition flow chart of a method for object recognition based on multi-band infrared images in some embodiments of the present application. According to an embodiment of the present invention, capturing a multi-band infrared image of a target object using a multi-band infrared camera specifically includes: S201, selecting a multi-band infrared camera and configuring camera parameters, where the camera parameters include at least one of the following: operating band, gain, and integration time; S202, capturing a scene image based on camera parameters and analyzing object positions within the scene image; S203, determining whether the object is located at the center of the scene based on the object position; S204, if it is not at the center of the scene, adjust the camera acquisition angle; S205: If the subject is at the center of the scene, an image background is set and a multi-band infrared image is collected in real time. The multi-band infrared image includes an infrared image of a first band, an infrared image of a second band, and an infrared image of a third band. The wavelength of the first band is smaller than that of the second band, and the wavelength of the second band is smaller than that of the third band.

[0030] It should be noted that the wavelength setting includes determining the specific band in which the camera works, which can be selected according to the radiation characteristics of the target object and the recognition purpose, and a specific infrared band that is sensitive to the object's characteristics can be selected.

[0031] Gain adjustment involves adjusting the gain based on the radiation intensity of the scene. If the target object radiates weakly, the gain can be increased to enhance the signal. However, excessive gain can introduce noise, so proper adjustment is required.

[0032] Integration time settings, including the integration time, affect image brightness and signal-to-noise ratio. For fast-moving targets, a shorter integration time is recommended to avoid smearing. For stationary targets with weaker radiation, a longer integration time can be used to capture more signal.

[0033] Clearly identify the target object's location, ensuring it is in the center of the camera's field of view or within the effective imaging range. Use auxiliary equipment such as laser pointers to assist with positioning. Choose a suitable background to create a clear radiometric difference between the target and the background, facilitating subsequent image analysis and target recognition. For example, use a dark, uniform background to highlight the target object. Collect multiple sets of images as needed, such as from different angles and at different times, to obtain more comprehensive information about the target object.

[0034] Please refer to Figure 3 , Figure 3 This is a flowchart of a multi-band infrared image denoising, registration, and normalization processing method for an object recognition method based on multi-band infrared images in some embodiments of the present application. According to an embodiment of the present invention, preprocessing of the collected multi-band infrared images specifically includes: S301, setting the neighborhood window size, obtaining the pixel values ​​of all pixels in the neighborhood window, and calculating the average value of the pixel values ​​of all pixels to obtain the pixel average value; S302, taking the pixel average as the pixel value of the center point of the neighborhood window to obtain a denoised multi-band infrared image; S303, extracting feature points from the multi-band infrared image, and finding corresponding feature point pairs between images of different bands based on a feature point matching algorithm; S304, calculating transformation parameters based on the corresponding feature point pairs, the transformation parameters including translation parameters, rotation parameters, and scaling parameters, and performing registration transformation on the multi-band infrared image based on the transformation parameters; S305 , performing a linear normalization operation on the infrared image of each band after the registration transformation to obtain a pre-processed multi-band infrared image.

[0035] It should be noted that median filtering is a nonlinear filtering method that replaces the value of each pixel in an image with the median of the pixels in its neighborhood. For multi-band infrared images, median filtering can be performed on each band separately to effectively remove salt-and-pepper noise while effectively preserving edge information. For example, for a 3×3 neighborhood window, the pixel values ​​within the window are sorted by size and the median value is used to replace the value of the center pixel in the window.

[0036] Feature point extraction algorithms include SIFT (Scale Invariant Feature Transform), SURF (Speeded Up Robust Features), etc. Feature point matching algorithms include the nearest neighbor matching algorithm based on Euclidean distance.

[0037] Linear normalization maps the pixel values ​​of an image to a fixed range, usually [0, 1] or [0, 255]. For multi-band infrared images, linear normalization is performed on each band separately.

[0038] According to an embodiment of the present invention, the image features of each band of infrared images are extracted based on a convolutional neural network, specifically including: Select the infrastructure, configure network parameters, and add target layers, including attention and pooling layers. The preprocessed multi-band infrared images are used as training sets to train the infrastructure and obtain a convolutional neural network model. Based on the convolutional neural network model, a layer that meets the preset conditions is selected as the feature extraction layer, and the infrared image of each band is input into the trained convolutional neural network model; The output is obtained based on the feature extraction layer to obtain a corresponding feature vector, which includes at least one of the following: texture features, shape features, and temperature distribution features of the multi-band infrared image.

[0039] It should be noted that the classic CNN architectures such as LeNet, AlexNet, VGGNet, and ResNet are used as the basis. For example, LeNet has a simple structure and is suitable for relatively simple image feature extraction tasks. ResNet solves the problem of vanishing gradients in deep network training through residual connections and can effectively extract complex features.

[0040] Adjust the model based on the characteristics of the infrared image and the number of bands. If the infrared image resolution is low, the number of convolutional layers or the size of the convolution kernel can be appropriately reduced. For multi-band scenarios, the number of channels in the input layer can be set to correspond to the number of bands, allowing the model to process image information from multiple bands simultaneously.

[0041] To better extract infrared image features, some target layers can be added. For example, an attention mechanism layer can be added to focus the model on features in key areas of the infrared image; or a pooling layer can be added to reduce the dimension of the feature map and the computational effort while retaining the main features.

[0042] The training set data is fed into the constructed CNN model. The output is calculated through forward propagation. The error between the predicted value and the true value is then calculated using a loss function. The model parameters are updated using a backpropagation algorithm, and parameters such as the convolution kernel weights are continuously adjusted to gradually enhance the model's ability to extract infrared image features. During training, the validation set is regularly used to evaluate model performance and avoid overfitting.

[0043] The feature map output by the convolutional layer retains the spatial structure information of the image and can intuitively reflect the local features of the image; the feature map output by the fully connected layer is a vector after integrating global information, which is more suitable for feature representation of tasks such as classification.

[0044] According to an embodiment of the present invention, the image features of each band of infrared images are fused based on a weighted fusion strategy, specifically including: Obtain the feature vector of the multi-band infrared image and analyze the dimension of the feature vector; Analyze the variance of each band’s eigenvector based on its dimension; Normalize the variance of each band’s feature vector and use it as the weight; The image features are weighted fused based on the weight of each band to obtain the fused image features.

[0045] During the specific calculation, the corresponding element of the eigenvector of each band is multiplied by its weight, and then the results of all bands are added together to obtain the fused eigenvector.

[0046] According to an embodiment of the present invention, the fused image features are classified based on a classifier to obtain an object recognition result, which specifically includes: Obtain the fused image features, associate the fused image features with the corresponding object category labels to obtain labeled data; Set hyperparameters and train a classifier based on labeled data to obtain training results. The classifier includes a support vector machine or a fully connected neural network. Adjust the hyperparameters based on the training results to obtain the trained classifier; Classify the fused image features based on the trained classifier to obtain the classification prediction results; The accuracy of the classification prediction results is evaluated based on the evaluation indicators to obtain the prediction accuracy; The classifier's hyperparameters are adjusted secondary based on the prediction accuracy.

[0047] It's important to verify the fused features and evaluate their performance on the target task (such as object recognition or classification). You can use the test set data to calculate metrics such as accuracy, recall, and F1 score. If performance is unsatisfactory, you can re-adjust the weight calculation method or try other fusion strategies to optimize the fusion process and improve the effect of feature fusion and model performance.

[0048] During training, use the validation set to evaluate the classifier's performance and adjust hyperparameters based on the evaluation results. For example, if the classifier's accuracy on the validation set is low, you can try adjusting hyperparameters such as the SVM penalty parameter C or the neural network's learning rate to improve performance.

[0049] A support vector machine (SVM) is a supervised learning model that classifies data by finding an optimal hyperplane. Using fused image features, the SVM can find a hyperplane in high-dimensional space that maximizes the distinction between object classes. SVMs perform particularly well when the sample size is relatively small and the feature dimensionality is high. For example, when processing multi-band infrared image features, SVMs can effectively distinguish different types of targets, such as identifying different types of equipment in military applications.

[0050] In a second aspect, an embodiment of the present application provides an object recognition system based on multi-band infrared images, the system comprising: a memory and a processor, the memory including a program for a method for object recognition based on multi-band infrared images, and the program for the method for object recognition based on multi-band infrared images, when executed by the processor, performing the following steps: Collect multi-band infrared images of target objects based on a multi-band infrared camera; Preprocessing the collected multi-band infrared image, wherein the preprocessing includes at least one of the following: denoising, registration and normalization processing, to obtain a preprocessed multi-band infrared image; Extract image features of each band of infrared images based on convolutional neural network; Based on the weighted fusion strategy, the image features of each band of infrared images are fused to obtain the fused image features; The fused image features are classified based on the classifier to obtain the object recognition results.

[0051] It should be noted that automatic feature extraction using convolutional neural networks (CNNs), such as ResNet and DenseNet, can extract deep semantic features of images layer by layer, effectively capturing complex feature information of objects and showing good adaptability to different scenes and objects.

[0052] According to an embodiment of the present invention, collecting a multi-band infrared image of a target object based on a multi-band infrared camera specifically includes: Select a multi-band infrared camera and configure camera parameters, where the camera parameters include at least one of the following: operating band, gain, and integration time; Capture scene images based on camera parameters and analyze object positions within the scene images; Determine whether the object is in the center of the scene based on the object's position; If it is not in the center of the scene, adjust the camera acquisition angle; If it is at the center of the scene, the image background is set and multi-band infrared images are collected in real time. The multi-band infrared images include infrared images of the first band, infrared images of the second band and infrared images of the third band. The wavelength of the first band is smaller than the wavelength of the second band, and the wavelength of the second band is smaller than the wavelength of the third band.

[0053] It should be noted that the wavelength setting includes determining the specific band in which the camera works, which can be selected according to the radiation characteristics of the target object and the recognition purpose, and a specific infrared band that is sensitive to the object's characteristics can be selected.

[0054] Gain adjustment involves adjusting the gain based on the radiation intensity of the scene. If the target object radiates weakly, the gain can be increased to enhance the signal. However, excessive gain can introduce noise, so proper adjustment is required.

[0055] Integration time settings, including the integration time, affect image brightness and signal-to-noise ratio. For fast-moving targets, a shorter integration time is recommended to avoid smearing. For stationary targets with weaker radiation, a longer integration time can be used to capture more signal.

[0056] Clearly identify the target object's location, ensuring it is in the center of the camera's field of view or within the effective imaging range. Use auxiliary equipment such as laser pointers to assist with positioning. Choose a suitable background to create a clear radiometric difference between the target and the background, facilitating subsequent image analysis and target recognition. For example, use a dark, uniform background to highlight the target object. Collect multiple sets of images as needed, such as from different angles and at different times, to obtain more comprehensive information about the target object.

[0057] According to an embodiment of the present invention, preprocessing of the collected multi-band infrared images specifically includes: Set the neighborhood window size, obtain the pixel values ​​of all pixels in the neighborhood window, and calculate the average value of the pixel values ​​of all pixels to obtain the pixel average value; The pixel average value is used as the pixel value of the center point of the neighborhood window to obtain the denoised multi-band infrared image; Extract feature points from multi-band infrared images and find corresponding feature point pairs between images of different bands based on feature point matching algorithm; Calculating transformation parameters based on corresponding feature point pairs, including translation parameters, rotation parameters, and scaling parameters, and performing registration transformation on the multi-band infrared images based on the transformation parameters; The infrared image of each band after registration transformation is linearly normalized to obtain the preprocessed multi-band infrared image.

[0058] It should be noted that median filtering is a nonlinear filtering method that replaces the value of each pixel in an image with the median of the pixels in its neighborhood. For multi-band infrared images, median filtering can be performed on each band separately to effectively remove salt-and-pepper noise while effectively preserving edge information. For example, for a 3×3 neighborhood window, the pixel values ​​within the window are sorted by size and the median value is used to replace the value of the center pixel in the window.

[0059] Feature point extraction algorithms include SIFT (Scale Invariant Feature Transform), SURF (Speeded Up Robust Features), etc. Feature point matching algorithms include the nearest neighbor matching algorithm based on Euclidean distance.

[0060] Linear normalization maps the pixel values ​​of an image to a fixed range, usually [0, 1] or [0, 255]. For multi-band infrared images, linear normalization is performed on each band separately.

[0061] According to an embodiment of the present invention, the image features of each band of infrared images are extracted based on a convolutional neural network, specifically including: Select the infrastructure, configure network parameters, and add target layers, including attention and pooling layers. The preprocessed multi-band infrared images are used as training sets to train the infrastructure and obtain a convolutional neural network model. Based on the convolutional neural network model, a layer that meets the preset conditions is selected as the feature extraction layer, and the infrared image of each band is input into the trained convolutional neural network model; The output is obtained based on the feature extraction layer to obtain a corresponding feature vector, which includes at least one of the following: texture features, shape features, and temperature distribution features of the multi-band infrared image.

[0062] It should be noted that the classic CNN architectures used as the foundation are LeNet, AlexNet, VGGNet, and ResNet. For example, LeNet has a simple structure and is suitable for relatively simple image feature extraction tasks; ResNet solves problems such as vanishing gradients in deep network training through residual connections and can effectively extract complex features.

[0063] Adjust the model based on the characteristics of the infrared image and the number of bands. If the infrared image resolution is low, the number of convolutional layers or the size of the convolution kernel can be appropriately reduced. For multi-band scenarios, the number of channels in the input layer can be set to correspond to the number of bands, allowing the model to process image information from multiple bands simultaneously.

[0064] To better extract infrared image features, some target layers can be added. For example, an attention mechanism layer can be added to focus the model on features in key areas of the infrared image; or a pooling layer can be added to reduce the dimension of the feature map and the computational effort while retaining the main features.

[0065] The training set data is fed into the constructed CNN model. The output is calculated through forward propagation. The error between the predicted value and the true value is then calculated using a loss function. The model parameters are updated using a backpropagation algorithm, and parameters such as the convolution kernel weights are continuously adjusted to gradually enhance the model's ability to extract infrared image features. During training, the validation set is regularly used to evaluate model performance and avoid overfitting.

[0066] The feature map output by the convolutional layer retains the spatial structure information of the image and can intuitively reflect the local features of the image; the feature map output by the fully connected layer is a vector after integrating global information, which is more suitable for feature representation of tasks such as classification.

[0067] According to an embodiment of the present invention, the image features of each band of infrared images are fused based on a weighted fusion strategy, specifically including: Obtain the feature vector of the multi-band infrared image and analyze the dimension of the feature vector; Analyze the variance of each band’s eigenvector based on its dimension; Normalize the variance of each band’s feature vector and use it as the weight; The image features are weighted fused based on the weight of each band to obtain the fused image features.

[0068] During the specific calculation, the corresponding element of the eigenvector of each band is multiplied by its weight, and then the results of all bands are added together to obtain the fused eigenvector.

[0069] According to an embodiment of the present invention, the fused image features are classified based on a classifier to obtain an object recognition result, which specifically includes: Obtain the fused image features, associate the fused image features with the corresponding object category labels to obtain labeled data; Set hyperparameters and train a classifier based on labeled data to obtain training results. The classifier includes a support vector machine or a fully connected neural network. Adjust the hyperparameters based on the training results to obtain the trained classifier; Classify the fused image features based on the trained classifier to obtain the classification prediction results; The accuracy of the classification prediction results is evaluated based on the evaluation indicators to obtain the prediction accuracy; The classifier's hyperparameters are adjusted secondary based on the prediction accuracy.

[0070] It's important to verify the fused features and evaluate their performance on the target task (such as object recognition or classification). You can use the test set data to calculate metrics such as precision, recall, and F1 score. If performance is unsatisfactory, you can re-adjust the weight calculation method or try other fusion strategies to optimize the fusion process and improve the effect of feature fusion and model performance.

[0071] During training, use the validation set to evaluate the classifier's performance and adjust hyperparameters based on the evaluation results. For example, if the classifier's accuracy on the validation set is low, you can try adjusting hyperparameters such as the SVM penalty parameter C or the neural network's learning rate to improve performance.

[0072] A support vector machine (SVM) is a supervised learning model that classifies data by finding an optimal hyperplane. Using fused image features, the SVM can find a hyperplane in high-dimensional space that maximizes the distinction between object classes. SVMs perform particularly well when the sample size is relatively small and the feature dimensionality is high. For example, when processing multi-band infrared image features, SVMs can effectively distinguish different types of targets, such as identifying different types of equipment in military applications.

[0073] A third aspect of the present invention provides a computer-readable storage medium, which includes a program for an object recognition method based on multi-band infrared images. When the program for an object recognition method based on multi-band infrared images is executed by a processor, the steps of any of the above-mentioned methods for object recognition based on multi-band infrared images are implemented.

[0074] The present invention discloses an object recognition method, system and medium based on multi-band infrared images. The method comprises the following steps: collecting multi-band infrared images of a target object by using a multi-band infrared camera; preprocessing the collected multi-band infrared images, wherein the preprocessing includes at least one of the following: denoising, registration and normalization processing to obtain preprocessed multi-band infrared images; extracting image features of each band infrared image based on a convolutional neural network; fusing the image features of each band infrared image based on a weighted fusion strategy to obtain fused image features; and classifying the fused image features based on a classifier to obtain an object recognition result. By fusing the features of the multi-band infrared images, the accuracy and adaptability of object recognition are improved.

[0075] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of units is merely a logical function division. In actual implementation, there may be other division methods, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or units can be electrical, mechanical or other forms.

[0076] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units; they may be located in one place or distributed across multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the scheme of this embodiment.

[0077] In addition, all functional units in the embodiments of the present invention may be integrated into one processing unit, or each unit may be separately used as a unit, or two or more units may be integrated into one unit; the above-mentioned integrated units may be implemented in the form of hardware or in the form of hardware plus software functional units.

[0078] Those skilled in the art will appreciate that all or part of the steps of the above-mentioned method embodiments may be implemented by hardware associated with program instructions, and the aforementioned program may be stored in a readable storage medium. When the program is executed, the program executes the steps of the above-mentioned method embodiments. The aforementioned storage medium includes various media that can store program codes, such as mobile storage devices, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical disks.

[0079] If the integrated units described above are implemented as software modules and sold or used as standalone products, they can also be stored on a readable storage medium. Based on this understanding, the technical solutions of the embodiments of the present invention, or the portion that contributes to the prior art, can be embodied in the form of a software product. This software product, stored on a storage medium, includes instructions for enabling a computer device (such as a personal computer, server, or network device) to execute all or part of the methods described in the various embodiments of the present invention. The aforementioned storage media include various media capable of storing program code, such as removable storage devices, ROM, RAM, magnetic disks, or optical disks.

Claims

1. A method for object recognition based on multi-band infrared images, characterized in that: include: Collect multi-band infrared images of target objects based on a multi-band infrared camera; Preprocessing the collected multi-band infrared image, wherein the preprocessing includes at least one of the following: denoising, registration and normalization processing, to obtain a preprocessed multi-band infrared image; Extract image features of each band of infrared images based on convolutional neural network; Based on the weighted fusion strategy, the image features of each band of infrared images are fused to obtain the fused image features; The fused image features are classified based on the classifier to obtain the object recognition results.

2. The object recognition method based on multi-band infrared images according to claim 1, characterized in that: The multi-band infrared image of the target object is collected based on the multi-band infrared camera, specifically including: Select a multi-band infrared camera and configure camera parameters, wherein the camera parameters include at least one of the following: operating band, gain, and integration time; Capture scene images based on camera parameters and analyze object positions within the scene images; Determine whether the object is in the center of the scene based on the object's position; If it is not in the center of the scene, adjust the camera acquisition angle; If it is at the center of the scene, the image background is set and a multi-band infrared image is collected in real time. The multi-band infrared image includes an infrared image of a first band, an infrared image of a second band and an infrared image of a third band. The wavelength of the first band is smaller than that of the second band, and the wavelength of the second band is smaller than that of the third band.

3. The object recognition method based on multi-band infrared images according to claim 2, characterized in that: Preprocess the collected multi-band infrared images, including: Set the neighborhood window size, obtain the pixel values ​​of all pixels in the neighborhood window, and calculate the average value of the pixel values ​​of all pixels to obtain the pixel average value; The pixel average value is used as the pixel value of the center point of the neighborhood window to obtain the denoised multi-band infrared image; Extract feature points from multi-band infrared images and find corresponding feature point pairs between images of different bands based on feature point matching algorithm; The transformation parameters are calculated based on the corresponding feature point pairs. The transformation parameters include translation parameters, rotation parameters and scaling parameters. Based on the transformation parameters, the multi-band infrared images are registered and transformed. The infrared image of each band after registration transformation is linearly normalized to obtain the preprocessed multi-band infrared image.

4. The object recognition method based on multi-band infrared images according to claim 3, characterized in that: The image features of each band of infrared images are extracted based on convolutional neural networks, including: Select the infrastructure, configure network parameters, and add target layers, including attention and pooling layers. The preprocessed multi-band infrared images are used as training sets to train the infrastructure and obtain a convolutional neural network model. Based on the convolutional neural network model, a layer that meets the preset conditions is selected as the feature extraction layer, and the infrared image of each band is input into the trained convolutional neural network model; The output is obtained based on the feature extraction layer to obtain a corresponding feature vector, which includes at least one of the following: texture features, shape features, and temperature distribution features of the multi-band infrared image.

5. The object recognition method based on multi-band infrared images according to claim 4, characterized in that: The image features of each band of infrared images are fused based on the weighted fusion strategy, specifically including: Obtain the feature vector of the multi-band infrared image and analyze the dimension of the feature vector; Analyze the variance of each band’s eigenvector based on its dimension; Normalize the variance of each band’s feature vector and use it as the weight; The image features are weighted fused based on the weight of each band to obtain the fused image features.

6. The object recognition method based on multi-band infrared images according to claim 5, characterized in that: The fused image features are classified based on the classifier to obtain the object recognition results, which include: Obtain the fused image features, associate the fused image features with the corresponding object category labels to obtain labeled data; Setting hyperparameters, training a classifier based on the labeled data, and obtaining a training result, wherein the classifier includes a support vector machine or a fully connected neural network; Adjust the hyperparameters based on the training results to obtain the trained classifier; Classify the fused image features based on the trained classifier to obtain the classification prediction results; The accuracy of the classification prediction results is evaluated based on the evaluation indicators to obtain the prediction accuracy; The classifier's hyperparameters are adjusted secondary based on the prediction accuracy.

7. An object recognition system based on multi-band infrared images, characterized in that: The system includes: a memory and a processor, wherein the memory includes a program of an object recognition method based on multi-band infrared images, and when the program of the object recognition method based on multi-band infrared images is executed by the processor, the following steps are implemented: Collect multi-band infrared images of target objects based on a multi-band infrared camera; Preprocessing the collected multi-band infrared image, wherein the preprocessing includes at least one of the following: denoising, registration and normalization processing, to obtain a preprocessed multi-band infrared image; Extract image features of each band of infrared images based on convolutional neural network; Based on the weighted fusion strategy, the image features of each band of infrared images are fused to obtain the fused image features; The fused image features are classified based on the classifier to obtain the object recognition results.

8. The object recognition system based on multi-band infrared images according to claim 7, characterized in that: The multi-band infrared image of the target object is collected based on the multi-band infrared camera, specifically including: Select a multi-band infrared camera and configure camera parameters, wherein the camera parameters include at least one of the following: operating band, gain, and integration time; Capture scene images based on camera parameters and analyze object positions within the scene images; Determine whether the object is in the center of the scene based on the object's position; If it is not in the center of the scene, adjust the camera acquisition angle; If it is at the center of the scene, the image background is set and a multi-band infrared image is collected in real time. The multi-band infrared image includes an infrared image of a first band, an infrared image of a second band and an infrared image of a third band. The wavelength of the first band is smaller than that of the second band, and the wavelength of the second band is smaller than that of the third band.

9. The object recognition system based on multi-band infrared images according to claim 8, characterized in that: Preprocess the collected multi-band infrared images, including: Set the neighborhood window size, obtain the pixel values ​​of all pixels in the neighborhood window, and calculate the average value of the pixel values ​​of all pixels to obtain the pixel average value; The pixel average value is used as the pixel value of the center point of the neighborhood window to obtain the denoised multi-band infrared image; Extract feature points from multi-band infrared images and find corresponding feature point pairs between images of different bands based on feature point matching algorithm; Calculating transformation parameters based on corresponding feature point pairs, including translation parameters, rotation parameters, and scaling parameters, and performing registration transformation on the multi-band infrared images based on the transformation parameters; The infrared image of each band after registration transformation is linearly normalized to obtain the preprocessed multi-band infrared image.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium includes an object recognition method program based on multi-band infrared images. When the object recognition method program based on multi-band infrared images is executed by a processor, the steps of the object recognition method based on multi-band infrared images as described in any one of claims 1 to 6 are implemented.