Infrared camera image processing method and system based on timing characteristics
By employing an infrared camera image processing method based on time-series features, local difference analysis and prediction verification are performed, and the simplification process is dynamically adjusted. This solves the problems of information redundancy and loss of key information in infrared camera image sequences, achieving efficient data compression and information preservation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHAANXI INST OF ZOOLOGY NORTHWEST INSTOF ENDANGERED ZOOLOGICAL SPECIES
- Filing Date
- 2026-04-17
- Publication Date
- 2026-07-07
AI Technical Summary
Existing methods for simplifying infrared camera image sequences fail to fully consider temporal correlation characteristics, resulting in excessive information redundancy or loss of key dynamic information. Furthermore, they lack quantitative evaluation of image prediction capabilities, making it difficult to achieve a balance between data compression efficiency and information integrity.
By employing an infrared camera image processing method based on temporal features, local adjacent image difference analysis and overall image difference analysis are performed. The first number of simplified images is configured, and the simplification process is dynamically adjusted through a closed-loop mechanism of image prediction and matching verification to achieve adaptive image simplification.
It improves the accuracy and intelligence of simplified processing of infrared image sequences, ensuring information integrity, while optimizing data compression efficiency and improving the extraction accuracy of key monitoring images.
Smart Images

Figure CN122066978B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, specifically to an infrared camera image processing method and system based on temporal features. Background Technology
[0002] With the rapid development of infrared thermal imaging technology, infrared cameras have been widely used in nighttime surveillance, industrial inspection, and security early warning systems. However, because infrared cameras typically require long periods of continuous operation, the amount of image sequence data they acquire is enormous, posing challenges to storage, transmission, and subsequent analysis and processing.
[0003] However, traditional image simplification methods are mainly based on the quality assessment of a single frame or simple inter-frame difference threshold judgment, which fails to fully consider the temporal correlation characteristics of the image sequence. This results in excessive redundancy in the simplified image set or loss of key dynamic information, making it difficult to achieve an effective balance between data compression efficiency and information integrity.
[0004] In existing technologies, some solutions employ fixed-interval frame extraction or keyframe extraction methods based on motion vectors. However, these methods are not well-suited to the low contrast and high noise characteristics unique to infrared images, and lack quantitative evaluation of image prediction capabilities, thus failing to ensure that the simplified image set can effectively support subsequent intelligent analysis tasks.
[0005] Furthermore, most existing methods separate image simplification from image verification, failing to form a closed-loop optimization mechanism, making it difficult to dynamically adjust the simplification strategy based on the actual prediction results. Summary of the Invention
[0006] This application provides an infrared camera image processing method and system based on time-series features, which solves the technical problem that the existing infrared image sequence simplification process lacks representativeness, resulting in inaccurate extraction of key monitoring images.
[0007] The technical solution to the above-mentioned technical problems in this application is as follows:
[0008] In a first aspect, this application provides an infrared camera image processing method based on temporal features, the method comprising:
[0009] The original image sequence acquired by the infrared camera is obtained, and local adjacent image difference analysis and overall image difference analysis are performed to obtain the local difference degree sequence and the overall image change degree. The number of first simplified images is configured according to the overall image change degree.
[0010] The first simplified image sequence is obtained by filtering based on the local difference sequence and the number of first simplified images;
[0011] Within the first simplified image sequence, multiple random combinations of image sets are selected for image prediction to obtain multiple predicted images. Multiple indexed original images are indexed within the original image sequence and matched and verified with the multiple predicted images to obtain multiple prediction accuracy coefficients. The number of second simplified images is then configured.
[0012] The image prediction accuracy coefficients of multiple first simplified images are statistically analyzed. Combined with the local difference sequence, the images are filtered according to the number of second simplified images to obtain a second simplified image set, which is used as the image processing result.
[0013] Secondly, this application provides an infrared camera image processing system based on time-series features, comprising:
[0014] The image acquisition module is used to acquire the original image sequence based on the infrared camera, perform local adjacent image difference analysis and overall image difference analysis to obtain the local difference degree sequence and the overall image change degree, and configure the first simplified image number according to the overall image change degree.
[0015] The image filtering module is used to filter and obtain a first simplified image sequence based on the local difference sequence and the number of first simplified images;
[0016] The image verification module is used to randomly select and combine image sets multiple times within the first simplified image sequence, perform image prediction, obtain multiple predicted images, index multiple indexed original images within the original image sequence, match and verify with the multiple predicted images, obtain multiple prediction accuracy coefficients, and configure the number of second simplified images.
[0017] The result acquisition module is used to statistically analyze the image prediction accuracy coefficients of multiple first simplified images, combine them with the local difference sequence, and filter them according to the number of second simplified images to obtain a second simplified image set, which is used as the image processing result.
[0018] This application provides one or more technical solutions, which have at least the following technical effects or advantages:
[0019] This application provides an infrared camera image processing method and system based on time-series features. First, it performs difference analysis on the original image sequence to obtain a local difference sequence and an overall image variability, quantitatively evaluating the dynamic changes in the image sequence and avoiding the insufficient adaptability issues caused by fixed thresholds or fixed ratios for simplification. Second, through a closed-loop mechanism of image prediction and matching verification, the prediction accuracy coefficient is used as a quantitative indicator of image representativeness. This ensures that the simplification process considers not only the difference features of the images themselves but also the role of the images in time-series prediction, effectively improving the support capability of the simplified image set for subsequent intelligent analysis tasks. Third, through a two-stage hierarchical simplification strategy, coarse screening is performed first based on local difference, followed by fine screening based on prediction accuracy, optimizing data compression efficiency while ensuring information integrity. Finally, a complete processing chain is formed from difference analysis, quantity configuration, prediction verification to result screening. Each link supports each other and dynamically adjusts, making the simplification processing of infrared image sequences more intelligent and precise, improving the extraction accuracy of key monitoring images.
[0020] Through the above technical solution, this application realizes adaptive image simplification based on temporal features, which reduces the amount of data while ensuring the information representativeness and prediction reliability of the image set. Attached Figure Description
[0021] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 This is a schematic flowchart of the infrared camera image processing method based on time-series features provided in the embodiments of this application;
[0023] Figure 2 This is a schematic diagram of the structure of an infrared camera image processing system based on time-series features provided in an embodiment of this application.
[0024] The components represented by each number in the attached diagram are explained below:
[0025] Image acquisition module 11, image filtering module 12, image verification module 13, and result acquisition module 14. Detailed Implementation
[0026] This application provides an infrared camera image processing method and system based on time-series features to address the technical problem that the screening methods used in the simplification of existing infrared image sequences are not representative enough, resulting in inaccurate extraction of key monitoring images.
[0027] Example 1, as Figure 1 As shown, this application provides an infrared camera image processing method based on time-series features, including:
[0028] S10: Obtain the original image sequence based on the infrared camera, perform local adjacent image difference analysis and overall image difference analysis to obtain the local difference degree sequence and the overall image change degree, and configure the first simplified image number according to the overall image change degree;
[0029] In this embodiment of the application, an image sequence is first acquired based on an infrared camera, and the difference between each image and the previous image is obtained by calculating the proportion of pixels with different gray values in adjacent images, and then combined to obtain a local difference sequence.
[0030] Specifically, under infrared cameras, if the ambient brightness is low, there will be more changes in image noise, which will lead to large differences in pixels between adjacent images. If the ambient brightness is high, the image quality will be stable and the differences in pixels between adjacent images will be small. Therefore, the mean of the local difference sequence is used as the overall image variability, and the number of first simplified images is configured according to the overall image variability.
[0031] This involves acquiring the original image sequence based on an infrared camera, performing local neighboring image difference analysis and overall image difference analysis to obtain the local difference degree sequence and the overall image change degree, including:
[0032] Acquire the raw image sequence based on the infrared camera, and perform preprocessing and grayscale processing;
[0033] Calculate the percentage of pixels in the original image sequence whose grayscale value difference between each original image and its adjacent original images is greater than the grayscale difference threshold, and obtain multiple local differences as a local difference sequence.
[0034] The mean of the local difference sequence is calculated as the overall image variability.
[0035] In this embodiment, firstly, based on the original image sequence acquired by the infrared camera, the original image sequence is preprocessed, including noise filtering and contrast enhancement. The multi-channel original infrared image is then converted into a single-channel grayscale image to reduce subsequent computational complexity. Specifically, median filtering is used to suppress the inherent salt-and-pepper noise and Gaussian noise in infrared images, and histogram equalization is used to improve the visibility of details in low-contrast areas.
[0036] Subsequently, each frame in the original image sequence is traversed, and its grayscale value is compared pixel-by-pixel with the adjacent previous frame. The number of pixels whose grayscale value difference exceeds a preset grayscale difference threshold is counted, and their proportion in the total number of pixels is calculated. This proportion is the local difference degree corresponding to that frame. The local difference degrees of each frame are arranged in chronological order to obtain the local difference degree sequence. The grayscale difference threshold is set considering the noise characteristics and dynamic change sensitivity of infrared images, and is usually taken as 5% to 10% of the grayscale range. For example, if the grayscale difference threshold is 10, if the grayscale value difference is greater than 10, then that grayscale value is included in the difference pixel statistics.
[0037] Finally, the arithmetic mean of the local difference sequence is calculated, and the result is the overall image variability, which characterizes the overall fluctuation characteristics of the image sequence during this period. The magnitude of the overall image variability reflects the illumination stability of the infrared camera's working environment.
[0038] Specifically, when the overall image variability is high, it indicates that the ambient lighting conditions fluctuate drastically or the infrared camera is subject to a lot of noise interference, and the differences between adjacent frames are significant. In this case, more image frames need to be retained to fully record the dynamic change process. Conversely, when the overall image variability is low, it indicates that the environment is relatively stable and the similarity between adjacent frames is high. The number of images retained can be appropriately reduced without losing key information.
[0039] Furthermore, the number of first simplified images is configured based on the overall image variability, including:
[0040] Obtain the maximum overall image change over a historical period;
[0041] The ratio of the overall image variability to the maximum overall image variability is calculated, and combined with the number of images in the original image sequence, the first simplified image number is obtained by rounding down.
[0042] In this embodiment of the application, firstly, the maximum overall image variation recorded by the infrared camera under the same working scenario is extracted from the historical database. The maximum overall image variation represents the upper limit of the inter-frame difference of the camera under extreme environmental conditions.
[0043] Subsequently, the ratio of the currently calculated overall image variability to the maximum overall image variability is calculated to obtain the relative variability coefficient. This coefficient is between 0 and 1. The closer the value is to 1, the more drastic the current environmental fluctuations are, and the closer it is to 0, the more stable the environment is. The relative variability coefficient is multiplied by the total number of images in the original image sequence, and the product is rounded up to obtain the first simplified image count.
[0044] For example, assuming the original image sequence contains 1000 frames, the historical maximum overall image change is 0.15, and the currently calculated overall image change is 0.09, then the relative change coefficient is 0.6, and the first number of simplified images is 600 frames rounded up. This configuration method achieves adaptive adjustment of the number of simplified images, which avoids excessive redundant storage when the environment is stable, and ensures the integrity of information when the environment fluctuates.
[0045] S20: Based on the local difference sequence and the number of first simplified images, a first simplified image sequence is obtained;
[0046] In this embodiment of the application, adaptive filtering is performed based on the degree of difference of each frame image in the local difference sequence. The difference values in the local difference sequence are sorted. The larger the difference, the more significant the change of the frame image relative to the previous frame. The first number of simplified image frames are selected in sequence according to the sorting result to form the first simplified image sequence.
[0047] Specifically, step S20 in the method includes:
[0048] Arrange the multiple local differences within the local difference sequence in descending order of magnitude;
[0049] The original images corresponding to the local differences of the first number of simplified images are selected and arranged to obtain the first simplified image sequence.
[0050] In this embodiment of the application, firstly, the difference values in the local difference sequence are sorted in descending order from largest to smallest to generate a sorted list with a time index. This sorted list records the local difference value of each frame of the original image and its position information in the original image sequence.
[0051] Subsequently, images are selected sequentially from the top of the sorted list until the number of selected images reaches the first number of simplified images. The original images corresponding to the time indices of each selected local difference are rearranged in chronological order to obtain the first simplified image sequence. This filtering strategy prioritizes the retention of the most drastically changing frames in the image sequence, effectively ensuring the capture of dynamic key information. At the same time, the upper limit of the number of images controls the initial data compression.
[0052] S30: Randomly select and combine image sets multiple times within the first simplified image sequence, perform image prediction, obtain multiple predicted images, and index multiple indexed original images within the original image sequence, match and verify with multiple predicted images to obtain multiple prediction accuracy coefficients, and configure the number of second simplified images.
[0053] In this embodiment of the application, several images are randomly selected multiple times within the first simplified image sequence constructed above and combined. Then, the next frame image of the last image is predicted, and a similarity analysis is performed with the image at the corresponding position in the original image sequence. The greater the similarity, the greater the prediction accuracy coefficient, and the greater the predictability of the image content. Therefore, the number of the second simplified images configured is smaller.
[0054] In this process, multiple combinations of images are randomly selected from the first simplified image sequence to perform image prediction, resulting in multiple predicted images, including:
[0055] According to a preset number of combinations, multiple sets of combined images are obtained by randomly combining a preset number of images within the first simplified image sequence in chronological order.
[0056] Call the image prediction agent;
[0057] Multiple combined image sets are input into the image prediction agent, and multiple predicted images are output.
[0058] In this embodiment of the application, firstly, a preset number of combinations is set, for example, 3. Three frames of images are randomly selected from the first simplified image sequence and arranged in chronological order to form a combined image set. This combined image set contains image information from historical moments and is used to predict the image content of future moments. The random selection operation is repeated multiple times to generate multiple different combined image sets.
[0059] Subsequently, a pre-trained image prediction agent is invoked. This agent is built on a convolutional network and has the ability to learn the spatiotemporal evolution rules from multiple historical images and generate predictions for future frames. Each combined image set is input into the image prediction agent in sequence. The agent analyzes the motion trends, target trajectories and background change features in the image sequence and outputs the corresponding prediction images. The prediction images represent the future scene information that can be inferred based on the simplified image set.
[0060] Specifically, the configuration steps of the image prediction agent include:
[0061] Based on image records from a historical time period, multiple sample image sets are collected, and the image of the next frame of the multiple sample image sets is collected as the multiple sample prediction image;
[0062] A temporal image prediction agent is constructed based on a convolutional neural network.
[0063] The training data and validation data are divided into multiple sample combination image sets and multiple sample prediction images. The image prediction agent is trained and validated under supervision using the sample combination image sets as training input and the sample prediction images as supervision. The image prediction agent is obtained after convergence.
[0064] In this embodiment, firstly, the image sequence recorded by the infrared camera under the same working scene is collected from the historical database. Continuous image frames are extracted according to a preset number of combinations to form a sample combination image set. The next frame image at the corresponding time point of the sample combination image set is collected as the sample prediction image. This sampling process is repeated to obtain a large number of training sample pairs.
[0065] Secondly, a temporal prediction model is constructed based on a convolutional neural network architecture, using an encoder-decoder structure. The encoder extracts spatial features through multiple convolutional layers, the temporal modeling module uses a long short-term memory network or a temporal convolutional network to capture the dynamic relationship between frames, and the decoder generates the prediction image through deconvolution. The model input is a sequence of historical images with a preset number of combinations, and the output is a single-frame prediction image.
[0066] Furthermore, during the training phase, the collected sample image sets and corresponding sample prediction images are divided into training and validation sets in a 7:3 ratio. The sample image sets are used as model inputs, and the sample prediction images are used as supervision signals. Mean squared error is used for supervised training. The network parameters are updated through the backpropagation algorithm. The backpropagation algorithm calculates the gradient of the loss function with respect to the weights of each layer using the chain rule. The error signal is propagated forward layer by layer from the output layer, and the network parameters of the encoder, temporal modeling module and decoder are updated sequentially.
[0067] For example, the number of nodes in the input layer equals the dimension of the input features. For instance, if the sample image set has 3 features, the input layer contains 3 nodes. One to three hidden layers are set, with the number of nodes in each layer adjusted experimentally (e.g., 64, 32). The ReLU activation function is used. The output layer generally does not use an activation function; for example, if the output takes 2 nodes, continuous values are directly output. During training, the Adam optimizer and mean squared error loss function are used to construct the training framework. The batch size is set to 32, the total training epochs to 50, and an early stopping mechanism (patience=5) is introduced. When the validation set loss does not decrease for 10 consecutive epochs, the early stopping strategy is used to prevent overfitting, save the optimal model parameters, and obtain an image prediction agent with temporal prediction capabilities.
[0068] Further, multiple original images are indexed within the original image sequence, matched and verified with multiple predicted images to obtain multiple prediction accuracy coefficients, and a second number of simplified images is configured, including:
[0069] Within the original image sequence, the images of the next frame of multiple combined image sets are indexed as multiple indexed original images;
[0070] Multiple indexed original images and multiple predicted images are input into an image verification agent constructed based on a Siamese network, and multiple image matching scores are output as multiple prediction accuracy coefficients.
[0071] Calculate the average of multiple prediction accuracy coefficients, combine them with the first number of simplified images, round down to obtain the second number of simplified images, and then configure them.
[0072] In this embodiment of the application, firstly, in the original image sequence, the next frame image is located according to the time index of each combined image set, and the frame image is extracted as the index original image. The index original image represents the real scene information at that moment and is used to compare and verify with the predicted image.
[0073] Subsequently, an image verification agent based on a Siamese network is invoked. This agent contains two feature extraction branches with shared weights, which receive the original indexed image and the predicted image as inputs, respectively. Multi-scale image features are extracted through convolutional layers, the Euclidean distance between the two feature vectors is calculated, and the image matching degree is output as the prediction accuracy coefficient. The higher the matching degree value, the closer the predicted image is to the real image, and the better the prediction accuracy.
[0074] Repeat the above verification process to perform matching verification on the predicted images corresponding to multiple combined image sets, obtain multiple prediction accuracy coefficients, and calculate the arithmetic mean of the multiple prediction accuracy coefficients to obtain the overall prediction accuracy level.
[0075] Specifically, when the overall prediction accuracy is high, it indicates that the image content at intermediate moments can be effectively inferred based on the first simplified image sequence, and the information redundancy of the image sequence is large, so the simplification ratio can be further increased; conversely, when the prediction accuracy is low, it indicates that the information density of the first simplified image sequence has approached the critical value, and the amount of further simplification should be reduced to retain more information.
[0076] Furthermore, the formula for configuring the second number of simplified images is (1 - the average of multiple prediction accuracy coefficients) × the first number of simplified images, and then rounded down to obtain the second number of simplified images.
[0077] For example, assuming the first number of simplified images is 600 frames and the average of multiple prediction accuracy coefficients is 0.75, then the second number of simplified images is (1-0.75)×600=150 frames. This configuration achieves secondary optimization of the simplification level through prediction accuracy feedback, maximizing compression efficiency while ensuring information recoverability.
[0078] The configuration steps of the image verification agent include:
[0079] Based on image records from a historical period, multiple sample image combinations are collected. Each sample image combination includes two sample images, and similarity is labeled for each sample image combination to obtain a sample image matching degree set.
[0080] Based on Siamese networks, an image verification agent is constructed, wherein the image verification agent includes the same two convolutional neural network paths;
[0081] The image verification agent is trained and verified under supervision using the combination of multiple sample images and the sample image matching degree set. The training configuration is completed after the accuracy meets the requirements.
[0082] In this embodiment, firstly, image pairs recorded by the infrared camera under the same working scene are collected from the historical database to construct a sample image combination. Specifically, this includes: selecting image pairs with a time interval less than a preset time threshold from a continuous image sequence as high similarity samples, image pairs with a time interval greater than the preset time threshold as low similarity samples, and image pairs generated by rotating, scaling, and adjusting the brightness of the same image through data augmentation as completely similar samples. The similarity of each type of sample is manually or automatically labeled to form a sample image matching degree set. The matching degree value ranges from 0 to 1, with a higher value indicating a higher degree of image similarity.
[0083] Secondly, an image verification agent is constructed based on a Siamese network architecture. This network contains two convolutional neural network branches with identical structures and shared weights. Each branch processes the input image independently and extracts multi-scale features of the image through multiple convolutional layers, batch normalization layers, and max pooling layers. The feature vectors output by the two branches are differentially processed and then input into a fully connected layer. Finally, the predicted image matching degree is output. The network design ensures symmetry, that is, swapping the order of the two input images does not affect the output result.
[0084] Furthermore, during the training phase, the collected sample images were combined and divided into training and validation sets in an 8:2 ratio. The sample image combinations were used as model inputs, and the labeled sample image matching degree was used as a supervision signal. A contrastive loss function was used for supervised training. This loss function optimizes network parameters by narrowing the feature distance between similar image pairs and widening the feature distance between dissimilar image pairs. During training, a stochastic gradient descent optimizer was used with an initial learning rate of 0.001. A learning rate decay strategy was introduced, reducing the learning rate to 0.1 times its original value every 10 training epochs. Dropout regularization was used to prevent overfitting, with a dropout probability of 0.5.
[0085] For example, the first convolutional layer of the convolutional neural network branch contains 32 3×3 convolutional kernels with a stride of 1 and padding of 1, followed by a ReLU activation function and a 2×2 max pooling layer; the second convolutional layer contains 64 3×3 convolutional kernels, configured as above; the third convolutional layer contains 128 3×3 convolutional kernels, followed by a global average pooling layer to compress the feature map into a 128-dimensional feature vector; after performing absolute difference operations on the feature vectors of the two branches, the input is a fully connected layer containing 256 nodes, and the final output is a single matching degree value. The training batch size is set to 64, and the total number of training rounds is 100. When the accuracy of the validation set does not improve for 15 consecutive rounds, an early stopping mechanism is triggered to save the optimal model parameters, resulting in an image validation agent with image similarity measurement capabilities.
[0086] S40: Calculate the image prediction accuracy coefficients of multiple first simplified images, combine them with the local difference sequence, and filter them according to the number of second simplified images to obtain a second simplified image set, which is used as the image processing result.
[0087] In this embodiment, for each frame of the first simplified image sequence, the prediction accuracy coefficient corresponding to the image as a component of the combined image set is calculated. Then, combined with the local difference, the first simplified image sequence is sorted in descending order according to the comprehensive screening index value from largest to smallest to generate a comprehensive sorting list. Images are selected sequentially from the top of the list until the number of selected images reaches the number of second simplified images. The selected images are then rearranged according to the original time order to obtain the second simplified image set.
[0088] Specifically, step S40 in the method includes:
[0089] Based on multiple prediction accuracy coefficients, the mean prediction accuracy coefficient of the combined image set corresponding to each selected first simplified image is calculated to obtain multiple image prediction accuracy coefficients.
[0090] Based on the image prediction accuracy coefficients and local dissimilarity of multiple first simplified images, the representativeness of multiple images is calculated.
[0091] The multiple images are sorted from largest to smallest representativeness, and the first simplified images corresponding to the representativeness of the second number of simplified images are selected to obtain the second simplified image set, which is used as the image processing result.
[0092] In this embodiment of the application, firstly, for each frame of the first simplified image sequence, its occurrence in multiple combined image sets is traced, the prediction accuracy coefficients corresponding to all prediction tasks in which the image participates are extracted, the arithmetic mean of the prediction accuracy coefficients is calculated, and the image prediction accuracy coefficient of the image is obtained.
[0093] Subsequently, the image representativeness is calculated based on the image prediction accuracy coefficient and the local difference, i.e., image representativeness = ((1 - image prediction accuracy coefficient) + local difference) / 2. The larger the image prediction accuracy coefficient, the more conventional the content is, and the smaller the image representativeness is. Conversely, the larger the local difference is, the larger the image representativeness is.
[0094] Furthermore, the image representativeness of all images in the first simplified image sequence is sorted in descending order of numerical value to generate a representativeness sorting list. Images are selected sequentially starting from the top of the list, and the original time index of each selected image is recorded until the number of selected images reaches the configured number of the second simplified images.
[0095] Finally, the selected images are sorted in ascending order according to their original time index to restore the chronological order, forming the final second simplified image set. This second simplified image set is output as the image processing result. It not only retains the key frames with the most significant dynamic changes in the original image sequence, but also ensures the integrity and predictability of the temporal information, realizing efficient compression and preservation of key information of infrared camera image data under the constraint of limited storage resources.
[0096] Furthermore, through the optimization of the two-level simplification strategy, the final output image set compresses the data volume to a very low proportion of the original size while ensuring that the behavior of the monitored target is traceable and the scene changes can be restored, thereby reducing the computational overhead of subsequent image storage, transmission and analysis processing.
[0097] In summary, compared with existing technologies, this application introduces a temporal prediction agent to quantitatively evaluate the predictability of image sequences, thereby realizing the transformation of the simplified strategy from static threshold screening to dynamic adaptive optimization.
[0098] In summary, the embodiments of this application have at least the following technical effects:
[0099] This application provides an infrared camera image processing method based on temporal features. First, it performs difference analysis on the original image sequence to obtain a local difference sequence and an overall image variability, quantitatively evaluating the dynamic changes in the image sequence and avoiding the insufficient adaptability issues caused by fixed thresholds or fixed ratios for simplification. Second, through a closed-loop mechanism of image prediction and matching verification, the prediction accuracy coefficient is used as a quantitative indicator of image representativeness. This ensures that the simplification process considers not only the difference features of the images themselves but also the role of the images in temporal prediction, effectively improving the support capability of the simplified image set for subsequent intelligent analysis tasks. Third, through a two-stage hierarchical simplification strategy, coarse screening is performed first based on local difference, followed by fine screening based on prediction accuracy, optimizing data compression efficiency while ensuring information integrity. Finally, a complete processing chain is formed from difference analysis, quantity configuration, prediction verification to result screening. Each link supports each other and dynamically adjusts, making the simplification processing of infrared image sequences more intelligent and precise, improving the extraction accuracy of key monitoring images.
[0100] Example 2, as Figure 2 As shown, based on the same inventive concept as the infrared camera image processing method based on temporal features provided in Embodiment 1, this application also provides an infrared camera image processing system based on temporal features, including:
[0101] Image acquisition module 11 is used to acquire the original image sequence based on the infrared camera, perform local adjacent image difference analysis and overall image difference analysis to obtain the local difference degree sequence and the overall image change degree, and configure the first simplified image number according to the overall image change degree.
[0102] Image filtering module 12 is used to filter and obtain a first simplified image sequence based on the local difference sequence and the number of first simplified images;
[0103] Image verification module 13 is used to randomly select and combine image sets multiple times within the first simplified image sequence, perform image prediction, obtain multiple predicted images, index multiple indexed original images within the original image sequence, match and verify with multiple predicted images, obtain multiple prediction accuracy coefficients, and configure the number of second simplified images.
[0104] The result acquisition module 14 is used to statistically analyze the image prediction accuracy coefficients of multiple first simplified images, combine them with the local difference sequence, and filter them according to the number of second simplified images to obtain a second simplified image set, which is used as the image processing result.
[0105] Further, in one embodiment, the process involves acquiring an original image sequence based on an infrared camera, performing local neighboring image difference analysis and overall image difference analysis to obtain a local difference sequence and an overall image variability sequence, including:
[0106] Acquire the raw image sequence based on the infrared camera, and perform preprocessing and grayscale processing;
[0107] Calculate the percentage of pixels in the original image sequence whose grayscale value difference between each original image and its adjacent original images is greater than the grayscale difference threshold, and obtain multiple local differences as a local difference sequence.
[0108] The mean of the local difference sequence is calculated as the overall image variability.
[0109] Furthermore, the number of first simplified images is configured based on the overall image variability, including:
[0110] Obtain the maximum overall image change over a historical period;
[0111] The ratio of the overall image variability to the maximum overall image variability is calculated, and combined with the number of images in the original image sequence, the first simplified image number is obtained by rounding down.
[0112] In one embodiment, the image filtering module 12 is specifically used for:
[0113] Arrange the multiple local differences within the local difference sequence in descending order of magnitude;
[0114] The original images corresponding to the local differences of the first number of simplified images are selected and arranged to obtain the first simplified image sequence.
[0115] Furthermore, in one embodiment, multiple random selections and combinations of image sets are performed within the first simplified image sequence to perform image prediction, resulting in multiple predicted images, including:
[0116] According to a preset number of combinations, multiple sets of combined images are obtained by randomly combining a preset number of images within the first simplified image sequence in chronological order.
[0117] Call the image prediction agent;
[0118] Multiple combined image sets are input into the image prediction agent, and multiple predicted images are output.
[0119] Furthermore, the configuration steps of the image prediction agent include:
[0120] Based on image records from a historical time period, multiple sample image sets are collected, and the image of the next frame of the multiple sample image sets is collected as the multiple sample prediction image;
[0121] A temporal image prediction agent is constructed based on a convolutional neural network.
[0122] The training data and validation data are divided into multiple sample combination image sets and multiple sample prediction images. The image prediction agent is trained and validated under supervision using the sample combination image sets as training input and the sample prediction images as supervision. The image prediction agent is obtained after convergence.
[0123] Further, multiple original images are indexed within the original image sequence, matched and verified with multiple predicted images to obtain multiple prediction accuracy coefficients, and a second number of simplified images is configured, including:
[0124] Within the original image sequence, the images of the next frame of multiple combined image sets are indexed as multiple indexed original images;
[0125] Multiple indexed original images and multiple predicted images are input into an image verification agent constructed based on a Siamese network, and multiple image matching scores are output as multiple prediction accuracy coefficients.
[0126] Calculate the average of multiple prediction accuracy coefficients, combine them with the first number of simplified images, round down to obtain the second number of simplified images, and then configure them.
[0127] Furthermore, in one embodiment, the configuration step of the image verification agent includes:
[0128] Based on image records from a historical period, multiple sample image combinations are collected. Each sample image combination includes two sample images, and similarity is labeled for each sample image combination to obtain a sample image matching degree set.
[0129] Based on Siamese networks, an image verification agent is constructed, wherein the image verification agent includes the same two convolutional neural network paths;
[0130] The image verification agent is trained and verified under supervision using the combination of multiple sample images and the sample image matching degree set. The training configuration is completed after the accuracy meets the requirements.
[0131] In one embodiment, the result acquisition module 14 is specifically used for:
[0132] Based on multiple prediction accuracy coefficients, the mean prediction accuracy coefficient of the combined image set corresponding to each selected first simplified image is calculated to obtain multiple image prediction accuracy coefficients.
[0133] Based on the image prediction accuracy coefficients and local dissimilarity of multiple first simplified images, the representativeness of multiple images is calculated.
[0134] The multiple images are sorted from largest to smallest representativeness, and the first simplified images corresponding to the representativeness of the second number of simplified images are selected to obtain the second simplified image set, which is used as the image processing result.
Claims
1. An infrared camera image processing method based on temporal features, characterized in that, The method includes: The original image sequence acquired by the infrared camera is obtained, and local adjacent image difference analysis and overall image difference analysis are performed to obtain the local difference degree sequence and the overall image change degree. The number of first simplified images is configured according to the overall image change degree. The first simplified image sequence is obtained by filtering based on the local difference sequence and the number of first simplified images; Within the first simplified image sequence, multiple random combinations of image sets are selected for image prediction to obtain multiple predicted images. Multiple indexed original images are then indexed within the original image sequence and matched with these predicted images to obtain multiple prediction accuracy coefficients. The number of second simplified images is then configured, including: Within the original image sequence, the images of the next frame of multiple combined image sets are indexed as multiple indexed original images; Multiple indexed original images and multiple predicted images are input into an image verification agent constructed based on a Siamese network, and multiple image matching scores are output as multiple prediction accuracy coefficients. Calculate the average of multiple prediction accuracy coefficients, combine them with the first number of simplified images, round down to obtain the second number of simplified images, and then configure them. The configuration steps of the image verification agent include: Based on image records from a historical period, multiple sample image combinations are collected. Each sample image combination includes two sample images, and similarity is labeled for each sample image combination to obtain a sample image matching degree set. Based on Siamese networks, an image verification agent is constructed, wherein the image verification agent includes the same two convolutional neural network paths; The image verification agent is trained and verified under supervision using the combination of multiple sample images and the sample image matching degree set. The training configuration is completed after the accuracy meets the requirements. The image prediction accuracy coefficients of multiple first simplified images are statistically analyzed. Combined with the local difference sequence, the images are filtered according to the number of second simplified images to obtain a second simplified image set, which is used as the image processing result.
2. The infrared camera image processing method based on time-series features according to claim 1, characterized in that, Obtain the original image sequence acquired by the infrared camera, perform local neighboring image difference analysis and global image difference analysis to obtain the local difference degree sequence and the global image variability, including: Acquire the raw image sequence based on the infrared camera, and perform preprocessing and grayscale processing; Calculate the percentage of pixels in the original image sequence whose grayscale value difference between each original image and its adjacent original images is greater than the grayscale difference threshold, and obtain multiple local differences as a local difference sequence. The mean of the local difference sequence is calculated as the overall image variability.
3. The infrared camera image processing method based on time-series features according to claim 1, characterized in that, The number of simplified images is configured based on the overall image variability, including: Obtain the maximum overall image change over a historical period; The ratio of the overall image variability to the maximum overall image variability is calculated, and combined with the number of images in the original image sequence, the first simplified image number is obtained by rounding down.
4. The infrared camera image processing method based on time-series features according to claim 1, characterized in that, Based on the local difference sequence and the number of first simplified images, a first simplified image sequence is obtained, including: Arrange the multiple local differences within the local difference sequence in descending order of magnitude; The original images corresponding to the local differences of the first number of simplified images are selected and arranged to obtain the first simplified image sequence.
5. The infrared camera image processing method based on time-series features according to claim 1, characterized in that, Within the first simplified image sequence, multiple sets of randomly selected and combined images are used for image prediction to obtain multiple predicted images, including: According to a preset number of combinations, multiple sets of combined images are obtained by randomly combining a preset number of images within the first simplified image sequence in chronological order. Call the image prediction agent; Multiple combined image sets are input into the image prediction agent, and multiple predicted images are output.
6. The infrared camera image processing method based on time-series features according to claim 5, characterized in that, The configuration steps for the image prediction agent include: Based on image records from a historical time period, multiple sample image sets are collected, and the image of the next frame of the multiple sample image sets is collected as the multiple sample prediction image; A temporal image prediction agent is constructed based on a convolutional neural network. The training data and validation data are divided into multiple sample combination image sets and multiple sample prediction images. The image prediction agent is trained and validated under supervision using the sample combination image sets as training input and the sample prediction images as supervision. The image prediction agent is obtained after convergence.
7. The infrared camera image processing method based on time-series features according to claim 1, characterized in that, The image prediction accuracy coefficients of multiple first simplified images are statistically analyzed. Combined with the local dissimilarity sequence, the images are filtered according to the number of second simplified images to obtain a second simplified image set, which serves as the image processing result, including: Based on multiple prediction accuracy coefficients, the mean prediction accuracy coefficient of the combined image set corresponding to each selected first simplified image is calculated to obtain multiple image prediction accuracy coefficients. Based on the image prediction accuracy coefficients and local dissimilarity of multiple first simplified images, the representativeness of multiple images is calculated. The multiple images are sorted from largest to smallest representativeness, and the first simplified images corresponding to the representativeness of the second number of simplified images are selected to obtain the second simplified image set, which is used as the image processing result.
8. An infrared camera image processing system based on time-series features, characterized in that, An infrared camera image processing method based on temporal features according to any one of claims 1-7, comprising: The image acquisition module is used to acquire the original image sequence based on the infrared camera, perform local adjacent image difference analysis and overall image difference analysis to obtain the local difference degree sequence and the overall image change degree, and configure the first simplified image number according to the overall image change degree. The image filtering module is used to filter and obtain a first simplified image sequence based on the local difference sequence and the number of first simplified images; The image verification module is used to randomly select and combine image sets multiple times within the first simplified image sequence, perform image prediction to obtain multiple predicted images, index multiple indexed original images within the original image sequence, match and verify with the multiple predicted images to obtain multiple prediction accuracy coefficients, and configure the number of second simplified images, including: Within the original image sequence, the images of the next frame of multiple combined image sets are indexed as multiple indexed original images; Multiple indexed original images and multiple predicted images are input into an image verification agent constructed based on a Siamese network, and multiple image matching scores are output as multiple prediction accuracy coefficients. Calculate the average of multiple prediction accuracy coefficients, combine them with the first number of simplified images, round down to obtain the second number of simplified images, and then configure them. The configuration steps of the image verification agent include: Based on image records from a historical period, multiple sample image combinations are collected. Each sample image combination includes two sample images, and similarity is labeled for each sample image combination to obtain a sample image matching degree set. Based on Siamese networks, an image verification agent is constructed, wherein the image verification agent includes the same two convolutional neural network paths; The image verification agent is trained and verified under supervision using the combination of multiple sample images and the sample image matching degree set. The training configuration is completed after the accuracy meets the requirements. The result acquisition module is used to statistically analyze the image prediction accuracy coefficients of multiple first simplified images, combine them with the local difference sequence, and filter them according to the number of second simplified images to obtain a second simplified image set, which is used as the image processing result.