A spatio-temporal domain-based infrared image processing method and related device

By constructing a unified background image and performing temporal analysis, combined with a kinematic model, the problem of repairing split-screen infrared images in dynamic scenes was solved, achieving temporal continuity and integrity of infrared image sequences, and improving the accuracy and efficiency of image repair.

CN121032861BActive Publication Date: 2026-03-24BEIJING DONGYU HONGDA TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-28
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

Existing technologies suffer from screen splitting issues when capturing infrared images of moving objects, due to the splicing of signals from different times. This causes delayed data misalignment to span multiple normal frames, and existing algorithms cannot correctly identify the misalignment pattern, leading to repair failure or image distortion.

Method used

A unified background image is constructed as a spatiotemporal reference. The complete trajectory of the moving object is determined by combining the temporal sequence. After splitting the screen image, background matching and key point extraction are performed separately. The current frame and the delayed frame are judged by the expected position distance between the key points and the trajectory prediction. The edge seamless fusion algorithm is used for stitching.

Benefits of technology

It effectively solves the problem of the inability to correctly identify misalignment patterns in existing technologies, realizes accurate repair of split-screen images in dynamic scenes, ensures the temporal continuity and integrity of infrared image sequences, and improves the reliability of frame judgment and the accuracy of background area registration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121032861B_ABST
    Figure CN121032861B_ABST
Patent Text Reader

Abstract

The application provides an infrared image processing method based on a space-time field and related equipment, and relates to the technical field of infrared image processing. The method comprises the following steps: acquiring continuous infrared images, constructing a unified background image, identifying key points of each target moving object and mapping the key points, determining an image sequence group according to a time sequence, and obtaining a complete motion trajectory curve according to the distribution of the key points. When a split-screen target infrared image is detected, the key points of each part are extracted after splitting, the two parts of the background are matched with the unified background image to determine an initial background area, the distance between the key points and the expected position predicted based on the trajectory, adjacent frames and motion rules is calculated, the current frame and the delayed frame are determined, and finally the complete sequence is obtained by processing the delayed frame. The method in the application can accurately repair the split-screen image in a dynamic scene, and ensure the time sequence coherence and integrity of the infrared image sequence.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of infrared image processing, and in particular to an infrared image processing method based on a space-time domain and related equipment. BACKGROUND

[0002] In the field of infrared imaging, sensors generate images by collecting infrared radiation emitted or reflected by objects. After the advent of the digital era, data integrity is mainly guaranteed through verification algorithms, and the recovery of damaged data mainly relies on redundant information. In the past, limited computing resources made it difficult to apply image feature-based repair methods, but with the advancement of hardware technology, using image inherent features for repair has become a feasible and practical solution.

[0003] Prior art such as CN112258427B proposes a repair method for infrared images, which identifies screen splitting features in the infrared image, such as target straight lines in the horizontal and / or vertical directions, and uses image stitching to repair the image. This technology effectively solves the problem of repairing damaged infrared images with misplaced data in static scenes, making damaged images that could only be discarded have value again.

[0004] However, when shooting infrared images of objects in motion, there may be screen splitting caused by different time signal splicing, such as multiple processing units processing different frames at the same time. If the synchronization mechanism between processing units fails, it may cause early processing unit data to be written into later frames incorrectly due to delay. When the misplacement spans multiple normal frames, existing algorithms may not be able to correctly identify this misplacement pattern, resulting in repair failure or more severe image distortion. SUMMARY

[0005] The present application provides an infrared image processing method based on a space-time domain and related equipment, which is used to solve the problem of repair failure or image distortion caused by incorrect identification of misplacement patterns when different time signal splicing screen splitting occurs due to multi-processing unit synchronization failure, etc. (including early delayed data written into later frames and misplacement spanning multiple normal frames) in dynamic scenes.

[0006] In a first aspect, this application provides a spatiotemporal domain-based infrared image processing method applied to a server. The method includes: after acquiring multiple consecutive infrared images, preprocessing to extract and construct an overall unified background image, simultaneously identifying the key point coordinates of a moving target object in each frame of the infrared image, mapping and marking the key point coordinates on the corresponding positions in the unified background image; determining an infrared image sequence group according to the time series, and determining the complete motion trajectory curve of the moving target object based on the distribution of key point coordinates on the unified background image; when a target infrared image containing split-screen features is detected, dividing the image into a first part and a second part based on horizontal or vertical boundary line features, and extracting the key point coordinates of the moving target object in each part; and then dividing the first part into a second part. Feature point matching and spatial registration are performed on the background regions of one part of the image and the second part of the image with a unified background map to determine the first initial background region and the second initial background region of the two parts in the unified background map. The distance between the coordinates of key points in the first initial background region and the expected position in the unified background is calculated. The image with the smaller distance is determined to be the current frame image that conforms to the current time sequence, and the image with the larger distance is determined to be the delayed frame image. The expected position is predicted by time series analysis and kinematic model based on the complete motion trajectory curve, combined with the coordinates of key points in the adjacent infrared images of the target infrared image and the motion law of the target moving object. Infrared image processing is performed on the delayed frame image to obtain a complete infrared image sequence.

[0007] By adopting the above technical solution, a unified background image is constructed as a spatiotemporal reference. The complete trajectory of the moving object is determined by combining the time sequence. After splitting the screen image, background matching and key point extraction are performed separately. The distance between the key point and the expected position predicted by the trajectory is used to determine the current frame and the delayed frame. Therefore, the technical means of effectively solving the problem that the existing technology cannot correctly identify the splicing of time signals across multiple normal frames in the presence of moving objects (such as the delay data misalignment caused by the failure of the processing unit synchronization) is not possible. Thus, the technical effect of accurately repairing the screen image in dynamic scenes and ensuring the temporal continuity and integrity of the infrared image sequence is achieved.

[0008] In conjunction with some embodiments of the first aspect, in some embodiments, the infrared image processing based on the delayed frame image to obtain a complete infrared image sequence specifically includes: for the delayed frame image, determining the original time point corresponding to the delayed frame image by finding the best matching position of the corresponding key point coordinates on the complete motion trajectory curve on a unified background image, and retrieving the original infrared image corresponding to the original time point; determining the similarity between the delayed frame image and the original infrared image, and marking the region with a similarity value exceeding a preset threshold as an overlapping region; accurately removing the overlapping region from the delayed frame image, and stitching the remaining non-overlapping part with the original infrared image using an edge seamless fusion algorithm to obtain a complete and temporally consistent infrared image sequence.

[0009] By employing the above technical solution, timing alignment is ensured through original time point positioning, similarity judgment guarantees the accuracy of overlapping areas, and seamless fusion eliminates splicing artifacts. These steps work synergistically to correct timing deviations in delayed frames while maintaining the integrity and visual coherence of image content, resulting in a processed sequence that more accurately reflects the actual motion process.

[0010] In conjunction with some embodiments of the first aspect, in some embodiments, before the step of dividing the image into a first part image and a second part image based on horizontal or vertical boundary line features and extracting the coordinates of key points of the target moving object in the images respectively, the method further includes: if there are no key points in the first part image and the second part image, then redetermining new key points based on features common to the target moving object in the first part image and the second part image; and redetermining the complete motion trajectory curve based on the new key points.

[0011] By adopting the above technical solution, the supplementary steps avoid the problem of trajectory interruption due to the loss of key points caused by screen splitting. By continuing the key point recognition logic through shared features, the trajectory curve can continuously and accurately reflect the motion state, ensuring that subsequent frame judgment and processing based on the trajectory have a reliable basis, and enhancing the robustness of the solution in dealing with complex screen splitting scenarios.

[0012] In some embodiments of the first aspect, the calculation of the distance between the key point coordinates in the first and second initial background regions and the expected position specifically includes: extracting the key point coordinate sequences of the first N frames and the last M frames of the target infrared image based on the complete motion trajectory curve and the determined temporal relationship, where N and M are preset positive integers; applying a kinematic model to the key point coordinate sequence for curve fitting to obtain the predicted motion direction and velocity parameters of the target moving object at the target timestamp corresponding to the target infrared image; calculating the theoretical coordinate position of the target moving object at the target timestamp on the complete motion trajectory curve based on the predicted motion direction and velocity parameters and the target timestamp, and determining the theoretical coordinate position as the expected position.

[0013] By employing the above technical solution, when calculating the expected position, a sequence of key points from multiple preceding and following frames is first extracted. Then, a kinematic model is used to fit the predicted direction and velocity, and the theoretical coordinates are determined by combining the timestamps. The preceding and following frame data provide the basis for the motion trend, the kinematic model quantifies the motion patterns, and the timestamps ensure accurate correspondence in the time dimension. The combination of these three elements makes the calculated expected position more closely resemble the actual physical characteristics of motion, improves the accuracy of comparison with the actual coordinates of key points, and thus more accurately distinguishes between the current frame and delayed frames, enhancing the reliability of the judgment.

[0014] In conjunction with some embodiments of the first aspect, in some embodiments, the background regions of the first part image and the second part image are respectively matched with feature points and spatially registered with a unified background image to determine the first initial background region and the second initial background region of the two parts in the unified background image. Specifically, this includes: removing feature points within the target moving object region in the first part image and the second part image, and retaining the feature points of the background region; matching the background region feature points of the first part image and the second part image with the background of all infrared images preceding the target infrared image in the infrared image sequence group; and determining the first initial background region and the second initial background region based on the background region with the highest similarity.

[0015] By adopting the above technical solution, during background registration, feature points of the target region are first removed while retaining background features. Then, the region is matched with historical images, and the initial region is determined based on similarity. Removing target feature points avoids interference from moving objects in background matching, while historical image matching utilizes temporal continuity to provide a reference benchmark, and the principle of maximizing similarity ensures matching accuracy. These steps synergistically improve the accuracy of background region registration, providing a more reliable spatial positioning basis for subsequent frame judgment and reducing positioning errors in split-screen processing.

[0016] In conjunction with some embodiments of the first aspect, in some embodiments, when a target infrared image containing split-screen features is detected, before dividing the image into a first part image and a second part image based on horizontal or vertical boundary line features, the method further includes: detecting the degree of abrupt changes in brightness and texture row by row or column by column in the horizontal and vertical directions of the infrared image, and using boundary lines with a degree of abrupt changes greater than a preset threshold as candidate boundary lines; comparing the similarity between the regions on both sides of the candidate boundary lines in the infrared image and the corresponding regions in the preceding and following frames, and when the similarity between one side region and the infrared image adjacent to the current frame time is greater than a preset similarity threshold while the similarity between the other side and other infrared images is greater than a preset similarity threshold, then it is confirmed that a split-screen feature exists, and the other infrared images refer to images in the infrared image sequence group other than the infrared image adjacent to the current frame time.

[0017] By employing the above technical solutions, brightness and texture abrupt change detection quickly locates potential split-screen boundaries, while similarity comparison verifies the split-screen characteristics from a temporal perspective, eliminating misjudgments based solely on texture abrupt changes. The combination of these two steps forms a complementary verification, significantly improving the accuracy of split-screen feature recognition, avoiding misjudging normal images as split-screen images, and ensuring the targeted and efficient processing flow.

[0018] In conjunction with some embodiments of the first aspect, in some embodiments, after performing infrared image processing based on the delayed frame image to obtain a complete infrared image sequence, the method further includes: determining the number of normal images between the current frame image and the delayed frame image from multiple infrared image processing history records; determining the number of normal images that appears most frequently as the number of common error intervals; and after determining a new target infrared image with split-screen features, preferentially matching the corresponding first part image and second part image with the infrared image before the number of common error intervals to obtain the corresponding initial background region.

[0019] By employing the above technical solution, the number of frequently error-prone intervals is determined using historical records. During the processing of new split-screen images, the initial region is matched preferentially according to this interval. Historical data statistics reveal common error patterns, and the priority matching strategy applies these patterns to new scenes, reducing the number of blind matches. This iterative optimization mechanism enables subsequent split-screen processing to find the accurate initial region more quickly, improving processing efficiency.

[0020] In a second aspect, this application provides a server comprising: one or more processors and a memory; the memory being coupled to the one or more processors, the memory being used to store computer program code including computer instructions, the one or more processors invoking the computer instructions to cause the server to perform the methods described in the first aspect and any possible implementation thereof.

[0021] Thirdly, this application provides a computer-readable storage medium including instructions that, when executed on a server, cause the server to perform the method described in the first aspect and any possible implementation thereof.

[0022] Fourthly, this application provides a computer program product that, when run on a server, causes the server to perform the method described in the first aspect and any possible implementation thereof.

[0023] One or more technical solutions provided in the embodiments of this application have at least the following technical effects or advantages:

[0024] 1. By using a unified background image as a spatiotemporal reference, combining time sequence to determine the complete trajectory of moving objects, splitting the screen image and performing background matching and key point extraction separately, and judging the current frame and delayed frame by the distance between the key points and the expected position predicted by the trajectory, the technology effectively solves the problem that existing technologies cannot correctly identify the splicing of time signals across multiple normal frames in scenes with moving objects. Thus, it achieves the technical effect of accurately repairing split images in dynamic scenes and ensuring the temporal continuity and integrity of infrared image sequences.

[0025] 2. By employing the technique of extracting key point sequences from multiple frames of the target's infrared image, obtaining predicted motion direction and velocity parameters through kinematic model fitting, and combining the timestamp to calculate the theoretical coordinates of the expected position, the problem of large judgment errors between the current frame and delayed frames caused by inaccurate calculation of the expected position in existing technologies is effectively solved. This results in improved frame judgment reliability and makes the expected position more consistent with the actual motion state of the target.

[0026] 3. By employing techniques to remove feature points of moving objects in split-screen images while retaining background feature points, and then matching the background with historical infrared images to determine the initial background region based on similarity, the problem of large registration errors in the background region of split-screen images caused by moving object interference in existing technologies is effectively solved. This results in improved background region registration accuracy and provides a reliable spatial positioning basis for frame judgment. Attached Figure Description

[0027] Figure 1 This is a schematic flowchart of an infrared image processing method based on the spatiotemporal domain in an embodiment of this application;

[0028] Figure 2 This is another schematic flowchart of the spatiotemporal domain-based infrared image processing method in the embodiments of this application;

[0029] Figure 3 This is a schematic diagram of the physical device structure of a server in an embodiment of this application. Detailed Implementation

[0030] The terminology used in the following embodiments of this application is for the purpose of describing particular embodiments only and is not intended to be limiting of this application. As used in the specification and appended claims of this application, the singular expressions “a,” “an,” “the,” “the,” “the,” and “this” are intended to include the plural expressions as well, unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in this application refers to and includes any or all possible combinations of one or more of the listed items.

[0031] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as implying or suggesting relative importance or implicitly indicating the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature, and in the description of the embodiments of this application, unless otherwise stated, "multiple" means two or more.

[0032] For ease of understanding, the method provided in this implementation is described in process below. Please refer to [link / reference]. Figure 1This is a flowchart illustrating an infrared image processing method based on the spatiotemporal domain in an embodiment of this application.

[0033] S101. After acquiring multiple consecutive infrared images, a unified background image is extracted and constructed through preprocessing. At the same time, the key point coordinates of the moving object in each frame of the infrared image are identified, and the key point coordinates are mapped and marked on the corresponding positions in the unified background image.

[0034] Among them, "infrared image" refers to an image formed by capturing thermal radiation information through infrared imaging equipment, used for observing targets in environments with insufficient visible light; "continuous infrared image" refers to a sequence of multiple infrared images acquired sequentially in time, forming a temporally coherent image data stream; "preprocessing" refers to basic processing operations such as noise reduction, enhancement, and normalization on the original infrared images to improve image quality and analyzability; "overall unified background map" refers to a static scene image extracted from multiple infrared images that does not contain moving targets, used as a reference benchmark for subsequent analysis; "target moving object" refers to a heat source target whose position changes in the infrared image sequence, distinguished from the static background; "key point coordinates" refers to the spatial coordinates of important location points that can characterize the features of the target moving object, used to track and analyze the target's motion state; "mapping" refers to the process of transforming the key point coordinates in each frame of the image to the unified background map coordinate system; "marking" refers to the operation of displaying or recording the key point positions in a recognizable manner on the unified background map.

[0035] After acquiring multiple consecutive infrared images from infrared sensor arrays, infrared cameras, or other infrared imaging devices, the server first performs basic preprocessing on these images. Preprocessing includes, but is not limited to, image denoising, contrast enhancement, size normalization, and grayscale normalization to improve image quality and lay the foundation for subsequent processing. The server then uses background modeling algorithms to extract static background information from the multiple frames of infrared images, constructing a unified background image that does not contain moving targets. This process typically employs background modeling methods such as temporal median filtering and Gaussian mixture models. Through statistical analysis of each pixel over time, stable and unchanging pixels are classified as background, gradually building a complete background model.

[0036] Meanwhile, the server performs target detection and keypoint extraction operations on each frame of infrared image. The target detection process first uses background subtraction to compare the current frame with a uniform background image at the pixel level, extracting potential foreground target regions. Then, it combines thresholding, connected component analysis, and morphological processing to further optimize the foreground region boundaries, removing noise and small false detection areas. After determining the target region, the server applies feature point detection algorithms (such as SIFT, SURF, ORB, or deep learning feature extraction networks) to detect stable and discriminative keypoints within the target region. These keypoints typically include inflection points on the target contour, high-contrast regions inside the target, and parts with distinct shape features, effectively representing the target's position and pose information. For each detected keypoint, the server records its precise coordinates (x, y) in the current image coordinate system, along with necessary descriptive information (such as keypoint type, response intensity, and main direction).

[0037] After extracting key points from a single frame, the server needs to transform the coordinates of these key points from the coordinate system of their respective images to the coordinate system of a unified background image. This mapping process first determines the geometric transformation relationship (such as translation, rotation, scaling, etc.) between the current frame and the unified background image, typically achieved through feature matching and transformation matrix estimation. The server uses background region feature points in the current frame to match with the unified background image, calculates precise coordinate transformation parameters, and then applies these parameters to transform the target key point coordinates to the unified background image coordinate system. The transformed key point coordinates are then marked on the corresponding positions in the unified background image. The marking can take the form of dots of a specific color, numbered markers, or other easily identifiable visual elements. Each key point marker usually also includes a timestamp for subsequent time-series-based analysis. This marking method allows key points in images from different time points and different frames to be compared and analyzed in the same coordinate system, providing a foundation for establishing the target motion trajectory.

[0038] S102. Determine the infrared image sequence group according to the time series, and determine the complete motion trajectory curve of the target moving object based on the key point coordinate distribution on the unified background map.

[0039] Among them, "infrared image sequence group" refers to a set of infrared images with temporal continuity and logical correlation; "key point coordinate distribution" represents the spatial distribution of all key points in a unified background image coordinate system, reflecting the spatial information of the target's motion; "complete motion trajectory curve" refers to the complete motion path of the target object in the spatiotemporal dimension, usually represented in the form of a curve, used to describe the motion law and characteristics of the target.

[0040] After completing the background image construction and key point marking in the previous step, the server first needs to organize and manage the acquired infrared images according to time sequence. This process begins by extracting timestamp information from the infrared image metadata, including the year, month, day, hour, minute, second, and even millisecond-level precision of the time stamp. If the image metadata does not contain time information or the time information is unreliable, the server will estimate the timestamp based on the image reception order and system time allocation. The server then sorts all images according to the timestamps, forming an image sequence arranged strictly in chronological order.

[0041] After time-series sorting, the server further divides these images into different infrared image sequence groups. Sequence grouping can be based on various rules, including but not limited to: fixed time windows (e.g., images every 10 seconds form a group), scene change points (new groups are created when significant scene changes are detected), and target appearance and disappearance points (groups are divided based on target entry and exit from the monitoring area). In practical applications, the server typically employs a hybrid strategy for sequence grouping, considering both specific task requirements and scene characteristics. For example, in continuous target tracking tasks, a sequence group might be formed based on the target's complete lifecycle (from appearance to disappearance); while in scene monitoring tasks, it might be divided according to fixed time windows or scene changes. The server assigns a unique identifier to each sequence group and establishes logical relationships between images within the sequence group. These relationships facilitate subsequent time-series analysis and anomaly detection.

[0042] After the infrared image sequence is determined, the server processes the coordinates of marked key points on the unified background image, analyzing the spatial distribution patterns of these points to determine the complete motion trajectory curve of the target moving object. First, the server extracts the three-dimensional information (x, y, t) of each key point, where x and y are the spatial coordinates of the key point on the unified background image, and t is the timestamp of the corresponding image. The server then classifies and clusters these data points, grouping key points belonging to the same target together. After clustering, each cluster represents a potential target motion trajectory. Finally, the server generates a continuous and smooth motion trajectory curve for each target, accurately reflecting the target's movement path and speed changes within the monitored area.

[0043] S103. When a target infrared image containing split-screen features is detected, the image is divided into a first part image and a second part image based on the horizontal or vertical dividing line features, and the key point coordinates of the target moving object in the image are extracted respectively.

[0044] Among them, "split-screen feature" refers to the characteristic of image content from different time periods appearing simultaneously in the same frame in an infrared image. It is usually manifested as a clear horizontal or vertical dividing line in the image, with time differences in the image content on both sides of the dividing line; "target infrared image" refers to a specific infrared image that has been detected to contain split-screen features and requires special processing; "horizontal or vertical dividing line feature" refers to the dividing line feature in the split-screen phenomenon, which may be a clear boundary line in the horizontal or vertical direction; "first part image" and "second part image" respectively represent the image regions on both sides of the dividing line, and these two parts may come from image content at different time points; "target moving object" refers to a heat source target that changes position in the infrared image; "key point coordinates" refers to the spatial coordinates of important position points that can characterize the features of the target moving object.

[0045] In processing infrared image sequences, it is necessary to identify and handle potential screen splitting issues. Screen splitting refers to the simultaneous inclusion of image content from different points in time within a single image frame. In this application, this phenomenon typically stems from factors such as signal transmission delays in the imaging device, frame buffer read / write errors, or hardware malfunctions. When the server detects that a frame of infrared image exhibits screen splitting characteristics, this processing step is triggered.

[0046] Once the presence of screen-splitting features in the target infrared image is confirmed, the server begins image segmentation based on the boundary line features. First, the server needs to accurately locate the boundary line. This process typically employs a combination of image processing techniques, including but not limited to: edge detection to identify strong edges in the image; row / column brightness analysis to calculate the average brightness value of each row or column and identify abrupt changes; texture discontinuity analysis to detect abrupt changes in texture features using texture description operators; and image structural similarity analysis to calculate the structural similarity index between local regions of the image and identify the boundaries of similarity abrupt changes. The server then integrates these features, using weighted voting or probabilistic fusion to determine the most probable boundary line location.

[0047] After determining the boundary line, the server divides the image into a first image and a second image. The division typically follows these principles: if the boundary line is horizontal, the area above it is defined as the first image, and the area below it as the second image; if the boundary line is vertical, the area to its left is defined as the first image, and the area to its right as the second image. During the division process, the server considers the ambiguity of the boundary line and transition areas, potentially excluding narrow bands (typically a few pixels wide) near the boundary line from both images to avoid using unreliable data in transition areas. The two divided images are treated as independent image content from different points in time and require separate processing.

[0048] The server then performs target moving object detection and key point extraction on the first and second image parts respectively. This process is similar to the key point extraction in step S101, but the following points require special attention: image incompleteness caused by screen splitting. Since each image part only contains a portion of the original complete field of view, the server needs to adjust the parameters of the target detection algorithm to adapt to the characteristics of the partial images; the target may be truncated at the dividing line, resulting in an incomplete target shape, and the server needs to implement robust partial target detection capabilities; the two images come from different time points, and the same target may appear simultaneously in both parts but in different positions, and the server needs to distinguish this situation to avoid mistaking it for two different targets.

[0049] The server applies an object detection algorithm to identify target regions in each image segment, and then extracts keypoints within those regions. Keypoint extraction can be based on geometric features (such as contour inflection points and centroids), texture features, or semantic features (such as specific parts of the target). For each detected keypoint, the server records its precise position (x, y) in the current image segment's coordinate system, along with necessary additional information (such as keypoint type and confidence level). Considering the special case caused by screen splitting, the server also records the image segment to which each keypoint belongs (first or second segment), information crucial for subsequent time-series analysis.

[0050] In some embodiments, prior to this step, if there are no keypoints in the first and second partial images, new keypoints are redefined based on common features of the moving target object in the first and second partial images; the complete motion trajectory curve is then redefined based on these new keypoints. Here, the new keypoints represent target feature location points redefined by analyzing common features, used to replace previously undetected keypoints. When processing infrared images of targets containing split-screen features, the server may encounter situations where no predefined keypoints are successfully detected in either the first or second partial image, or where only one partial image has a keypoint but the other partial image does not. To address this issue, the server needs to implement an alternative solution: redefined keypoints based on common features of the moving target object.

[0051] First, the server comprehensively analyzes the image content in both the first and second image sets, attempting to identify the presence of the same moving object in both sets. This identification is no longer limited to predefined keypoints but extends to a broader set of features, including the overall contour of the target, thermal imaging feature distribution, texture patterns, shape features, and contrast with the background. The server employs various image processing and analysis techniques, such as edge detection, region segmentation, hotspot analysis, texture analysis, and morphological processing, to extract feature information of the moving object from different perspectives. During feature extraction, the server pays particular attention to common features that can be stably identified in both the first and second image sets; these common features provide a reliable basis for determining the consistency of the target. Once the common features are identified, the server redefines new keypoints based on these features. The principle for defining new keypoints is to select feature points that can be stably identified under different poses and viewpoints and accurately reflect the target's position and posture. For example, for human targets, even if predefined skeletal joints cannot be detected, the server can define new keypoints such as the head center and torso center based on thermal imaging features; for vehicle targets, new keypoints such as the vehicle body center and the midpoints of the front and rear boundaries can be defined based on contour features. When defining new keypoints, the server considers their spatial distribution to ensure they effectively describe the target's overall position and pose. The server assigns each new keypoint a clear semantic meaning and physical location interpretation for comparison and correspondence with keypoints in other images during subsequent processing. After redefining the new keypoints, the server needs to reconstruct the complete motion trajectory curve based on them. This process first requires the server to back-analyze previously processed infrared image sequences, extracting corresponding new keypoints from these images according to the newly defined criteria to ensure consistency of keypoints throughout the sequence. Then, the server arranges these new keypoints in a time series and applies appropriate kinematic models and curve fitting algorithms to reconstruct the complete motion trajectory curve of the moving target.

[0052] S104. The background regions of the first part image and the second part image are respectively matched with the unified background image by feature point matching and spatial registration, and the first initial background region and the second initial background region of the two parts in the unified background image are respectively determined.

[0053] Here, "background region" refers to the static part of the image that does not contain the moving target object, and is the main region for feature matching; "feature point matching" refers to the process of finding corresponding feature point pairs by comparing the feature point descriptors in two images; "spatial registration" refers to the process of aligning two images to the same coordinate system through geometric transformation; "unified background image" refers to the overall scene background image constructed in step S101, which serves as a spatial reference benchmark; "first initial background region" and "second initial background region" respectively represent the background regions corresponding to the first part of the image and the second part of the image in the unified background image, which are used for subsequent analysis.

[0054] After dividing the screen images and extracting key points, it is necessary to determine the precise spatial positions of these two image parts within a unified background image. This is a crucial step in determining their temporal attribution. First, the server needs to separate the background region from the first and second image parts. Background region separation is typically based on foreground-background segmentation techniques, combined with the target moving object region information identified in step S103. The server uses a masking operation to mask the identified target regions (usually represented by bounding boxes or outlines) in the original image, retaining the non-target regions as the background.

[0055] Once the background region is determined, the server begins extracting feature points from that region. These feature points will be used for matching against a unified background image. Feature point extraction typically employs local feature description algorithms, which identify well-discriminative and repetitive local features in the image and generate feature vectors (descriptors) describing these features. For the characteristics of infrared images, the server may need to adjust the parameters of the feature extraction algorithm, such as lowering the contrast threshold or increasing the number of feature points, to adapt to the low contrast and limited detail of infrared images. For the background feature point sets extracted from the first and second parts of the image, the server then matches them against feature points in the unified background image. The matching process first requires extracting feature points and descriptors from the unified background image as well. Since the unified background image is typically a complete scene image, while the first and second parts of the image are only partial scenes, the server faces a part-to-whole matching problem. To improve matching efficiency and accuracy, the server may use a spatial index structure to organize the feature points in the unified background image, accelerating the nearest neighbor search process. For each feature point to be matched, the server uses a nearest neighbor search algorithm to find the point with the smallest descriptor distance in the feature point set of the unified background image as a matching candidate.

[0056] After the initial matching is completed, the server enters the spatial registration stage, which aims to determine the precise location and extent of the two image parts within a unified background image. Spatial registration first requires estimating a geometric transformation model based on the matching point pairs. Depending on the characteristics of the infrared imaging equipment and the application scenario, this transformation model may be a similarity transformation (including translation, rotation, and uniform scaling), an affine transformation (allowing non-uniform scaling and shearing), or a more complex perspective transformation (handling changes in viewpoint). The server uses the RANSAC (Random Sample Consensus) algorithm to remove outliers (false matches) from the matching point pairs, while simultaneously estimating the optimal transformation parameters. The core idea of ​​the RANSAC algorithm is to find the model that best explains the data through multiple random samplings. In specific implementation, the server first randomly selects a minimum number of samples from all matching point pairs (e.g., 3 pairs for affine transformation, 4 pairs for perspective transformation), and calculates a preliminary transformation model based on these samples. Then, the server transforms all matching points through this model and calculates the error between the transformed points and the target points (usually measured using Euclidean distance). If the error is less than a preset threshold (e.g., 2-5 pixels), the point is considered an interior point (a point that conforms to the model); otherwise, it is considered an exterior point (an outlier). The server repeats this process hundreds of times (the exact number depends on the required success probability and the expected proportion of inliers), ultimately selecting the model with the highest number or proportion of inliers as the optimal model. To further improve accuracy, the server typically recalculates the final transformation model based on all inliers.

[0057] After the transformation model is determined, the server applies these transformation parameters to map the first and second parts of the image onto the coordinate system of the unified background image, respectively, determining their precise positions and coverage areas within the unified background image, i.e., the first and second initial background regions. These two regions are typically represented as rectangles or quadrilaterals, containing the coordinates of the four corner points of the transformed image boundary, fully describing the spatial position and coverage area of ​​each part of the image within the unified background image. During spatial registration, the server calculates and records registration quality metrics, such as the inlier ratio (the proportion of matching points conforming to the transformation model out of the total matching points, typically expected to be greater than 60%) and the average reprojection error (the average distance between the transformed point and the ideal position, typically expected to be less than 2 pixels). Ultimately, the server accurately determines the first and second initial background regions, which accurately reflect the spatial position and range of the two parts of the split-screen image within the unified background image, laying the foundation for subsequent temporal analysis and image reconstruction. This region information includes not only spatial position but also registration quality metrics and transformation parameters, facilitating the evaluation of the reliability of the results and making necessary adjustments during subsequent processing.

[0058] S105. Calculate the distance between the key point coordinates in the first initial background region and the second initial background region and the expected position in the unified background. The smaller distance is determined to be the current frame image that conforms to the current time sequence, and the larger distance is determined to be the delayed frame image. The expected position is based on the complete motion trajectory curve, combined with the key point coordinates in the adjacent infrared images of the target infrared image and the motion law of the target moving object, and is predicted by time series analysis and kinematic model.

[0059] The term "first initial background region" refers to the corresponding position of the background region in the first image portion within the unified background image, representing the location of a portion of the split-screen image within the overall background. "Second initial background region" refers to the corresponding position of the background region in the second image portion within the unified background image, representing the location of the other portion of the split-screen image within the overall background. "Expected position" refers to the theoretical position of the target object at the current time point, predicted based on the motion trajectory curve and kinematic model, representing a coordinate position conforming to the laws of physical motion. "Complete motion trajectory curve" represents the continuous motion path of the target moving object throughout the entire image sequence, representing the positional relationship of the object over time. "Adjacent infrared images" refer to image frames adjacent to the target infrared image in the time series, representing temporally continuous image data. "Current frame image" represents the portion of the image that corresponds to the current moment in the time series, representing a time-correct image. "Delayed frame image" represents the portion of the historical image incorrectly written to the current frame due to processing unit synchronization failure, representing a time-incorrect image.

[0060] After determining the first and second initial background regions, the core step of split-screen image timing judgment begins. The goal of this step is to determine which part of the image corresponds to the normal display at the current moment (the current frame image) and which part is content from an earlier moment but displayed with a delay (the delayed frame image). In this step, the server needs to transform the keypoint coordinates of the moving objects in the first and second part images to the unified background image coordinate system and then compare them with the expected positions. This is a coordinate system transformation and distance calculation process that requires meaningful comparisons within the same coordinate system. First, the server does not directly use the keypoint coordinates in the first and second part images; instead, it needs to map these local coordinates to the global coordinate system of the unified background image. Specifically, the server uses the "first initial background region" and "second initial background region" determined in the previous steps as references to establish a mapping relationship from the local coordinate system of the split-screen image to the global coordinate system of the unified background image. For example, suppose the keypoint coordinates of a human head are detected in the first part image as (x1, y1), and the first part image is determined to correspond to a certain region in the unified background image through feature matching, and there exists a transformation matrix M1 that maps the first part image to this region. The server applies the transformation matrix M1 to the keypoint coordinates (x1, y1) to obtain the keypoint's global coordinates (X1, Y1) in the unified background image. Similarly, for the keypoint coordinates (x2, y2) in the second image, the server applies the transformation matrix M2 to obtain its global coordinates (X2, Y2) in the unified background image. In this way, the server successfully transforms the local keypoint coordinates in both images to the same reference system—the global coordinate system of the unified background image—creating the conditions for subsequent distance calculations.

[0061] Next, the server needs to calculate the expected position. The expected position is the theoretical position that the target object should appear at the corresponding time point in the target infrared image with a split screen, based on the complete motion trajectory curve and kinematic model prediction. This will be described in detail in steps S201-S203 and will not be repeated here. Next, the server calculates the distance D1 between the converted global coordinates (X1, Y1) of the first part of the keypoints and the expected position (Xe, Ye), and the distance D2 between the global coordinates (X2, Y2) of the second part of the keypoints and the expected position (Xe, Ye). These distance calculations typically use Euclidean distance:

[0062] After calculating the distance value, the server compares the magnitudes of D1 and D2. The part of the image with the smaller distance (e.g., D1 < D2) is determined to be the current frame image that conforms to the current time sequence because the target position in it is closer to the expected position; while the part with the larger distance (e.g., D2 > D1) is determined to be the delayed frame image. The server finally obtains an expected position coordinate (Xe, Ye), which represents the position where the target object should theoretically appear at the current moment, and this position is also represented in the unified background image coordinate system.

[0063] Because when there is a split screen problem, a part of the image may be from the current moment, while another part may be from an earlier moment (delayed frame). The object position in the delayed frame reflects the state at a certain past moment, rather than the state at the current moment. Since the object is in continuous motion, there will be an obvious gap between the object position in the delayed frame and the expected position at the current moment, and the size of this gap depends on the length of the delay time and the motion speed of the object. Therefore, through this distance comparison method based on the continuity of physical motion, the system can reliably determine which part of the image belongs to the current frame and which part belongs to the delayed frame.

[0064] S106. Perform infrared image processing on the delayed frame image to obtain a complete infrared image sequence.

[0065] Among them, the "delayed frame image" refers to the part of the image in the target infrared image that is determined to have a time sequence inconsistent with the current frame, and it refers to the part of the historical image that is wrongly written into the current frame due to the synchronization failure of the processing unit.

[0066] After determining the timing of the split-screen images and identifying which part is the current frame and which is the delayed frame, the infrared images need to be systematically processed based on the identified delayed frame information to restore a complete and time-consistent image sequence. First, the server needs to determine the original time point corresponding to the delayed frame image; this is a crucial prerequisite for the entire processing. Using the complete motion trajectory curve constructed in the previous steps, the server searches for the position on the trajectory that best matches the coordinates of the key points of the moving object in the delayed frame image. Specifically, the server calculates the Euclidean or Mahalanobis distance between the key point coordinates in the delayed frame image (already transformed to a unified background coordinate system) and the key point coordinates at each time point on the complete motion trajectory curve, selecting the time point with the smallest distance as the original time point of the delayed frame image. After determining the original time point, the server retrieves the original infrared image corresponding to that original time point from the infrared image sequence group. Next, the server needs to determine the precise correspondence between the delayed frame image and the original infrared image and identify overlapping areas. The server uses a local feature matching algorithm to extract feature points and descriptors from the two images. Then, the server establishes the correspondence between the feature points of the two images through feature descriptor matching. Based on these correspondences, the server can estimate the geometric transformation between two images, such as affine transformation or homography transformation.

[0067] Based on feature matching, the server calculates the similarity between corresponding regions. Similarity calculation can be based on various metrics, such as normalized cross-correlation, structural similarity index, or feature space distance. The server sets an appropriate similarity threshold and marks regions with similarity exceeding the threshold as overlapping regions. Accurate identification of overlapping regions is crucial for subsequent image fusion, as it determines which information is redundant and needs to be removed, and which information is complementary and needs to be retained.

[0068] After identifying overlapping regions, the server precisely removes these regions from the delayed frame image. The removal process needs to consider smoothing the region boundaries to avoid generating hard edges. After removing the overlapping regions, the remaining non-overlapping parts of the delayed frame image contain information that may be missing or incomplete in the original infrared image; this information needs to be appropriately fused back into the original infrared image.

[0069] The server uses an edge-seamless fusion algorithm to stitch the non-overlapping portions of the delayed frame image with the original infrared image. This process considers not only geometric alignment but also the continuity of brightness, contrast, and texture to ensure that the stitched image appears visually natural and coherent.

[0070] Because there may be varying degrees of overlap between the delayed frame image and the original infrared image:

[0071] Partial overlap: A portion of the content in the delayed frame image overlaps with the original infrared image, while another portion contains information not present in the original image;

[0072] Complete overlap: The delayed frame image is completely contained within the original infrared image, without containing any new information. By calculating the similarity of local regions, the system can flexibly identify various overlap patterns, accurately locating the boundaries of overlapping regions, whether partial or complete.

[0073] Modern infrared imaging systems typically consist of multiple processing units, each responsible for processing different parts or frames of an image. When the synchronization mechanism between these processing units fails: a processing unit may delay data processing; delayed data may be incorrectly written into later frames; this results in the final image containing data from different points in time. Through similarity analysis, the system can accurately identify these varying degrees of overlap, providing precise region segmentation for subsequent image fusion, thereby achieving high-quality image restoration. This method not only solves simple split-screen problems but also handles multi-frame delays and aliasing in complex scenes.

[0074] In this embodiment, since a unified background image is constructed and a complete motion trajectory curve is determined, the current frame and delayed frame in the split-screen image can be accurately determined based on spatiotemporal domain analysis. This effectively solves the problem of repairing split-screen errors in infrared images in dynamic scenes, thereby achieving high-quality reconstruction and temporal consistency maintenance of infrared image sequences containing moving objects.

[0075] In some embodiments, after step S106, the number of normal images between the current frame image and the delayed frame image can be determined from multiple infrared image processing history records; the number of normal images that appears most frequently is determined as the number of common error intervals; after a new target infrared image with split-screen characteristics is determined next time, the corresponding first part image and second part image are preferentially matched with the infrared image before the number of common error intervals to obtain the corresponding initial background region. Here, the number of normal images represents the number of infrared image frames that should exist in chronological order between the current frame image and the delayed frame image, used to represent the interval between two frames under normal timing conditions. The number of common error intervals represents the fixed number of delayed frames that lag behind the current frame in the most frequent split-screen phenomenon, used to represent the inherent pattern of delay phenomena in the system. The new target infrared image represents a newly detected infrared image containing split-screen characteristics, referring to the latest image with temporal anomaly characteristics that needs to be processed.

[0076] Specifically, during the continuous processing of infrared image sequences, the server gradually accumulates a large amount of image processing history. These records contain information about the current frame image and delayed frame image determined each time the split-screen feature images are processed, as well as the temporal relationship between them. Step S501 uses these historical records to identify possible fixed delay patterns in the system and utilizes these patterns to improve processing efficiency in subsequent processing. First, the server needs to determine the number of normal images between the current frame image and delayed frame image from multiple infrared image processing history records. In this process, the server systematically analyzes the results of each split-screen feature image processing and extracts image pairs that have been clearly identified as the current frame image and delayed frame image. For each such image pair, the server queries their index position or timestamp in the original infrared image sequence and calculates the number of normal image frames that should exist between them. For example, if the index of the current frame image is i and the index of the delayed frame image is ik, then the number of normal images is k-1. The server performs this calculation for each split-screen processing result and records the results in a dedicated statistical data structure. During the calculation process, the server considers factors such as the frame rate of the image acquisition device and system processing latency to ensure that the calculated number of normal images accurately reflects the temporal relationship. As the system runs longer, this statistical data structure accumulates more and more samples, providing a rich data foundation for subsequent analysis. Next, the server analyzes the statistical distribution of these normal image counts and determines the most frequent count as the common error interval.

[0077] Specifically, the server performs frequency statistics on the number of normal images in the statistical data structure, calculating the frequency of each different value and identifying the most frequent value. This most frequent value is determined as the most likely fixed latency pattern in the system, i.e., the number of common error intervals. This statistical analysis is based on an important assumption: under specific hardware and software environments, the latency behavior of a system often exhibits certain regularities, and the most common latency pattern usually reflects the inherent characteristics of the system. After determining the number of common error intervals, the server applies this information to subsequent image processing, especially when a new split-screen feature image is detected. When a new target infrared image with split-screen features is identified again, the server no longer tries to match all possible historical images from scratch. Instead, it prioritizes matching the first and second parts of the image with the infrared image before the specified number of common error intervals. Specifically, if the current image being processed is an infrared image with index j and the number of common error intervals is m, the server will prioritize attempting to match the first and second parts of the split-screen image with the infrared image with index jm-1 for background matching and feature point comparison. This targeted priority matching strategy greatly improves processing efficiency and reduces unnecessary computation. During the actual matching process, the server employs multiple image features and matching algorithms to ensure accuracy. If a primary match fails, the server gradually expands the search range, but still prioritizes intervals close to the number of frequently failed intervals. In this way, the server can quickly obtain the initial background regions corresponding to each part of the split-screen image, laying the foundation for subsequent temporal analysis and image reconstruction. This historical statistical optimization strategy fully utilizes the regularity of system latency, significantly improving the efficiency and accuracy of processing split-screen feature images, especially when processing large sequences of continuous infrared images.

[0078] Based on the above, the following is a more detailed description of the process provided in this implementation. Please refer to [link / reference]. Figure 2 This is another flowchart illustrating the spatiotemporal domain-based infrared image processing method in this application embodiment.

[0079] S201. Based on the complete motion trajectory curve and the determined temporal relationship, extract the key point coordinate sequence of the first N frames and the last M frames of the target infrared image, where N and M are preset positive integers.

[0080] Among them, "complete motion trajectory curve" represents the continuous motion path of the target moving object in the entire infrared image sequence, used to represent the positional relationship of the object as time changes; "temporal relationship" refers to the arrangement of each frame of the infrared image sequence according to time order, used to represent the temporal order between images; "target infrared image" represents the infrared image containing split-screen features to be processed, used to represent the specific image that needs to be analyzed and repaired; "previous N frames" refers to the N consecutive image frames that are located before the target infrared image in the time series, used to represent the historical image data of the target image; "last M frames" refers to the M consecutive image frames that are located after the target infrared image in the time series, used to represent the future image data of the target image.

[0081] After determining the complete motion trajectory curve and detecting the screen-split feature in the target infrared image, in practical applications, relying solely on the information from the target infrared image itself is often insufficient to accurately predict the expected position of the moving object, especially when the target image has screen-split features, as the motion information it contains may have been corrupted or misaligned. Therefore, the server needs to utilize consecutive frames before and after the target image to provide a more complete motion context. Based on the constructed complete motion trajectory curve and the determined temporal relationship, the server locates the position of the target infrared image within the entire time series and then determines the time window range to be extracted.

[0082] First, the server needs to determine appropriate N and M values, i.e., the number of frames extracted forward and backward. The selection of these values ​​requires balancing several factors: on the one hand, larger N and M values ​​provide richer historical and future information, helping to capture complex motion patterns, especially for nonlinear motion and speed changes; on the other hand, excessively large values ​​may introduce irrelevant data, increase computational complexity, and may even include other interfering factors. In practical implementation, the values ​​of N and M are usually dynamically adjusted based on factors such as the characteristics of the moving target, the infrared image acquisition frequency, and the system's processing capabilities. For example, for fast-moving targets, smaller N and M values ​​(e.g., N=5, M=3) may be needed to avoid significant changes in motion patterns; while for slowly changing targets, larger values ​​(e.g., N=15, M=10) can be chosen to obtain more stable motion trends.

[0083] After determining the time window, the server retrieves the first N frames and last M frames of the target infrared image from the infrared image sequence group in chronological order. Under normal circumstances, this process is relatively straightforward; the server can locate the target based on the image's timestamp or sequence index. For each retrieved frame, the server needs to extract the coordinates of key points of the moving object. These key points are usually identified and mapped to a uniform background image in previous steps. The server extracts the key point coordinates for the corresponding time point from the uniform background image, ensuring that all coordinates are in the same reference coordinate system for subsequent motion analysis.

[0084] After extracting keypoints from the first N frames and the last M frames, the server organizes these coordinates into an ordered sequence according to time. For motion in a two-dimensional plane, each keypoint typically contains (x, y) coordinates; for motion in three-dimensional space, each keypoint contains (x, y, z) coordinates. Furthermore, the server associates each keypoint with its corresponding timestamp, forming a complete spatiotemporal sequence of data. This keypoint coordinate sequence serves as the foundation for subsequent kinematic modeling and prediction.

[0085] It is worth noting that for complex targets containing multiple keypoints (such as the human body), the server may need to process multiple keypoint sequences. In this case, the server can select the most representative keypoint (such as the center of gravity of the human body) for primary analysis, or model multiple keypoints separately and then perform a comprehensive analysis. The motion characteristics of different keypoints may differ, and comprehensive analysis helps to obtain a more comprehensive understanding of motion. Finally, the server outputs a time-ordered sequence of keypoint coordinates, containing motion information from the previous N frames and the following M frames of the target's infrared image.

[0086] S202. Apply a kinematic model to the key point coordinate sequence to perform curve fitting, and obtain the predicted motion direction and velocity parameters of the target moving object at the target timestamp corresponding to the target infrared image.

[0087] After extracting the keypoint coordinate sequence, the extracted spatiotemporal data is analyzed and fitted using an appropriate kinematic model to predict the motion state of the target moving object at a specific moment. The server utilizes mathematical modeling and parameter estimation techniques to extract motion patterns from discrete observation data and construct a mathematical model that can describe and predict the motion behavior of the target object.

[0088] In practical applications of infrared image processing, moving targets may exhibit various complex motion patterns, such as uniform linear motion, accelerated motion, curvilinear motion, or a combination of these. The server needs to select a kinematic model suitable for the target's actual motion characteristics and determine the model parameters through a fitting algorithm to accurately predict the target's motion state at a given timestamp. This kinematic model-based prediction method not only considers the target's historical position information but also incorporates physical motion laws, effectively handling complex situations such as changes in acceleration and direction during motion.

[0089] First, the server needs to select a suitable kinematic model based on the characteristics of the moving target object and the distribution features of the key point coordinate sequence. Common kinematic models include uniform linear motion models, suitable for targets that move at approximately uniform linear speeds for short periods; uniformly accelerated motion models, suitable for targets with significant acceleration or deceleration; and higher-order polynomial models, suitable for complex nonlinear motion.

[0090] When selecting a model, the server considers several factors: First, the physical characteristics of the target, such as human movement which usually involves acceleration and deceleration, while a vehicle may be approximately at a constant speed for a short period of time; second, the quality and quantity of the observation data, where a simple model should be chosen to avoid overfitting when there are few data points; and third, the time span of the prediction, where a local linear model can be used for short-term predictions, while a more complex nonlinear model needs to be considered for long-term predictions.

[0091] Based on the characteristics of the target moving object, the server selects an appropriate kinematic model type. In practical applications, the server may simultaneously construct multiple candidate models and determine the most suitable model for the current target moving object's characteristics through cross-validation. For most scenarios, the server uses a polynomial fitting method, treating time as the independent variable and the x and y coordinates as dependent variables, constructing two independent polynomial functions x(t) and y(t). The order of the polynomial is dynamically adjusted according to the complexity of the motion; simple linear motion may use polynomials of order 1-2, while complex nonlinear motion may require polynomials of order 3-5. The server optimizes the polynomial coefficients using the least squares method to minimize the mean square error between the fitted curve and the actual observation points. During the fitting process, the server assigns higher weights to coordinate points with more recent time, ensuring that the prediction is closer to the latest state of the target moving object. After fitting, the server evaluates the fitting quality, calculating indicators such as the average error, maximum error, and variance between the fitted curve and the actual observation points. If the fitting quality is unsatisfactory, the server attempts to adjust the model type or parameters and refit. Once a satisfactory fitting curve is obtained, the server differentiates the fitting functions x(t) and y(t) at the target timestamp t0 to obtain the instantaneous velocity vector (dx / dt|t=t0, dy / dt|t=t0), where the direction of the vector is the predicted motion direction, and the magnitude of the vector is the predicted velocity magnitude. The server further analyzes the acceleration characteristics of the target moving object, using the second derivatives (d²x / dt²|t=t0, d²y / dt²|t=t0) to assess the acceleration or deceleration trend of the motion, as well as the curvature change of the trajectory. Finally, the server outputs a set of parameters containing the predicted motion direction angle θ and the velocity magnitude v, which will be used in subsequent steps to calculate the theoretical coordinate position of the target moving object at the target timestamp.

[0092] S203. Based on the predicted motion direction and velocity parameters, and combined with the target timestamp, calculate the theoretical coordinate position of the target moving object at the target timestamp on the complete motion trajectory curve, and determine the theoretical coordinate position as the expected position.

[0093] After obtaining the predicted motion direction and velocity parameters in step S202, the theoretical coordinate position of the moving object at the target timepoint needs to be calculated on the complete motion trajectory curve based on these parameters and the target timepoint. The server selects the known keypoint coordinates that are closest in time to the target timepoint as the starting reference point, denoted as (x_ref, y_ref, t_ref). Then, the server calculates the time difference Δt = t_target - t_ref between the target timepoint t_target and the reference timepoint t_ref. This time difference will be used for subsequent position extrapolation calculations. The server selects an appropriate calculation method based on the type of kinematic model. For a simple uniform linear motion model, the server uses vector extrapolation, that is, based on the predicted velocity vector v = (v_x, v_y) and the time difference Δt, it calculates the displacement vector Δs = v·Δt, and then superimposes the displacement vector onto the reference point coordinates to obtain the theoretical coordinate position (x_theory, y_theory) = (x_ref + Δs_x, y_ref + Δs_y). After calculating the theoretical coordinates, the server verifies whether the position falls on the complete motion trajectory curve or within its reasonable vicinity. If the deviation is too large, the server re-evaluates the motion model parameters or adjusts the calculation method. Finally, the server determines the verified theoretical coordinates (x_theory, y_theory) as the expected position, which will be used as an important reference benchmark for judging the current frame image and the delayed frame image in subsequent steps.

[0094] In this embodiment, since the key point coordinate sequence is extracted from the N frames before and M frames after the target infrared image and the kinematic model is applied for curve fitting, and the theoretical coordinate position of the target timestamp is calculated by combining the complete motion trajectory curve, the position and motion state of the target moving object at any time point can be accurately predicted. This effectively solves the problem that traditional methods cannot accurately distinguish between the current frame and the delayed frame when processing infrared images containing split-screen features. As a result, intelligent identification and automatic correction of spatiotemporal anomalies in the infrared image sequence are realized, which greatly improves the temporal accuracy and spatial consistency of infrared image processing.

[0095] The server in the embodiments of this invention is described below from the perspective of hardware processing. Please refer to [link / reference]. Figure 3 This is a schematic diagram of the physical device structure of a server in an embodiment of this application.

[0096] It should be noted that, Figure 3 The server structure shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of the present invention.

[0097] like Figure 3As shown, the server includes a Central Processing Unit (CPU) 301, which can perform various appropriate actions and processes based on a program stored in Read-Only Memory (ROM) 302 or a program loaded from storage portion 308 into Random Access Memory (RAM) 303, such as performing the methods described in the above embodiments. The RAM 303 also stores various programs and data required for system operation. The CPU 301, ROM 302, and RAM 303 are interconnected via a bus 304. An Input / Output (I / O) interface 305 is also connected to the bus 304.

[0098] The following components are connected to I / O interface 305: input section 306 including audio input devices, push-button switches, etc.; output section 307 including a liquid crystal display (LCD) and audio output devices, indicator lights, etc.; storage section 308 including a hard disk, etc.; and communication section 309 including a network interface card such as a LAN (Local Area Network) card, modem, etc. Communication section 309 performs communication processing via a network such as the Internet. Drive 310 is also connected to I / O interface 305 as needed. Removable media 311, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., are installed on drive 310 as needed so that computer programs read from them can be installed into storage section 308 as needed.

[0099] In particular, according to embodiments of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present invention include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing computer programs for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 309, and / or installed from removable medium 311. When the computer program is executed by central processing unit (CPU) 301, it performs the various functions defined in the present invention.

[0100] It should be noted that specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this invention, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0101] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. Each block in a flowchart or block diagram may represent a module, program segment, or portion of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those shown in the drawings.

[0102] Specifically, the server in this embodiment includes a processor and a memory. The memory stores a computer program. When the computer program is executed by the processor, it implements the spatiotemporal domain-based infrared image processing method provided in the above embodiment.

[0103] In another aspect, the present invention also provides a computer-readable storage medium, which may be included in the server described in the above embodiments; or it may exist independently and not assembled into the server. The storage medium carries one or more computer programs that, when executed by a processor of the server, cause the server to implement the spatiotemporal domain-based infrared image processing method provided in the above embodiments.

[0104] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit it. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.

[0105] As used in the above embodiments, depending on the context, the term "when..." can be interpreted as meaning "if...", "after...", "in response to determining...", or "in response to detecting...". Similarly, depending on the context, the phrase "when determining..." or "if (the stated condition or event) is interpreted as meaning "if determining...", "in response to determining...", "when (the stated condition or event) is detected", or "in response to detecting (the stated condition or event)".

[0106] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This program can be stored in a computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM or random access memory (RAM), magnetic disks, or optical disks.

Claims

1. A spatiotemporal domain-based infrared image processing method, applied to a server, characterized in that, The method includes: After acquiring multiple consecutive infrared images, a unified background image is extracted and constructed through preprocessing. At the same time, the key point coordinates of the moving object in each frame of the infrared image are identified, and the key point coordinates are mapped and marked on the corresponding positions in the unified background image. Infrared image sequence groups are determined according to time series, and the complete motion trajectory curve of the target moving object is determined based on the key point coordinate distribution on the unified background image. When a target infrared image containing split-screen features is detected, the image is divided into a first part and a second part based on the horizontal or vertical dividing line features, and the key point coordinates of the moving object in the image are extracted respectively. The background regions of the first and second parts of the image are respectively matched with the unified background image by feature point matching and spatial registration, and the first and second initial background regions of the two parts in the unified background image are respectively determined. Calculate the distance between the key point coordinates in the first and second initial background regions and the expected position in a unified background. The smaller distance is determined to be the current frame image that conforms to the current time sequence, and the larger distance is determined to be the delayed frame image. The expected position is based on the complete motion trajectory curve, combined with the key point coordinates in the adjacent infrared images of the target infrared image and the motion law of the target moving object, and is predicted by time series analysis and kinematic model. Infrared image processing is performed on the delayed frame image to obtain a complete infrared image sequence.

2. The method according to claim 1, characterized in that, The step of performing infrared image processing based on the delayed frame image to obtain a complete infrared image sequence specifically includes: For the delayed frame image, the original time point corresponding to the delayed frame image is determined by finding the best matching position of the corresponding key point coordinates on the complete motion trajectory curve on the unified background image, and the original infrared image corresponding to the original time point is retrieved. Determine the similarity between the delayed frame image and the original infrared image, and mark the regions with similarity values ​​exceeding a preset threshold as overlapping regions; The overlapping areas are precisely removed from the delayed frame image, and the remaining non-overlapping parts are stitched together with the original infrared image using an edge seamless fusion algorithm to obtain a complete and time-consistent infrared image sequence.

3. The method according to claim 1, characterized in that, Before the step of dividing the image into a first part and a second part based on horizontal or vertical boundary line features, and extracting the key point coordinates of the moving object in the image respectively, the method further includes: If there are no key points in the first part of the image and the second part of the image, new key points are determined based on the common features of the moving object in the first part of the image and the second part of the image. The complete motion trajectory curve is redefined based on the new key points.

4. The method according to claim 1, characterized in that, The calculation of the distance between the coordinates of key points in the first and second initial background regions and the expected locations specifically includes: Based on the complete motion trajectory curve and the determined temporal relationship, extract the key point coordinate sequence of the first N frames and the last M frames of the target infrared image, where N and M are preset positive integers; By applying a kinematic model to the key point coordinate sequence for curve fitting, the predicted motion direction and velocity parameters of the target moving object at the target timestamp corresponding to the target infrared image are obtained. Based on the predicted motion direction and velocity parameters, and combined with the target timestamp, the theoretical coordinate position of the target moving object at the target timestamp is calculated on the complete motion trajectory curve, and this theoretical coordinate position is determined as the expected position.

5. The method according to claim 1, characterized in that, The step of performing feature point matching and spatial registration between the background regions of the first and second parts of the image and a unified background image, respectively, to determine the first and second initial background regions of the two parts in the unified background image, specifically includes: Remove feature points within the target moving object region from the first and second part of the image, while retaining feature points in the background region; The background region feature points of the first part of the image and the second part of the image are respectively matched with the background of all infrared images preceding the target infrared image in the infrared image sequence group; The first and second initial background regions are determined based on the background regions with the highest similarity.

6. The method according to claim 1, characterized in that, Before the step of dividing the target infrared image into a first part and a second part based on horizontal or vertical boundary line features when a split-screen feature is detected, the method further includes: The degree of abrupt changes in brightness and texture is detected row by row or column by column in the horizontal and vertical directions of the infrared image, and the boundary line with the degree of abrupt change greater than a preset threshold is used as the candidate boundary line. Compare the similarity between the regions on both sides of the candidate boundary line in the infrared image and the corresponding regions in the preceding and following frames. If the similarity between one side of the region and the infrared image adjacent to the current frame in time is greater than a preset similarity threshold, while the similarity between the other side and other infrared images is greater than a preset similarity threshold, then the presence of screen splitting features is confirmed. The other infrared images refer to images in the infrared image sequence group other than the infrared image adjacent to the current frame in time.

7. The method according to claim 1, characterized in that, After the step of performing infrared image processing on the delayed frame image to obtain a complete infrared image sequence, the method further includes: Determine the number of normal images between the current frame image and delayed frame images from multiple infrared image processing history records; The number of times the normal image appears most frequently is determined as the number of common error intervals; After a new target infrared image with split-screen characteristics is identified, the corresponding first part image and second part image are preferentially matched with the infrared image before the specified number of error intervals to obtain the corresponding initial background area.

8. A server, characterized in that, The server includes: one or more processors and a memory; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code including computer instructions, and the one or more processors call the computer instructions to cause the server to perform the method as described in any one of claims 1-7.

9. A computer-readable storage medium comprising instructions, characterized in that, When the instructions are executed on the server, the server causes the server to perform the method as described in any one of claims 1-7.

10. A computer program product, characterized in that, When the computer program product is run on the server, the server performs the method as described in any one of claims 1-7.

Citation Information

Patent Citations

  • A method and apparatus for restoring infrared images

    CN112258427B

  • Multi-target vehicle trajectory extraction method based on pixel-level image fusion

    CN115457080A

  • Target object motion trail processing method and system based on video data

    CN117474959A