Tibetan antelope behavior analysis system based on multi-modal image fusion and optimized transmission method

Through dual feature modeling of temperature texture synergistic anomaly index and spatial dynamic anomaly index, combined with polynomial regression model, the problem of misjudgment of non-target high-temperature objects in infrared images is solved, and high-precision and anti-interference recognition of Tibetan antelope behavior analysis are achieved, which improves the system's resource scheduling efficiency and intelligence level.

CN120356244AActive Publication Date: 2025-07-22SHAANXI INST OF ZOOLOGY NORTHWEST INSTOF ENDANGERED ZOOLOGICAL SPECIES

Patent Information

Application Number
CN202510846853.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-07-22
Estimated Expiration
2045-06-24

AI Technical Summary

Technical Problem

In the existing multimodal image fusion technology, non-target high-temperature objects in infrared images are easily misjudged as Tibetan antelope targets, resulting in weakening or loss of animal characteristics, affecting the accuracy of behavior recognition and monitoring.

Method used

The dual-eigen modeling method of temperature texture synergistic anomaly index and spatial dynamic abnormality index is adopted, combined with polynomial regression model, and the high-temperature areas in infrared images are marked in grades and optimized transmission processing to reduce interference effects.

Benefits of technology

The fidelity and behavior recognition accuracy of target animal characteristics in the fusion image are improved, and the system's resource scheduling efficiency and intelligence level in complex environments are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356244A_ABST
    Figure CN120356244A_ABST
Patent Text Reader

Abstract

The invention discloses a Tibetan antelope behavior analysis system based on multi-modal image fusion and an optimized transmission method, and belongs to the technical field of image fusion. Through synchronously acquiring an infrared image and a visible light image, performing registration, denoising and normalization processing on the infrared image and the visible light image, and extracting a temperature texture collaborative anomaly index and a spatial dynamic transaction index, the optimal transmission of the Tibetan antelope behavior analysis system is realized; constructing a polynomial regression model for identifying a non-target heat source, and outputting a high-temperature non-target object analysis value; the system divides a high-temperature area in the infrared image into a credible target level, an uncertain level and an interference level, and calculates a fusion interference risk value for the uncertain area in combination with the response strength and weight distribution of an image fusion algorithm; finally, Tibetan antelope behavior recognition is executed based on the fused image, the recognition confidence coefficient is dynamically adjusted in combination with the interference risk, and high-precision and anti-interference recognition of Tibetan antelope real behaviors is achieved. According to the method, the behavior recognition accuracy and ecological monitoring intelligence level of multi-modal image fusion in a complex environment are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image fusion, and particularly to a Tibetan antelope behavior analysis system for multimodal image fusion and an optimized transmission method. Background Art

[0002] Tibetan antelope behavior analysis refers to the systematic study and analysis of the behavior characteristics and activity patterns of Tibetan antelopes in their natural environment. Such analysis usually includes their migration routes, foraging habits, reproductive behaviors, social structures, and responses to external disturbances (such as human activities or climate change), aiming to reveal the ecological adaptation mechanisms and survival strategies of Tibetan antelopes and provide a scientific basis for species protection and ecological management.

[0003] The existing technologies have the following deficiencies: During the multimodal image fusion process, non-target high-temperature objects often appear in infrared images, such as bare rocks heated by sunlight in the afternoon of summer. These areas appear bright in thermal imaging and are extremely likely to be misjudged as target animals such as Tibetan antelopes by fusion algorithms (especially models based on attention mechanisms), thus wrongly guiding the allocation of fusion weights. As a result, the true animal features may be weakened, obscured, or even completely lost in the fused image, seriously interfering with the accuracy of behavior recognition and monitoring. Summary of the Invention

[0004] The purpose of the present invention is to provide a Tibetan antelope behavior analysis system for multimodal image fusion and an optimized transmission method to solve the deficiencies in the background art.

[0005] To achieve the above purpose, the present invention provides the following technical solutions: A Tibetan antelope behavior analysis system for multimodal image fusion, including an image acquisition module, an image preprocessing module, a feature extraction module, an interference analysis model construction module, and a grading module: The image acquisition module synchronously acquires multimodal image data including infrared images and visible light images; The image preprocessing module preprocesses the multimodal image data; The feature extraction module extracts temperature texture collaborative features and spatial dynamic anomaly features from the infrared images; The interference analysis model construction module trains a machine learning model based on the temperature texture collaborative features and spatial dynamic anomaly features, and outputs the analysis values of high-temperature non-target objects corresponding to each high-temperature area; The grading module classifies and marks the high-temperature areas in the infrared images according to the analysis values of the high-temperature non-target objects, divides them into a credible target level, an uncertain level, and an interference target level, and performs corresponding processing.

[0006] Preferably, after analyzing the extracted temperature-texture collaborative features, calculate the temperature-texture collaborative anomaly index, including: extracting the high-temperature region from the infrared image and setting the threshold as to obtain the preliminary target mask: ; extracting fine-grained edge / texture features from the infrared image : PST is an edge detection method implemented through Fourier domain convolution and non-linear phase operation. represents the scale control parameter in PST, represents the phase non-linearity control parameter in PST, represents the pixel value of the image at position (x, y) after PST processing, and construct the temperature-texture collaborative feature map , which represents the coupling relationship between the heat intensity and the texture response, and is defined as: ; represents the pixel value of the infrared image after normalization; represents the texture response image after PST transformation; in the high-temperature mask region , calculate the statistical mean and the standard deviation of , and define the temperature-texture collaborative anomaly index as the degree of dispersion and intensity coupling of the texture-temperature collaborative features. The expression is: ; in the formula, represents the temperature-texture collaborative anomaly index.

[0007] Preferably, after analyzing the extracted spatial dynamic anomaly features, generate the spatial dynamic anomaly index, specifically including: extracting the high-temperature region from the infrared image to obtain the preliminary target mask , a total of T frames; represents the thermal mask, which is a binary image extracted according to the pixel temperature in the infrared image exceeding the set threshold ; select a spatial position as the candidate region; For the pixels in the region R, calculate the structural similarity value between each pair of consecutive frames (It, It+1), and generate a set of temporal structural similarity values, and calculate the change amount of the adjacent frame SSIM. The expression is: ; in the formula, represents the structural similarity value between frame t and t+1 in the region R, represents the structural similarity value between frame t+1 and t+2 in the region R; T represents the number of frames in the infrared image sequence. Define the spatial dynamic anomaly index using a combined index of the volatility of SSIM and the inverse of the mean. The expression is: ; wherein, represents the average value of all SSIM values, and SDI represents the spatial dynamic anomaly index.

[0008] Preferably, the interference analysis model construction module specifically includes: Convert the temperature texture collaborative anomaly index and the spatial dynamic anomaly index into a comprehensive feature vector, and use the comprehensive feature vector as the input of the machine learning model; The machine learning model takes predicting the high-temperature non-target object analysis value label corresponding to each high-temperature region for each group of comprehensive feature vectors as the prediction target, and takes minimizing the sum of the prediction errors of the high-temperature non-target object analysis value labels corresponding to all high-temperature regions as the training target, and trains the machine learning model until the sum of the prediction errors reaches convergence and then stops the model training; Determine the high-temperature non-target object analysis value corresponding to each high-temperature region according to the model output result, wherein the machine learning model is a polynomial regression model.

[0009] Preferably, the grading module specifically includes: Compare the obtained high-temperature non-target object analysis value with the gradient threshold, the gradient threshold includes a first threshold and a second threshold, and the first threshold is less than the second threshold, and compare the high-temperature non-target object analysis value with the first threshold and the second threshold respectively; If the high-temperature non-target object analysis value is greater than the second threshold, it is determined as the interference target level; perform a down-weighting process on the region in the fused image; If the high-temperature non-target object analysis value is greater than or equal to the first threshold and less than or equal to the second threshold, it is determined as the uncertain level; mark the region as a region to be further evaluated; If the high-temperature non-target object analysis value is less than the first threshold, it is determined as the credible target level, assign a high fusion weight to the region, and directly participate in the subsequent Tibetan antelope behavior analysis and processing.

[0010] The present invention also provides an optimized transmission method for Tibetan antelope behavior analysis in multi-modal image fusion, including: Collect multi-modal image data including infrared images and visible light images, and preprocess the multi-modal image data; Extract the temperature texture collaborative feature and the spatial dynamic anomaly feature from the infrared image; Train a machine learning model based on the temperature texture collaborative feature and the spatial dynamic anomaly feature, and output the high-temperature non-target object analysis value corresponding to each high-temperature region; Perform hierarchical marking on the high-temperature regions in the infrared image according to the high-temperature non-target object analysis value, divide them into the credible target level, the uncertain level and the interference target level, and perform corresponding processing; For the high-temperature region in the uncertainty level, during the image fusion process, combine the response intensity and weight distribution of the fusion algorithm to analyze the interference degree on the expression of target features in the fused image, and output the fusion interference risk value; Combine the fusion interference risk value to perform optimization transmission and intelligent scheduling processing on the high-temperature region of the uncertainty level.

[0011] Preferably, for the high-temperature region in the uncertainty level, during the image fusion process, extract the intermediate response map information: the key feature map of the pre-fusion image ; the target region feature map of the post-fusion image ; extract the region from the fusion algorithm of the fusion weight value ; calculate the change degree of the key feature maps before and after fusion on the region , and the expression is: ; in the formula, ; represents the number of region pixels; the comprehensive distortion degree and the region fusion weight , define the fusion interference risk value , and the expression is: . .

[0012] Preferably, combine the fusion interference risk value to perform optimization transmission and intelligent scheduling processing on the image data, specifically including: Compare the fusion interference risk value of each high-temperature region of the uncertainty level with the preset scheduling threshold; If the fusion interference risk value is lower than the first scheduling threshold, mark the corresponding image region data as low priority and delay or compress the transmission; If the fusion interference risk value is between the first scheduling threshold and the second scheduling threshold, mark the region as medium priority and enter the asynchronous fusion cache queue; If the fusion interference risk value is higher than the second scheduling threshold, mark the region as high priority, immediately transmit it to the central analysis end, and enable the enhanced fusion mode for key processing.

[0013] In the above technical solution, the technical effects and advantages provided by the present invention: 1. The multi-modal image fusion Tibetan antelope behavior analysis system and optimization transmission method provided by the present invention innovatively propose a dual-feature modeling method of temperature texture collaborative anomaly index and spatial dynamic anomaly index for the problem that non-target high-temperature interference objects in infrared images are easily misjudged. Combine the polynomial regression model to realize the quantitative analysis of the interference risk of high-temperature regions, and accurately divide the credible, interference and uncertain regions through the hierarchical marking mechanism, effectively improving the fidelity of target animal features in the fused image and the accuracy of behavior recognition.

[0014] 2. By integrating the interference risk value to guide the optimized transmission and intelligent scheduling of the uncertain area, the present invention dynamically allocates data priorities under the conditions of limited bandwidth or edge computing power, ensuring that key areas are preferentially fused and enhanced, and improving the resource scheduling efficiency and environmental adaptability of the system. The overall solution not only improves the robustness of image fusion and target recognition, but also enhances the practicality and intelligence level of the system in complex wild ecological monitoring scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments recorded in the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings.

[0016] Figure 1 It is the system module mind map of the present invention.

[0017] Figure 2 It is the method mind map of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0018] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0019] Example 1. Refer to Figure 1 As shown, the multi-modal image fusion Tibetan antelope behavior analysis system in this embodiment includes an image acquisition module, an image preprocessing module, a feature extraction module, an interference analysis model construction module, and a grading module: The image acquisition module synchronously acquires multi-modal image data including infrared images and visible light images; The image preprocessing module preprocesses the multi-modal image data; The feature extraction module extracts temperature texture collaborative features and spatial dynamic anomaly features from the infrared images; The interference analysis model construction module trains a machine learning model based on the temperature texture collaborative features and spatial dynamic anomaly features, and outputs the analysis value of high-temperature non-target objects corresponding to each high-temperature area; The grading module classifies and marks the high-temperature regions in the infrared image according to the high-temperature non-target object analysis value, divides them into a credible target level, an uncertain level, and an interfering target level, and performs corresponding processing.

[0020] The image acquisition module is used to synchronously acquire multi-channel image information of the target area in the natural field environment, specifically including: Multi-modal sensor deployment: The module includes an infrared thermal imaging sensor and a visible light image sensor installed at a fixed observation point or an unmanned monitoring device, and may further include a depth camera, a night vision enhancer, or a radar sensor to enhance the imaging ability in low light and complex terrain; Time synchronization control mechanism: The acquisition of the infrared image and the visible light image is controlled by a unified timestamp, and a high-precision synchronous signal triggering mechanism is adopted to ensure that different modal images are acquired in the same time window, thereby reducing the impact of timing deviation on the subsequent registration and fusion accuracy; Spatial perspective consistency design: Various sensors are installed at fixed relative positions and in a unified orientation. The external parameter matrix between sensors is obtained through calibration to achieve spatial perspective consistency between images, providing a geometric basis for subsequent image registration; Environmental adaptability: The acquisition module is equipped with an automatic exposure control, an anti-reflection filter, and a temperature compensation device to cope with extreme natural conditions such as high-intensity light, low temperature, and high wind speed on the plateau, ensuring the stability of image quality; Data caching and uploading mechanism: The acquired multi-modal image data is temporarily stored through a local cache unit and uploaded to the backend processing system through a wireless communication module (such as LoRa, 4G / 5G, satellite link) under network permission conditions.

[0021] The image preprocessing module is used to perform structured processing on the multi-modal image data acquired by the image acquisition module to improve the effectiveness of subsequent fusion and recognition, specifically including: Align different modal images (such as infrared images and visible light images) in the spatial dimension so that corresponding pixel points have the same geographical or target meaning. The registration process includes: Geometric transformation calculation: Calculate the affine transformation matrix or the projection matrix based on the sensor calibration parameters or key points extracted from the scene (such as SIFT, SURF, ORB); Inter-modal feature matching: Adopt a multi-modal alignment algorithm, such as image registration based on mutual information, to adapt to image pairs with obvious information differences such as infrared and visible light; Deep learning-assisted registration: Optionally, use an end-to-end image registration model based on a neural network (such as RegNet or deep registration CNN) to improve the registration accuracy, especially suitable for images in areas without obvious texture.

[0022] Eliminate the noise signals generated during the acquisition process and enhance the image clarity and the expressiveness of edge features, including: Spatial filtering methods: such as median filtering, Gaussian filtering, bilateral filtering, etc., which are used to remove random noise; Frequency domain denoising methods: such as wavelet transform denoising, which is suitable for suppressing high-frequency noise while retaining the image structure information; Modal adaptive strategy: Aiming at the problem of low contrast between the high-temperature area and the background in infrared images, dynamically adjust the denoising intensity to retain the key thermal feature areas.

[0023] Unify the differences in the numerical range, dynamic range, and scale expression of multi-modal images for fusion processing. The normalization processing includes: Gray normalization: Normalize the pixel values of different modal images to a unified gray range (such as [0,1] or [0,255]); Thermal intensity normalization: Perform contrast stretching or histogram equalization on the heat values of infrared images to enhance the contrast of high-temperature areas; Spatial resolution unification: Make different modal images have a consistent resolution scale through image resampling methods (such as bilinear interpolation, nearest neighbor interpolation); Structural consistency alignment: Use edge-preserving normalization methods (such as the Retinex algorithm) to ensure that the images have good structure perception characteristics before fusion.

[0024] Through the above preprocessing steps, the image preprocessing module has achieved a high degree of consistency in the spatial position, feature expression, and data scale of different modal images, providing a standardized and comparable input basis for subsequent feature extraction, fusion, and recognition.

[0025] In the feature extraction module, the extraction of temperature-texture collaborative features (TTC) is to solve the problem that relying only on thermal intensity in traditional infrared images may cause target recognition errors, especially when there are non-animal high-temperature objects (such as sunlit rocks). By considering both the infrared intensity distribution in the high-temperature area and its texture edge structure in the image, TTC can effectively distinguish the real target areas with biological structure features (such as clear contours and closed edges) from the non-target heat source areas with irregular textures and scattered edges, thereby improving the recognition accuracy of high-temperature interference targets and providing a structural discrimination basis for subsequent fusion judgment.

[0026] After analyzing the extracted temperature-texture collaborative features, calculate the temperature-texture collaborative anomaly index, including: Extract the high-temperature area from the infrared image Set the threshold as to obtain the preliminary target mask: ; For the infrared image Extract fine-grained edge / texture features: ; PST is an edge detection method implemented through Fourier domain convolution and non-linear phase operations, which can enhance weak texture edges and is suitable for low-contrast infrared images; represents the scale control parameter in PST, which adjusts the width of the filter in the frequency domain and affects the accuracy of detail extraction; a smaller value is used to enhance fine edges; represents the phase non-linearity control parameter in PST, which controls the non-linearity degree of the enhancement effect and determines the amplitude and response mode of the edge response; represents the pixel value of the image at position (x, y) after PST processing, which reflects the texture / edge response intensity at that place.

[0027] Construct a temperature-texture collaborative feature map , which represents the coupling relationship between thermal intensity and texture response, and is defined as: ; represents the pixel value of the infrared image after normalization (such as linearly normalized to [0, 1]), which is used to unify the numerical scale of thermal intensity; represents the texture response image after PST transformation, and the pixel value at (x, y) after normalization processing; where the two images are first processed by standard normalization (such as 0-1 normalization) to ensure the same numerical scale.

[0028] In the high-temperature mask area On, for Calculate the statistical mean And standard deviation Define the temperature-texture collaborative anomaly index as the degree of discreteness and intensity coupling of the texture-temperature collaborative feature, and the expression is: ; In the formula, represents the temperature-texture collaborative anomaly index.

[0029] If the TTCAI is high, it indicates that the texture in the high-temperature area is inconsistent and the distribution is abnormal, which may be a non-animal target (such as a rock); if the TTCAI is low, it means that the boundary of the heat source area is clear and consistent with the thermal intensity, with typical animal target characteristics.

[0030] Extracting spatial dynamic anomaly features (SDM) aims to utilize the dynamic change information of image frames in the time series to model the spatial behavior patterns of high-temperature areas, and then identify whether they have the typical motion characteristics of biological targets. Since animals such as Tibetan antelopes usually show directional, continuous and deformable motion trajectories in the image sequence, while static heat sources (such as rocks or ground heat spots) remain relatively stable in shape and position, SDM can be used as a key discriminant feature in the dynamic dimension to assist in identifying the real target area with obvious motion consistency and suppressing the weight influence of static interference areas in the fusion process.

[0031] After analyzing the extracted spatial dynamic anomaly features, a spatial dynamic anomaly index is generated, which specifically includes: Extract the high-temperature region from the infrared image to obtain a preliminary target mask , a total of T frames; Denote the thermal mask, which is a binary image extracted according to the pixel temperature in the infrared image exceeding the set threshold and is used to mark the high-temperature region. Select a spatial position as the candidate region (such as a sliding window or a target region).

[0032] For the pixels in region R, calculate the structural similarity value between each pair of consecutive frames (It, It+1), generate a set of temporal structural similarity values, and calculate the change amount of adjacent frame SSIM , and the expression is: ; in the formula, represents the structural similarity value between frame t and t+1 in region R, represents the structural similarity value between frame t+1 and t+2 in region R; T represents the number of frames in the infrared image sequence. Define the spatial dynamic anomaly index using a combined index of SSIM volatility and mean inverse ratio, and the expression is: ; in the formula, represents the average value of all SSIM values, representing the overall structural retention degree, and SDI represents the spatial dynamic anomaly index.

[0033] SDI is high (e.g., > 0.5): The regional structure changes greatly in time and has strong dynamics → may be targets such as Tibetan antelopes; SDI is low (close to 0): The structure remains stable → may be a stationary heat source such as a rock.

[0034] The interference analysis model construction module trains a machine learning model based on the temperature texture collaborative features and spatial dynamic anomaly features, and outputs the analysis value of the high-temperature non-target object corresponding to each high-temperature region, which specifically includes: Convert the temperature texture collaborative anomaly index and the spatial dynamic anomaly index into a comprehensive feature vector, and use the comprehensive feature vector as the input of the machine learning model; The machine learning model takes predicting the label of the analysis value of the high-temperature non-target object corresponding to each high-temperature region with each group of comprehensive feature vectors as the prediction target, and takes minimizing the sum of the prediction errors of the analysis value labels of the high-temperature non-target objects corresponding to all each high-temperature region as the training target, trains the machine learning model until the sum of the prediction errors reaches convergence, and then stops the model training; Determine the high-temperature non-target object analysis value corresponding to each high-temperature region according to the model output result, where the machine learning model is a polynomial regression model.

[0035] In practical applications, use the trained polynomial regression model to predict the comprehensive feature vector of the high-temperature region in the newly input image, and output the high-temperature non-target object analysis value corresponding to each high-temperature region. This value is used for subsequent high-temperature region grading and interference risk assessment.

[0036] The grading module grades and marks the high-temperature regions in the infrared image according to the high-temperature non-target object analysis value, divides them into a credible target grade, an uncertain grade, and an interference target grade, and performs corresponding processing, specifically including: Compare the obtained high-temperature non-target object analysis value with the gradient thresholds. The gradient thresholds include a first threshold and a second threshold, and the first threshold is less than the second threshold. Compare the high-temperature non-target object analysis value with the first threshold and the second threshold respectively; If the high-temperature non-target object analysis value is greater than the second threshold, it is determined to be the interference target grade, indicating that this region is highly likely to be a non-animal heat source; perform a weight reduction process on this region in the fused image, or exclude it from participating in the fused feature extraction to avoid interfering with the fusion decision; If the high-temperature non-target object analysis value is greater than or equal to the first threshold and less than or equal to the second threshold, it is determined to be the uncertain grade, indicating that this region has a certain interference risk but the features are not completely significant; Mark this region as a region that needs further evaluation; If the high-temperature non-target object analysis value is less than the first threshold, it is determined to be the credible target grade, indicating that this region is highly likely to be a real biological target; assign a higher fusion weight to this region and directly participate in the subsequent Tibetan antelope behavior analysis and processing flow.

[0037] Embodiment 2, please refer to Figure 2 As shown, the optimized transmission method for Tibetan antelope behavior analysis based on multi-modal image fusion in this embodiment includes: Collect multi-modal image data including infrared images and visible light images, and preprocess the multi-modal image data; Extract the temperature texture collaborative feature and the spatial dynamic abnormal feature from the infrared image; Train a machine learning model based on the temperature texture collaborative feature and the spatial dynamic abnormal feature, and output the high-temperature non-target object analysis value corresponding to each high-temperature region; Grade and mark the high-temperature regions in the infrared image according to the high-temperature non-target object analysis value, divide them into a credible target grade, an uncertain grade, and an interference target grade, and perform corresponding processing; For the high-temperature region in the uncertainty level, during the image fusion process, combine the response intensity and weight distribution of the fusion algorithm to analyze the degree of interference on the expression of target features in the fused image, and output the fusion interference risk value; Combine the fusion interference risk value to perform optimized transmission and intelligent scheduling processing on the high-temperature region of the uncertainty level.

[0038] Among them, for the high-temperature region in the uncertainty level, during the image fusion process (such as in deep fusion network or attention mechanism fusion), extract the intermediate response map information: The key feature map of the pre-fusion image , such as the edge map and texture direction map; The target region feature map of the post-fusion image , especially the part that overlaps with the high-temperature region in terms of spatial position.

[0039] Extract the fusion weight value of the region from the fusion algorithm , which can be sourced from: attention weight map; channel response intensity of the fusion feature map; pixel weighted average / transformation weight value.

[0040] Calculate the degree of change (feature distortion) of the key feature map before and after fusion in the region , and the expression is: ; where, represents the number of region pixels; the difference can be measured by methods such as L1 norm, structural similarity decline, and gradient direction change. The comprehensive distortion degree and the region fusion weight , define the fusion interference risk value , and the expression is: .

[0041] Receive the fusion interference risk value (FIR) of each high-temperature region of the uncertainty level output; The system presets two-level scheduling thresholds, namely the first scheduling threshold TL and the second scheduling threshold TH, satisfying TL < TH; Compare the FIR of each high-temperature region with TL and TH, and divide it into the following three types of scheduling levels: Low priority (FIR < TL); Medium priority (TL ≤ FIR ≤ T_H); High priority (FIR > TH).

[0042] ​Low-priority processing: The data in the corresponding area is marked as a delayed transmission area, and compression encoding, low-frequency sampling can be performed, or it can be uploaded when network resources are idle; the system can use the edge caching mechanism to temporarily store such data to reduce resource occupancy.

[0043] Medium-priority processing: The regional image data enters the fusion asynchronous processing queue, and the system dynamically allocates the processing time according to the fusion model load in the main thread or batch scheduling; at the same time, the fusion parameters are retained for reuse when the fusion model is adjusted later.

[0044] High-priority processing: The regional data is immediately marked as a key image block and preferentially uploaded to the central server or high-performance node; the system triggers an enhanced fusion mode, which may include using a higher-precision image registration algorithm, a stronger feature extraction model, or an explicit fusion enhancement strategy (such as a weighting strategy, local reconstruction) to reduce the interference effect.

[0045] Based on the historical scheduling effect, the values of TL and TH are dynamically adjusted through reinforcement learning or a sliding window statistical model to adapt to the system response requirements under different climate, lighting, or target density conditions, and to achieve the adaptive optimization of the scheduling strategy.

[0046] By implementing risk-aware scheduling management for uncertain areas before image fusion, the optimization scheduling module achieves the goals of resource priority allocation, dynamic adjustment of fusion processing, and self-optimization of system performance, and significantly improves the continuous monitoring efficiency and recognition accuracy of target behaviors such as Tibetan antelopes in complex ecological environments.

[0047] The above formulas are all dimensionless and take their numerical calculations. The formulas are obtained by collecting a large amount of data for software simulation to get a formula closest to the real situation. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.

[0048] It should be understood that the term "and / or" in this article is merely a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Among them, A and B can be singular or plural. In addition, the character " / " in this article generally represents an "or" relationship between the front and back associated objects, but it may also represent an "and / or" relationship, which can be specifically understood by referring to the context.

[0049] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0050] As described above, the above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed in this application can easily think of changes or substitutions, and all should be covered within the protection scope of this application.

Claims

1. A multi-modal image fusion-based Tibetan antelope behavior analysis system, characterized in that: It includes an image acquisition module, an image preprocessing module, a feature extraction module, an interference analysis model construction module, and a grading module: The image acquisition module synchronously acquires multi-modal image data including infrared images and visible light images; The image preprocessing module preprocesses the multi-modal image data; The feature extraction module extracts temperature-texture collaborative features and spatial dynamic anomaly features from the infrared images; The interference analysis model construction module trains a machine learning model based on the temperature-texture collaborative features and spatial dynamic anomaly features, and outputs the analysis value of high-temperature non-target objects corresponding to each high-temperature region; The grading module classifies and marks the high-temperature regions in the infrared image according to the analysis value of high-temperature non-target objects, divides them into a credible target level, an uncertain level, and an interference target level, and performs corresponding processing.

2. The multi-modal image fusion Tibetan antelope behavior analysis system according to claim 1, characterized in that: After analyzing the extracted temperature-texture collaborative features, calculate the temperature-texture collaborative anomaly index, including: extracting the high-temperature region from the infrared image and setting the threshold as to obtain the preliminary target mask: ; extracting the fine-grained edge / texture features from the infrared image : PST is an edge detection method implemented through Fourier domain convolution and non-linear phase operation, represents the scale control parameter in PST, represents the phase non-linearity control parameter in PST, represents the pixel value of the image at position (x, y) after PST processing, and construct the temperature-texture collaborative feature map , which represents the coupling relationship between the heat intensity and the texture response, and is defined as: ; represents the pixel value of the infrared image after normalization; represents the texture response image after PST transformation; in the high-temperature mask region , for , calculate the statistical mean and the standard deviation , and define the temperature-texture collaborative anomaly index as the degree of dispersion and intensity coupling of the texture-temperature collaborative features, and the expression is: ; in the formula, represents the temperature-texture collaborative anomaly index.

3. The multi-modal image fusion Tibetan antelope behavior analysis system according to claim 2, wherein: After analyzing the extracted spatial dynamic anomaly features, a spatial dynamic anomaly index is generated, which specifically includes: extracting high-temperature regions from infrared images to obtain a preliminary target mask , a total of T frames; denotes the thermal mask, which is a binary image extracted based on the pixel temperature in the infrared image exceeding the set threshold ; select a spatial position as the candidate region; For the pixels in region R, calculate the structural similarity value between each pair of consecutive frames (It, It+1), and generate a set of temporal structural similarity values. Calculate the change amount of the SSIM between adjacent frames. , and the expression is: ; where represents the structural similarity value between frames t and t+1 in region R, represents the structural similarity value between frames t+1 and t+2 in region R; T represents the number of frames in the infrared image sequence. Define the spatial dynamic anomaly index using a combined index of the volatility of SSIM and the inverse of the mean, and the expression is: ; where represents the average value of all SSIM values, and SDI represents the spatial dynamic anomaly index.

4. The Tibetan antelope behavior analysis system for multimodal image fusion according to claim 3, wherein: The interference analysis model construction module specifically includes: Converting the temperature-texture collaborative anomaly index and the spatial dynamic anomaly index into a comprehensive feature vector, and using the comprehensive feature vector as the input of the machine learning model; Taking the prediction of the analysis value label of high-temperature non-target objects corresponding to each high-temperature region by each group of comprehensive feature vectors as the prediction target of the machine learning model, and taking the minimization of the sum of prediction errors of the analysis value labels of high-temperature non-target objects corresponding to all each high-temperature region as the training target, training the machine learning model until the sum of prediction errors reaches convergence and then stopping the model training; Determining the analysis value of high-temperature non-target objects corresponding to each high-temperature region according to the model output result, where the machine learning model is a polynomial regression model.

5. The multi-modal image fusion Tibetan antelope behavior analysis system according to claim 4, characterized in that: The grading module specifically includes: Comparing the obtained analysis value of high-temperature non-target objects with the gradient thresholds, where the gradient thresholds include a first threshold and a second threshold, and the first threshold is less than the second threshold, and comparing the analysis value of high-temperature non-target objects with the first threshold and the second threshold respectively; If the analysis value of high-temperature non-target objects is greater than the second threshold, it is determined as the interference target level; perform a down-weighting process on the region in the fused image; If the analysis value of high-temperature non-target objects is greater than or equal to the first threshold and less than or equal to the second threshold, it is determined as the uncertain level; mark the region as a region to be further evaluated; If the analysis value of high-temperature non-target objects is less than the first threshold, it is determined as the credible target level, assign a high fusion weight to the region, and directly participate in the subsequent analysis and processing of the behavior of Tibetan antelopes.

6. An optimized transmission method for analyzing the behavior of Tibetan antelopes through multimodal image fusion, which is used to implement the multimodal image fusion-based Tibetan antelope behavior analysis system described in any one of claims 1-5, characterized in that: It includes: Collecting multi-modal image data including infrared images and visible light images, and preprocessing the multi-modal image data; Extracting temperature-texture collaborative features and spatial dynamic anomaly features from the infrared images; Training a machine learning model based on the temperature-texture collaborative features and spatial dynamic anomaly features, and outputting the analysis value of high-temperature non-target objects corresponding to each high-temperature region; Classifying and marking the high-temperature regions in the infrared image according to the analysis value of high-temperature non-target objects, dividing them into a credible target level, an uncertain level, and an interference target level, and performing corresponding processing; For the high-temperature area in the uncertainty level, during the image fusion process, combine the response intensity and weight distribution of the fusion algorithm, analyze the degree of interference on the expression of target features in the fused image, and output the fusion interference risk value; Combine the fusion interference risk value to perform optimized transmission and intelligent scheduling processing on the high-temperature area of the uncertainty level.

7. The optimized transmission method for analyzing the behavior of Tibetan antelopes by multimodal image fusion according to claim 6, characterized in that: Among them, For high temperature areas in the uncertainty level, during the image fusion process, the intermediate response map information is extracted: the key feature map of the image before fusion Feature map of the target area of the fused image ; Extract regions from fusion algorithm The fusion weight value ; In the area Calculate the change degree of the key feature map before and after fusion , the expression is: ; In the formula, Indicates the number of pixels in the area; the overall degree of distortion and regional fusion weights , define the fusion interference risk value , the expression is: .

8. The optimized transmission method for analyzing the behavior of Tibetan antelopes by multimodal image fusion according to claim 7, characterized in that: Combining the fusion interference risk value to perform optimized transmission and intelligent scheduling processing on the image data specifically includes: Compare the fusion interference risk value of the high-temperature area of each uncertainty level with the preset scheduling threshold; If the fusion interference risk value is lower than the first scheduling threshold, mark the data of the corresponding image area as low priority and delay or compress the transmission; If the fusion interference risk value is between the first scheduling threshold and the second scheduling threshold, mark the area as medium priority and enter the asynchronous fusion cache queue; If the fusion interference risk value is higher than the second scheduling threshold, mark the area as high priority, immediately transmit it to the central analysis end, and enable the enhanced fusion mode for key processing.

Citation Information

Patent Citations

  • Bird monitoring system with combination of dual-light camera carried by unmanned aerial vehicle and deep learning

    CN118196660A

  • Lottery store violation abnormity identification method and system

    CN119229260A

  • Target detection method and system based on multi-modal sensor information fusion

    CN119273964A

  • Hunting camera imaging quality optimization method and system based on multimode data fusion

    CN119444598A

  • Object-level infrared-and-visible-light image fusion method based on fully convolutional neural network

    WO2024174488A1

Cited By

  • Transparent package foreign matter detection method and system based on multispectral fusion and adaptive threshold, and medium

    CN120997600A