A method for processing partition data of an infrared image in a specific area
Through multimodal data fusion and objective function optimization processing methods, the problems of insufficient partition accuracy and poor reliability of traditional infrared image processing in complex backgrounds are solved, and high-precision and robust partitioning results are achieved, which are suitable for a variety of scenarios.
Patent Information
- Application Number
- CN202510214247.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-02-26
AI Technical Summary
Traditional infrared image processing methods have problems such as insufficient partitioning accuracy in target areas, insufficient fusion of multimodal data, and poor partitioning results reliability and multi-scene applicability under complex backgrounds.
Using a technical solution of multimodal data fusion, the objective function is constructed for partitioning processing by combining infrared images, visible light images and lidar data, including data consistency terms, boundary smoothing terms and modal interaction verification terms, and optimized by Lagrangian multiplication method and gradient descent method, and dynamically adjusting the modal weights.
It significantly improves the accuracy and robustness of partition results, solves the problem of single data source and limited processing accuracy, realizes precise partitioning in complex contexts, and improves the reliability and multi-scenario applicability of partition results.
Smart Images

Figure CN119722713B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of infrared image processing, and specifically provides a method for processing partition data of infrared images in a specific area. Background Art
[0002] With the rapid development of information technology, infrared imaging technology is widely used in multiple fields such as monitoring, security, and military reconnaissance because it is not restricted by light conditions and has good recognition performance at night or in complex environments. However, traditional infrared image processing methods often lack targeted data processing strategies. Especially when facing complex targets in a specific area, how to efficiently and accurately extract useful information has become a challenge.
[0003] Currently, common infrared image processing technologies mainly include global threshold segmentation method, edge detection technology, and deep learning-based object recognition method. The global threshold segmentation method distinguishes the target from the background by setting a unified standard, which is easy to operate but has poor accuracy in the case of complex and changeable backgrounds; the edge detection technology relies on the image gradient change to locate the object boundary and is difficult to effectively capture targets in low-contrast environments; although the deep learning-based method has a significant improvement in object recognition accuracy, it requires a large number of labeled samples for training and consumes a large amount of computing resources. These methods have their own advantages, but there are still obvious limitations in the fine data processing for specific areas. Summary of the Invention
[0004] Aiming at the deficiencies of the prior art, the present invention provides a method for processing partition data of infrared images in a specific area, which solves the problems of insufficient partition accuracy of the target area in complex backgrounds, insufficient multi-modal data fusion, and poor reliability and multi-scene applicability of the partition results in traditional infrared image processing.
[0005] To achieve the above objectives, the present invention is realized through the following technical solutions: A method for processing partition data of infrared images in a specific area, including the following steps:
[0006] Obtain data from multi-modal sensors;
[0007] Calibrate and preprocess the obtained multi-modal data;
[0008] Construct an objective function for partition processing, and the objective function is used to measure the characteristics of the partition area;
[0009] Optimize the objective function to determine the partition boundary;
[0010] Output the target area partition based on the optimization result.
[0011] Preferably, the multi-modal sensors include:
[0012] An infrared image sensor for collecting thermal radiation information;
[0013] A visible light image sensor for collecting texture and color information;
[0014] A lidar sensor for collecting depth and spatial information.
[0015] Preferably, the calibration of the multi-modal data includes:
[0016] Spatial registration and time synchronization of the three-modal data:
[0017] Spatial registration realizes the registration of the infrared image and the visible light image through a feature matching algorithm:
[0018] The lidar data is mapped to the image plane coordinate system through a point cloud projection matrix.
[0019] Preferably, the objective function includes a data consistency term, a boundary smoothing term, and a modal interaction verification term, and its expression is:
[0020] ;
[0021] Wherein, represents the total energy of the objective function; represents the partition region; represents the partition boundary; represents the data consistency term, which is used to measure the characteristic consistency of different modal data within the target region; represents the boundary smoothing term, which is used to constrain the smoothness of the partition boundary; represents the modal interaction verification term, which is used to measure the data consistency between modalities; represents the weight of the smoothing term; represents the weight of the modal interaction verification term.
[0022] Preferably, the optimization process of the objective function includes the following steps:
[0023] Based on the Lagrange multiplier method, the objective function is modeled, and a partition consistency constraint condition is introduced. The expression of the objective function is:
[0024] ;
[0025] Wherein, represents the Lagrange objective function; represents the consistency constraint condition, which is used to constrain the boundary consistency of the data between modalities;
[0026] The gradient descent method is used to iteratively solve the partition boundary, and the update expression of the partition boundary is:
[0027] ;
[0028] wherein, represents the partition boundary at the -th iteration; represents the learning rate; represents the gradient of the objective function;
[0029] Dynamically adjust the modal weights in the objective function. The modal weights are determined based on the noise level of the modal data. The calculation formula for the modal weights is:
[0030] ;
[0031] wherein, represents the weight of the -th modality; represents the noise level of the -th modality; represents the index variable; represents the -th standard deviation of the modality;
[0032] Continuously optimize the objective function until the energy converges to the optimal value, and generate the final partition boundary.
[0033] Preferably, the partition processing includes the following steps:
[0034] Determine the partition boundary of the target area and the set of pixels within the partition;
[0035] Generate corresponding characteristic data based on the partition result of the target area;
[0036] Analyze and process the characteristic data and then output it. The partition result is used for monitoring, security, or fire monitoring of the target area.
[0037] Preferably, the data consistency term is calculated according to the dynamic modal weights. The expression of the data consistency term is:
[0038] ;
[0039] wherein, represents the data consistency term; represents the data of the -th modality; represents the characteristic estimated value of the target area within the partition; represents the weight of the -th modality.
[0040] Preferably, the modal interaction verification term verifies the consistency of the segmentation region by analyzing the boundary probability distribution between modalities, and evaluates the modal boundary consistency based on the Bayesian inference model.
[0041] Preferably, the target area partitioning includes the following steps:
[0042] Construct an initial partitioning boundary for the target area;
[0043] Based on the characteristics of multi-modal data, perform data consistency analysis on the initial partitioning boundary;
[0044] Use an optimization algorithm to dynamically adjust the partitioning boundary so that the boundary conforms to the characteristic distribution of the target area;
[0045] Combine the modal interaction verification results to correct the partitioning boundary of the target area and generate the final partitioning result.
[0046] The present invention also provides a partitioning data processing system for infrared images of a specific area, including:
[0047] A multi-modal sensor module for collecting infrared images, visible light images, and lidar point cloud data;
[0048] A data calibration module for performing spatial registration and time synchronization on the collected multi-modal data:
[0049] A partitioning modeling module for constructing a partitioning processing objective function for the target area:
[0050] An optimization solving module for modeling the objective function based on the Lagrange multiplier method, and iteratively optimizing the partitioning boundary through the gradient descent method, dynamically adjusting the modal weights, and optimizing the target area partitioning;
[0051] A partitioning verification module for evaluating the reliability of the partitioning result through modal interaction verification:
[0052] A result output module for outputting the partitioning boundary of the target area and the pixel set within the partition, and the partitioning result can be used in monitoring, security, or fire monitoring scenarios of the target area.
[0053] The present invention provides a partitioning data processing method for infrared images of a specific area. It has the following beneficial effects:
[0054] 1. By adopting the technical solution of multi-modal data fusion, the present invention combines infrared images, visible light images, and lidar data, achieving the technical effect of significantly improving the accuracy and robustness of the partitioning result. Compared with the problem of insufficient characteristic information of the target area caused by single-modal data processing in the prior art, it solves the deficiencies of single data source and limited processing accuracy.
[0055] 2. The present invention adopts the technical solution of objective function modeling, and quantitatively evaluates the partitioned regions through data consistency terms, boundary smoothing terms, and modal interaction verification terms, achieving the technical effect of accurate partitioning in complex backgrounds. Compared with the prior art solutions of global threshold segmentation and edge detection with poor adaptability to low-contrast environments, the problem of insufficient partitioning accuracy in complex background situations is solved.
[0056] 3. The present invention adopts the technical solution of dynamic modal weight adjustment and optimization solution, and iteratively optimizes the partition boundary through Lagrangian modeling and gradient descent method, achieving the technical effect of dynamically adjusting the partition result according to the quality of modal data. Compared with the prior art solutions that do not fully consider the noise level of multi-modal data, the deficiency of inaccurate optimization of the partition boundary is solved.
[0057] 4. The present invention adopts the technical solution of modal interaction verification and Bayesian inference model to evaluate the reliability and consistency of the partition result, achieving the technical effect of effectively excluding abnormal partitioned regions. Compared with the prior art solutions where the reliability of the partition result is difficult to verify, the deficiencies of discontinuous partition boundaries and modal conflicts are solved.
[0058] 5. The present invention adopts the technical solution of multi-format partition result output, and generates partition boundary and regional characteristic data and supports multiple application scenarios, achieving the technical effect of improving the applicability and scalability of the partition result. Compared with the prior art solutions where the partition result is limited to single-scenario applications, the problem of insufficient multi-scenario applicability is solved. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] Figure 1 is the flowchart of the method of the present invention;
[0060] Figure 2 is the system architecture diagram of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0061] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0062] Please refer to Figure 1 , the embodiments of the present invention provide a method for processing partition data of an infrared image of a specific region, including the following steps:
[0063] S1. Obtain data from a multi-modal sensor;
[0064] S2. Calibrate and preprocess the acquired multimodal data;
[0065] S3. Construct an objective function for partition processing, where the objective function is used to measure the characteristics of the partition region;
[0066] S4. Optimize the objective function to determine the partition boundary;
[0067] S5. Output the target area partition based on the optimization result.
[0068] For S1, it is to obtain the data of the multimodal sensor. This step is the basis of the entire partition data processing process and directly affects the effects of subsequent calibration, partition modeling, and optimization processing. By reasonably selecting and configuring the multimodal sensor, sufficient raw data support can be provided for the accurate partition of the target area, and the integrity of the spatial, temporal, and characteristic information of the data can be ensured.
[0069] Generally, the data acquisition of the multimodal sensor includes multiple data sources, and each modality of data has different characteristics and application scenarios. For example, infrared image data mainly reflects the thermal radiation characteristics of the target area, visible light image data provides texture and color information, and lidar data captures the depth and spatial geometric characteristics of the target area. The comprehensive acquisition and use of multimodal data can significantly improve the accuracy and robustness of partition processing.
[0070] Specifically, in this embodiment, step S1 includes the following contents:
[0071] First, configure the multimodal sensor system. As an option, the following multimodal sensors are preferably used in the present invention:
[0072] Infrared image sensor, used to capture the thermal radiation characteristics of the target area, and its output data contains the infrared intensity information of the target area, which is used to detect the thermal signal distribution of the target object;
[0073] Visible light image sensor, used to capture the color and texture characteristics of the target area, and its output data contains the optical characteristics of the target area, which is used to refine the external boundary of the target object;
[0074] Lidar sensor, used to obtain the depth information of the target area, and its output data includes point cloud data, and the point cloud data can reflect the three-dimensional spatial geometric characteristics of the target area.
[0075] In one implementation, the multimodal sensor system realizes synchronous acquisition through a hardware trigger signal. Generally, the frame rate of sensor synchronous acquisition should meet the requirements of scene dynamics to ensure the data integrity and consistency of the target area during the sampling process.
[0076] In some embodiments, the acquisition range of the multimodal sensor should cover the entire target area. Specifically, the field of view angles of the infrared image sensor and the visible light image sensor need to be matched to avoid spatial misalignment problems between different modality data; the detection range of the lidar needs to cover all significant features of the target area, including the boundaries and internal regions of the target object.
[0077] As an option, the multimodal data output by the sensor can be stored in multiple formats. Infrared images and visible light images are usually stored in the form of two-dimensional matrices, and the pixel values in the matrices represent the thermal radiation intensity and optical characteristics of the target area in each modality respectively; the data output by the lidar is usually stored in the form of three-dimensional point clouds, and each point contains three-dimensional spatial coordinates and its reflected intensity value.
[0078] In one implementation, the acquisition of multimodal data can be carried out through the following process:
[0079] Collect the thermal signal distribution of the target area through the infrared image sensor and output it in the form of a grayscale image;
[0080] Collect the optical characteristics of the target area through the visible light image sensor and output it in the form of a color image or a grayscale image;
[0081] Scan the target area through the lidar sensor and output point cloud data, where each point contains three-dimensional spatial coordinates and its corresponding reflected intensity value.
[0082] Generally, there may be data noise problems during the acquisition process. For example, the infrared image may be interfered by the ambient temperature, the visible light image may be affected by the lighting conditions, and the lidar data may have problems such as sparse point clouds or local missing points. For this reason, as an implementation, a preprocessing mechanism can be introduced in the data acquisition stage, including operations such as noise filtering and dynamic range adjustment, to improve the quality of multimodal data.
[0083] In this embodiment, the acquired multimodal data can be represented by the following expressions:
[0084] Infrared image data: , where represents the thermal radiation intensity of the infrared image at the pixel point ;
[0085] Visible light image data: , where represents the optical characteristics of the visible light image at the pixel point ;
[0086] Lidar point cloud data: , where, represents the three-dimensional coordinates of the point cloud, represents the reflection intensity of the point cloud.
[0087] As an option, during the sensor configuration process, the acquisition parameters can be adjusted according to the requirements of the application scenario. For example, in a night surveillance scenario, the acquisition sensitivity of the infrared image sensor can be preferentially increased; in a complex lighting environment, high dynamic range imaging technology can be introduced to enhance the capture of details in visible light images; in a three-dimensional target segmentation scenario, the point cloud density of the lidar can be adjusted to ensure accurate scanning of the target area.
[0088] The data collected by the above sensors will be directly used in the subsequent calibration and preprocessing steps. Through high-quality multi-modal data input, reliable basic data support can be provided for the construction and partition optimization of the objective function. Due to the complementary characteristics of multi-modal data, infrared image data can provide the thermal characteristics of the target area, visible light image data can provide texture information, and lidar point cloud data can provide depth information. The combination of multi-modal data will effectively solve the deficiencies of single-modal data processing and improve the accuracy and robustness of the partition results of the target area.
[0089] In summary, through the implementation of step S1, comprehensive acquisition of multi-modal sensor data can be achieved, providing sufficient raw data support for the precise partitioning of the target area.
[0090] For S2, it involves calibrating and preprocessing the acquired multi-modal data, which is an essential key link after data acquisition. Through the calibration and preprocessing steps, the problems of differences in time, space, and characteristics of multi-modal data can be solved, laying a consistent foundation for the construction and optimization of the subsequent partition objective function. Generally, multi-modal data calibration includes spatial registration and time synchronization, and preprocessing includes operations such as noise filtering, data format standardization, and dynamic range adjustment. Through this step, the fusion accuracy between multi-modal data can be effectively improved, and the partition error caused by data differences can be reduced.
[0091] Specifically, in this embodiment, step S2 includes the following content:
[0092] In terms of spatial calibration, it is first necessary to achieve spatial alignment among infrared images, visible light images, and lidar point cloud data. Generally, there may be differences in the field of view angles between infrared images and visible light images, and the spatial coordinate systems of lidar point cloud data and image data are also inconsistent. As a possible implementation method, the spatial transformation matrix between infrared images and visible light images can be obtained through the calibration plate calibration method, thereby completing the registration of the two. Specifically, the calibration plate calibration method realizes the spatial alignment of the two modalities of data by detecting the feature points on the calibration plate and calculating the spatial transformation relationship between the feature points. In another implementation method, the corresponding feature points in the infrared image and the visible light image can be automatically detected through the ORB (Oriented FAST and Rotated BRIEF) feature point matching algorithm, and the registration of the two can be completed based on the homography matrix.
[0093] In the spatial mapping of point cloud data, the projection matrix of the lidar can be used to map the three-dimensional point cloud data to the image plane coordinate system. Generally, the projection matrix can be obtained through the geometric calibration of the lidar and the image sensor, and its mapping relationship can be expressed as:
[0094] ;
[0095] where, represents the image plane coordinates, represents the three-dimensional coordinates of the point cloud, represents the projection matrix.
[0096] Generally, in order to ensure data consistency, interpolation processing needs to be performed on the registered infrared images, visible light images, and point cloud data. For example, for sparse point cloud data, the bilinear interpolation method or the nearest neighbor interpolation method can be used to interpolate the point cloud data to the same resolution and grid structure as the image data for subsequent fusion processing.
[0097] In terms of time synchronization, to ensure the temporal consistency of multi-modal data, data synchronization can be achieved through hardware trigger signals or timestamp alignment methods. As an option, the hardware trigger signal controls the sampling times of each sensor through a unified trigger device, thereby ensuring the temporal consistency of infrared images, visible light images, and lidar data. In another implementation method, the timestamp-based alignment method records the acquisition timestamps of each modality of data and performs interpolation or cropping processing on the data according to the time difference, thereby achieving time synchronization.
[0098] In terms of preprocessing, it is first necessary to filter the noise in the multimodal data. Infrared image data is generally interfered by thermal noise, visible light image generally introduces random noise due to lighting conditions, and lidar point cloud data generally has spurious points or noise points. Specifically, Gaussian filters can be used to smooth infrared images to reduce the impact of thermal noise; median filters can be used to remove random noise from visible light images; statistical filtering methods or radius filtering methods can be used to remove outliers and spurious points from lidar point cloud data.
[0099] As an option, it is also possible to adjust the dynamic range of the multimodal data to enhance the contrast and feature saliency between different modal data. The adjustment of the dynamic range of infrared images can be achieved through histogram equalization. Specifically, the histogram equalization method expands the dynamic range of the image to the entire pixel value range by redistributing the probability distribution of pixel values, thereby enhancing the contrast of the image. Gamma correction can be used to adjust the brightness distribution of visible light images to adapt to the lighting conditions of the target area. The reflection intensity of lidar data can be standardized through linear normalization to enhance the expression ability of point cloud features.
[0100] In terms of data format standardization, all multimodal data needs to be converted into a unified format for subsequent processing. For example, infrared images and visible light images can be standardized into two-dimensional matrices with the same resolution and pixel depth; lidar data can be standardized into a three-dimensional point cloud structure with a fixed point cloud density.
[0101] In some embodiments, the calibration and preprocessing effects can be further optimized by combining the characteristics of multimodal data. For example, for infrared images with low resolution, the high-resolution characteristics of visible light images can be used for super-resolution reconstruction to improve the resolution of infrared images. In another possible implementation, the three-dimensional geometric characteristics of lidar point cloud data can be used to correct the boundaries of infrared images and visible light images, thereby improving the accuracy of multimodal data fusion.
[0102] Through the calibration and preprocessing in this step, the acquired multimodal data can be converted into well-consistent standardized data, providing reliable data support for the construction of the objective function and partition optimization. At the same time, the noise filtering and dynamic range adjustment of multimodal data can significantly improve the data quality, thereby enhancing the accuracy and robustness of target area partitioning. In summary, this step plays a crucial role in connecting the preceding and the following in the entire partition data processing method and is the basis for ensuring the smooth implementation of subsequent steps.
[0103] For S3, it is the core link to implement partition data processing. By constructing the objective function, the optimization objective of partition processing can be clarified, and a quantitative evaluation criterion can be provided to guide the subsequent optimization of the partition boundary. Generally, when designing the objective function, the consistency of the characteristics of multimodal data, the smoothness of the boundary, and the cross-modal verification need to be comprehensively considered to ensure the accuracy and reliability of the partition results of the target area.
[0104] Specifically, in this embodiment, step S3 includes the following content:
[0105] The construction of the objective function includes three parts: the data consistency term, the boundary smoothness term, and the cross-modal verification term, which are specifically described as follows:
[0106] Generally, the data consistency term is used to measure the consistency between the multimodal data in the partition area and the estimated values of the characteristics of the target area, so as to ensure that the partition area can accurately reflect the characteristics of the target area. The mathematical expression of the data consistency term is:
[0107] ;
[0108] Where, represents the data consistency term; represents the observation value of modality , which are respectively infrared image data, visible light image data, and lidar data in the present invention; represents the weight of modality , which is dynamically determined by the noise level of the modality data, and the calculation formula is:
[0109] ;
[0110] Where, represents the noise level of modality ; represents the weight of modality ; represents the index variable, which is used to traverse the data uncertainty of all modalities; represents the standard deviation of the th modality (or data channel), which measures the dispersion or uncertainty of the data of this modality.
[0111] As an option, the estimated value of the characteristics of the target area can be calculated by the weighted average value in the partition area, and the specific calculation formula is:
[0112] ;
[0113] The above calculation can effectively utilize the complementarity of multimodal data and reduce the interference of high-noise modalities on the partitioning results.
[0114] In the construction of the boundary smoothing term, the smoothness constraint of the partitioning boundary is mainly considered to avoid discontinuous or irregular boundary shapes in the partitioning results. The mathematical expression of the boundary smoothing term is:
[0115] ;
[0116] where, is denoted as the boundary smoothing term, is denoted as the gradient change of the partitioning boundary, which is used to describe the smoothness of the partitioning boundary.
[0117] In one implementation, the boundary smoothness of the partitioning results can be enhanced by increasing the weight parameter of the boundary smoothing term, so as to meet the partitioning requirements of the target region in complex scenarios. The modality interaction verification term is used to measure the data consistency between modalities. Through the comparative analysis of multimodal data, the reliability of the target region partitioning is verified. The mathematical expression of the modality interaction verification term is:
[0118] ;
[0119] where, is denoted as the modality interaction verification term; is denoted as the observation values of modality and modality ; is denoted as the interaction weight between modality and modality .
[0120] As an option, the interaction weight can be set according to the complementary characteristics between modalities to improve the effect of modality interaction verification. For example, infrared image data and visible light image data have strong complementarity in thermal radiation and physical characteristics, so a relatively high value can be assigned to their interaction weight.
[0121] In this embodiment, the total energy expression of the objective function is:
[0122] ;
[0123] where, is denoted as the total energy of the objective function; is denoted as the partitioning region; is denoted as the partitioning boundary; is denoted as the weight of the boundary smoothing term; is denoted as the weight of the modality interaction verification term.
[0124] Generally, the smaller the total energy value of the objective function, the more the partition region conforms to the characteristic requirements of multimodal data, the smoother the partition boundary, and the higher the modal interaction consistency. In one implementation, by adjusting and values, the importance of boundary smoothness and modal interaction verification can be controlled to meet the requirements of different scenarios.
[0125] This step discusses the construction of the objective function, clarifying the evaluation criteria and optimization objectives for partition optimization. In the entire partition data processing flow, the construction of the objective function directly affects the accuracy and reliability of the partition result, and also provides a mathematical basis for subsequent optimization processing. By reasonably designing the data consistency term, boundary smoothness term, and modal interaction verification term in the objective function, the characteristics of multimodal data can be fully utilized to significantly improve the accuracy and robustness of the target region partition. In summary, this step is a key link in implementing the specific region infrared image partition data processing method.
[0126] For S4, to optimize the objective function to determine the optimal partition boundary of the target region. The optimization process is the core step of partition data processing, and its main objective is to obtain the optimal partition result that conforms to the characteristics of multimodal data and partition requirements by minimizing the energy of the objective function. Generally, the optimization process needs to comprehensively consider multiple constraint conditions of the objective function, including data consistency, smoothness of the partition boundary, and modal interaction verification results, so as to ensure the accuracy and robustness of the partition.
[0127] Specifically, in this embodiment, step S4 includes the following content:
[0128] Generally, the optimization of the objective function needs to achieve a balance among data consistency, boundary smoothness, and modal interaction verification. To introduce the partition consistency constraint, in this embodiment, Lagrangian modeling is performed on the objective function. Specifically, the Lagrangian form of the objective function is expressed as:
[0129] ;
[0130] where represents the Lagrangian objective function; represents the original objective function, including the data consistency term, boundary smoothness term, and modal interaction verification term; represents the partition consistency constraint condition, which is used to ensure the consistency between multimodal data; represents the Lagrangian multiplier, which is used to adjust the trade-off between the objective function and the consistency constraint.
[0131] As an option, the partition consistency constraint condition It can be calculated based on the boundary consistency between modal data. The specific formula is:
[0132] ;
[0133] where, , , are the observation values of the infrared image, visible light image, and lidar respectively.
[0134] After the Lagrangian modeling is completed, the objective function can be iteratively optimized by the gradient descent method. Specifically, the update formula for the partition boundary is:
[0135] ;
[0136] where, represents the partition boundary at the -th iteration; represents the learning rate, which is used to control the step size of the gradient descent; represents the gradient of the Lagrangian objective function.
[0137] Generally, the choice of the learning rate has an important impact on the optimization convergence speed and stability. In one possible implementation, a method of dynamically adjusting the learning rate can be adopted, such as accelerating convergence by gradually decreasing the learning rate.
[0138] During the optimization process, it is also necessary to dynamically adjust the modal weights in the objective function to improve the balance of the influence of different modal data on the partition result. The dynamic adjustment formula for the modal weights is:
[0139] ;
[0140] where, represents the weight of the modality ; represents the noise level of the modality ; represents the index variable, which is used to traverse the data uncertainty of all modalities; represents the standard deviation of the -th modality (or data channel), which measures the dispersion or uncertainty of the data of this modality.
[0141] As an option, the modal noise level can be calculated from the standard deviation or variance estimate of the multi-modal data, so as to reflect the signal-to-noise ratio of different modal data.
[0142] In some embodiments, to improve the robustness of partition optimization, the partition boundary can be dynamically adjusted in combination with a boundary smoothing term. Specifically, the optimization objective of the boundary smoothing constraint is to minimize the gradient change of the partition boundary, and its optimization direction is:
[0143] ;
[0144] where, represents the smoothed boundary at the -th iteration; represents the learning rate of the smoothing term; represents the boundary smoothing term.
[0145] In another implementation, the partition result can be corrected in combination with a modal interaction verification term. The goal of modal interaction verification is to ensure the consistency of different modal data at the partition boundary, and its correction method is based on the boundary probability distribution between modalities. The specific formula is:
[0146] ;
[0147] where, is the posterior probability when the boundary condition is ; represents the likelihood probability of the modal data under the boundary condition; represents the prior probability of the boundary; represents the joint probability of the modal data.
[0148] Through modal interaction verification, abnormal regions in the partition result can be identified and corrected.
[0149] The optimization iteration process usually ends when the following stopping conditions are met:
[0150] The change amplitude of the partition boundary is less than a preset threshold;
[0151] The energy value of the objective function converges to a local or global minimum;
[0152] The number of iterations reaches a preset upper limit.
[0153] In summary, this step realizes the consistent fusion of multi-modal data, the optimization of the smoothness of the partition boundary, and the guarantee of the reliability of modal interaction verification through the optimization of the objective function. Through reasonable optimization strategies and parameter settings, the accuracy and robustness of the partition result can be ensured, providing a reliable basis for the subsequent output of the partition result.
[0154] For S5, it is to output the target area partition based on the optimization result. The main task of this step is to organize the optimized partition result and output it in an appropriate form for subsequent application scenarios, such as target area monitoring, security, and fire monitoring, etc. Generally, the output partition result includes the partition boundary and the pixel set within the partition, aiming to provide descriptions of the spatial location, shape, and target area characteristic data of the partition. This step needs to be closely connected with the previous optimization processing step (S4) and further improve the usability of the partition result.
[0155] Specifically, in this embodiment, step S5 includes the following content:
[0156] Generally, the optimized partition boundary is represented in the form of pixel-level boundary information. For the convenience of subsequent application processing, the partition boundary can be further converted into a polygon form or a point set form. In one implementation, the polygon representation of the partition boundary can be achieved through a boundary tracing algorithm. For example, an algorithm based on Canny edge detection can extract the connected point set of the partition boundary and further generate a polygon contour representation. Specifically, the partition boundary can be represented as:
[0157] ;
[0158] where, represents the vertex coordinates on the partition boundary; represents the number of vertices of the partition boundary.
[0159] As an option, the partition boundary can also be smoothed to further eliminate the boundary noise or jagged phenomenon that may occur during the optimization process. Generally, the smoothing process can adopt Gaussian filtering or B spline fitting methods. Specifically, Gaussian filtering eliminates local irregular points by performing a convolution operation on the coordinates of the boundary points, while B spline fitting describes the partition boundary by generating a smooth curve.
[0160] In the generation of partition area characteristic data, the main task is to analyze and statistically process the pixel set within the target area, so as to extract area features and associate them with the partition boundary. Generally, the partition area characteristic data includes the average characteristic value, standard deviation, and pixel distribution information within the area. For example, for infrared image data, the average thermal radiation intensity of the partition area can be calculated by the following formula:
[0161] ;
[0162] where, represents the average thermal radiation intensity of the target area; Expressed as the number of pixels within the partitioned region; Expressed as the observed value of the infrared image data at the pixel point location.
[0163] As an option, the distribution characteristics of the regional characteristics can be further extracted, such as calculating the variance or frequency distribution diagram of the pixel values within the region. Through the generation of these characteristic data, a more detailed description can be provided for the monitoring or analysis of the target region. In the result formatted output, the partition result needs to be converted into an output format suitable for the specific application scenario. For example, in the monitoring of the target region, the partition result can be output in the form of an image overlay, that is, the partition boundary is overlaid with the original infrared image or visible light image to generate a partition annotation map. In one implementation, the partition annotation map can be achieved through the following formula:
[0164] ;
[0165] where, represents the output image; the original image (infrared image or visible light image); represents the color value or intensity increment used to annotate the partition boundary.
[0166] In another implementation, the partition result can be output in the form of vector data, such as GeoJSON format or SVG format, for spatial analysis or visualization processing. The advantage of this formatted output is that it can be easily integrated with a geographic information system or other spatial data analysis tools.
[0167] Through the above representation of the partition boundary, generation of the regional characteristic data, and result formatted output, the efficient output and flexible application of the partition result are ultimately achieved. In some embodiments, the output partition result can be directly used for the classification, status monitoring, or risk assessment of the target region. For example, in the fire monitoring application, based on the distribution of the thermal radiation intensity in the partitioned region, the high-temperature hot spot regions can be quickly identified and warning signals can be issued.
[0168] In summary, this step not only enhances the usability of the partition data by converting the partition optimization result into a partition result form that is easy to apply and understand, but also provides reliable technical support for the practical application of the target region. Through the precise representation of the partition boundary and the generation of the regional characteristic data, this step is of great significance in the entire partition data processing method and is the key link to achieve the precise analysis and processing of the target region.
[0169] Please refer to Figure 2 , the present invention also provides a partition data processing system for an infrared image of a specific region, including:
[0170] Multimodal sensor module, used to collect infrared images, visible light images, and lidar point cloud data;
[0171] Data calibration module, used to perform spatial registration and time synchronization on the collected multimodal data:
[0172] Partition modeling module, used to construct the partition processing objective function for the target area:
[0173] Optimization and solution module, used to model the objective function based on the Lagrange multiplier method, and iteratively optimize the partition boundary through the gradient descent method, dynamically adjust the modal weights, and optimize the target area partition;
[0174] Partition verification module, used to evaluate the reliability of the partition result through modal interaction verification:
[0175] Result output module, used to output the partition boundary of the target area and the pixel set within the partition. The partition result can be used in scenarios such as monitoring, security, or fire monitoring of the target area.
[0176] Multimodal sensor module
[0177] Used to collect multimodal data of a specific area, including infrared images, visible light images, and lidar point cloud data, and is the core module for data input in the system. The infrared image sensor collects the thermal radiation information of the target area for monitoring the temperature distribution; the visible light image sensor obtains the color and texture characteristics of the area for identifying surface details; the lidar sensor captures the three-dimensional depth information of the target area for constructing the spatial geometric model of the area. Through the collaborative work of these sensors, it can provide comprehensive and multi-dimensional raw data support for subsequent data calibration and partition processing.
[0178] The data calibration module is responsible for performing spatial registration and time synchronization on the data collected by the multimodal sensors to ensure that different modal data can be processed and analyzed in the same coordinate system. The spatial transformation matrix between the infrared image and the visible light image is obtained through the calibration plate calibration method, and the lidar point cloud data is mapped to the image plane coordinate system through the projection matrix to achieve precise spatial alignment. At the same time, data time synchronization is achieved through the hardware trigger signal or timestamp alignment method, and combined with noise filtering and interpolation processing to improve the consistency and quality of the data, providing high-precision input data for subsequent partition modeling.
[0179] Partition modeling module
[0180] The partitioning processing objective function for constructing the target area is the mathematical basis for the system to achieve partitioning optimization. The objective function consists of three parts: a data consistency term, a boundary smoothness term, and a modal interaction verification term, which are used to measure the consistency of multimodal data within the partitioned area, constrain the smoothness of the partition boundary, and evaluate the interaction consistency of data between modalities, respectively. By reasonably setting the weight parameters of each item, the characteristic requirements of the partitioned area can be comprehensively reflected, providing a clear evaluation criterion and optimization direction for subsequent optimization processing.
[0181] Optimization Solving Module
[0182] Responsible for optimizing and solving the objective function to determine the optimal partition boundary of the target area. This module models the objective function through the Lagrange multiplier method, introduces partition consistency constraints, and uses the gradient descent method to iteratively optimize the partition boundary. At the same time, the module can dynamically adjust the modal weights according to the noise level of the modal data to optimize the influence of different modal data on the partition result. Through continuous iterative optimization until the energy value of the objective function converges to the optimum, a high-precision partition boundary is generated, laying a foundation for partition verification and result output.
[0183] Partition Verification Module
[0184] Used to evaluate the reliability of the partition result to ensure that the finally output partition boundary can accurately reflect the consistency of multimodal data. This module analyzes the performance of the partition boundary under different modalities through modal interaction verification, calculates the posterior probability distribution of the partition result using the Bayesian inference model, identifies abnormal areas, and provides correction suggestions, thereby improving the robustness and credibility of the partition result.
[0185] Result Output Module
[0186] Used to output the final partition result in a form suitable for the application scenario, including the partition boundary and the characteristic data within the partitioned area. The partition boundary can be output in the form of polygons, point sets, or vector data (such as GeoJSON or SVG formats) for visualization and further analysis. The characteristic data of the partitioned area includes thermal radiation intensity, texture characteristics, and depth information, etc., providing direct decision-making support for applications such as monitoring, security, and fire monitoring. Through flexible output formats and content, this module effectively enhances the practicality and expandability of the partition result.
[0187] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for processing partition data of infrared images of a specific area, characterized in that: The following steps are involved: Acquire data from multimodal sensors; Calibrate and preprocess the acquired multimodal data; constructing an objective function of the partitioning process, wherein the objective function is used to measure the characteristics of the partitioned area; The objective function includes a data consistency term, a boundary smoothing term and a modal interaction verification term, and its expression is: E(Ω,Γ)=∫ Ω f data (x)dx+λ∫ Γ f smooth (x)dx+μ∫ Ω f modal (x)dx; Where E(Ω, Γ) represents the total energy of the objective function; Ω represents the partition area; Γ represents the partition boundary; f data (x) represents the data consistency term, which is used to measure the characteristic consistency of different modal data in the target area; f smooth (x) represents the boundary smoothing term, which is used to constrain the smoothness of the partition boundary; f modal (x) represents the modal interaction verification term, which is used to measure the data consistency between modalities; λ represents the weight of the smoothing term; μ represents the weight of the modal interaction verification term; The data consistency term is calculated according to the dynamic modal weight, and the expression of the data consistency term is: Among them, f data (x) indicates data consistency item; M i (x) represents the data of mode i; f(x) represents the characteristic estimation value of the target area in the partition; α i represents the weight of mode i; Optimize the objective function and determine the partition boundaries; Output the target region partition based on the optimization results.
2. The method for processing regional data of infrared images of a specific area according to claim 1, characterized in that: The multimodal sensor comprises: Infrared image sensor, used to collect thermal radiation information; Visible light image sensor for collecting texture and color information; LiDAR sensor, used to collect depth and spatial information.
3. The method for processing regional data of infrared images of a specific area according to claim 1, characterized in that: The multimodal data calibration includes: Spatial registration and temporal synchronization of trimodal data: Spatial registration realizes the registration of infrared images and visible light images through feature matching algorithm: The LiDAR data is mapped to the image plane coordinate system through the point cloud projection matrix.
4. The method for processing regional data of infrared images of a specific area according to claim 1, characterized in that: The optimization process of the objective function comprises the following steps: The objective function is modeled based on the Lagrange multiplier method, and the partition consistency constraint is introduced. The expression of the objective function is: L(Ω,Γ,λ)=E(Ω,Γ)+λ∫ Ω g(x)dx; Among them, L(Ω, Γ, λ) represents the Lagrangian objective function; g(x) represents the consistency constraint condition, which is used to constrain the boundary consistency of data between modes; The partition boundary is iteratively solved using the gradient descent method. The update expression of the partition boundary is: Among them, Γ (k) represents the partition boundary at the kth iteration; η represents the learning rate; represents the gradient of the objective function; Dynamically adjust the modal weight in the objective function. The modal weight is determined according to the noise level of the modal data. The calculation formula of the modal weight is: Among them, α i represents the weight of mode i; σ i represents the noise level of mode i; j represents the index variable; σ j represents the standard deviation of the jth mode; The objective function is continuously optimized until the energy converges to the optimal value and the final partition boundary is generated.
5. The method for processing partition data of infrared images of a specific area according to claim 1, characterized in that: The partitioning process comprises the following steps: Determine the partition boundary of the target area and the pixel set within the partition; Generate corresponding characteristic data based on the partition result of the target area; The characteristic data is analyzed and processed and then outputted, and the partition results are used for monitoring, security or fire detection of the target area.
6. The method for processing regional data of infrared images of a specific area according to claim 1, characterized in that: The modal interaction verification item verifies the consistency of the segmented region by analyzing the boundary probability distribution between the modalities, and evaluates the consistency of the modal boundaries based on the Bayesian inference model.
7. The method for processing partition data of infrared images of a specific area according to claim 1, characterized in that: The target area partitioning comprises the following steps: Construct initial partition boundaries of the target area; Based on the characteristics of multimodal data, data consistency analysis is performed on the initial partition boundaries; Use optimization algorithms to dynamically adjust partition boundaries so that they match the characteristic distribution of the target area; The partition boundaries of the target area are modified in combination with the modal interaction verification results to generate the final partition results.
8. A partition data processing system for infrared images of a specific area, applied to a partition data processing method for infrared images of a specific area as claimed in any one of claims 1 to 7, characterized in that: include: Multimodal sensor module for collecting infrared images, visible light images and lidar point cloud data; Data calibration module, used to perform spatial registration and temporal synchronization on the collected multimodal data: Partition modeling module, used to construct the partition processing objective function of the target area: The optimization solution module is used to model the objective function based on the Lagrange multiplier method, and iteratively optimize the partition boundaries through the gradient descent method, dynamically adjust the modal weights and optimize the target area partitions; Partition verification module, used to evaluate the reliability of partition results through modal interaction verification: The result output module is used to output the partition boundary of the target area and the pixel set within the partition. The partition result can be used for monitoring, security or fire monitoring scenarios of the target area.
Citation Information
Patent Citations
Fan blade defect detection system and method based on multi-mode perception
CN119180793A