An abnormality identification method and system based on multi-modal sparse consistency representation
By mapping multimodal sensing data to a unified spatial region, constructing a sparse consistency description, and combining consistency indicators and temporal stability indicators, the problem of unstable local correspondence in multimodal anomaly identification is solved, achieving more stable anomaly identification and hierarchical risk assessment.
Patent Information
- Application Number
- CN202610604384.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-06
- Publication Date
- 2026-08-25
AI Technical Summary
Existing multimodal anomaly identification methods struggle to stably describe the local correspondence between different modalities, making the anomaly identification results susceptible to single-frame noise or modal fluctuations, and failing to adequately characterize cross-modal consistency.
By mapping different modal sensing data to a unified spatial region, a multimodal sparse consistency description is constructed, and anomaly risk judgment is made by combining cross-modal consistency indicators and temporal stability indicators.
It improves the stability and applicability of multimodal perception anomaly identification, reduces the impact of occasional interference on the identification results, and outputs a graded anomaly risk level for easier subsequent processing.
Smart Images

Figure CN122634136A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of multimodal sensing anomaly identification technology, specifically to an anomaly identification method and system based on multimodal sparse consistency representation. The method acquires different types of sensing data and performs spatial mapping, sparse feature extraction, consistency index calculation, and anomaly risk level determination to identify local or persistent inconsistencies in multimodal sensing data, providing technical support for the safe operation and anomaly handling of the sensing system. Background Technology
[0002] With the development of intelligent sensing technology, multimodal sensing systems have been widely used in scenarios such as drones, mobile robots, intelligent security, autonomous driving, and industrial inspection. Multimodal sensing systems typically perceive the external environment through various sensory information such as infrared images, depth vision images, point cloud data, or inertial measurement data, in order to compensate for the limitations of single sensors under conditions such as changes in lighting, occlusion, noise interference, or limited viewing angle.
[0003] Existing multimodal anomaly detection methods typically process sensory data using techniques such as feature stitching, weighted fusion, model reconstruction, or single-modal quality assessment. While these methods can improve anomaly detection capabilities to some extent, in practical applications, different modal sensory data often have different imaging mechanisms, spatial resolutions, and feature representations. For example, infrared data primarily reflects regional response intensity, thermal distribution characteristics, and significant regional changes, while visual data focuses more on spatial structure, distance variations, depth edges, and contour information. Due to the lack of a unified local event representation method among different modalities, existing methods struggle to stably describe the correspondence between different modalities within the same spatial region. This leads to anomaly detection results being easily affected by single-frame noise or modal fluctuations under conditions of local inconsistency, short-term interference, or persistent anomalies.
[0004] Therefore, it is necessary to propose a new anomaly identification method that maps different modal sensing data to a unified spatial region and uses sparse features, consistency indices, and continuous temporal variations to jointly describe the abnormal states between modalities, so as to improve the stability and applicability of multimodal sensing anomaly identification. Summary of the Invention
[0005] The purpose of this invention is to provide an anomaly identification method and system based on multimodal sparse consistency representation, which solves the problems of unstable local correspondence of different modal sensing data, insufficient cross-modal consistency characterization, and single-frame anomaly judgment being easily affected by occasional interference in existing multimodal anomaly identification methods.
[0006] To achieve the above objectives, the technical solution adopted by the present invention is as follows: This method proposes an anomaly identification approach based on multimodal sparse consistency representation. By mapping different modal sensing data to a unified spatial region, a multimodal sparse consistency description is constructed. This description is then combined with cross-modal consistency indices and temporal stability indices to assess anomaly risk, thereby enabling the identification of multimodal sensing anomalies. The method includes the following steps: S1, acquire different modal sensing data at the same acquisition time or within the same acquisition time window, and map the different modal sensing data to a unified spatial region; S2, sparse feature extraction is performed on different modal sensing data within the unified spatial region to construct a multimodal sparse consistency description; S3, calculate the cross-modal consistency index based on the multimodal sparse consistency description; S4. Calculate the time series stability index based on the changes in cross-modal consistency index at multiple consecutive acquisition times; S5. Determine the anomaly risk based on the cross-modal consistency index and the temporal stability index, and output the anomaly identification result.
[0007] S1, acquire infrared image data and visual image data at the same acquisition time, and map the infrared image data and visual image data to a unified spatial region.
[0008] Optionally, the infrared image data includes infrared thermal imaging data, near-infrared image data, infrared response intensity map, or infrared target detection results; the visual image data includes visible light image data, depth visual data, target contour map, edge map, texture map, depth map, or target detection results.
[0009] Optionally, the unified spatial region is divided according to any one of the following methods: image plane region, forward field of view region, bird's-eye view region, or target candidate region.
[0010] In S1, the sensing range is divided into multiple unified spatial regions, and each unified spatial region is assigned a corresponding region number.
[0011] Optionally, the infrared sparse features include infrared response intensity, infrared salient region area, infrared edge intensity, infrared response gradient, infrared region contrast, infrared target contour integrity, and infrared quality features; the visual sparse features include image edge intensity, texture sharpness, target contour integrity, visual salient region area, effective depth ratio, nearest depth, spatial structure edge, local structural abrupt change, hole ratio, and visual quality features.
[0012] In S2, the multimodal sparsity consistency description includes a unified spatial region, infrared sparsity features, visual sparsity features, infrared quality features, visual quality features, infrared-visual difference features, and acquisition time.
[0013] Optionally, the infrared-visual consistency index includes the overlap between the infrared salient region and the visual structural region, the correspondence between the infrared response mutation region and the visual structural mutation region, the shape consistency between the infrared target contour and the visual target contour, and the infrared-visual difference features.
[0014] In S3, the area number is The unified spatial region in the first The infrared-visual consistency index at each acquisition time is defined as follows: The calculation formula is as follows:
[0015] in, This indicates the overall quality status of infrared image data and visual image data within the same area. This indicates the degree of overlap between the infrared salient region and the visual structural region. This indicates the degree of correspondence between regions of abrupt changes in infrared response and regions of abrupt changes in visual structure. This indicates the similarity between infrared sparse features and visual sparse features. Indicates infrared-visual difference features. , , , These are preset weighting coefficients.
[0016] Optionally, the temporal stability index is determined based on the mean of the infrared-visual consistency index, the variance of the infrared-visual consistency index, the number of frames with abnormality, the change in the number of abnormal spatial regions, and the change in the location of abnormal spatial regions within multiple consecutive acquisition times.
[0017] In S4, the selected area number is A unified spatial region in continuous Infrared-visual consistency indices within each acquisition time point constitute a temporal consistency sequence. The time-consistency sequence includes sequences from the first... From the first data collection moment to the first... Infrared-visual consistency index at each acquisition time. Temporal stability index calculated based on the temporal consistency sequence. The calculation formula is as follows:
[0018] in, and The preset adjustment coefficient, Indicates continuity The variance of the infrared-visual consistency index within each acquisition time point This indicates the degree to which the infrared-visual consistency index remains below a preset consistency threshold.
[0019] In step S5, anomaly risk values are calculated based on the infrared-visual consistency index and the temporal stability index. The calculation formula is as follows:
[0020] in, , , These are preset weighting coefficients.
[0021] Optionally, the abnormal risk level includes a normal level, a local inconsistency level, and a persistent inconsistency level.
[0022] When the abnormal risk value is lower than the first risk threshold, the abnormal risk level is determined to be normal; when the abnormal risk value is higher than the first risk threshold but lower than the second risk threshold, the abnormal risk level is determined to be local inconsistency; when the abnormal risk value is higher than the second risk threshold, or when the same unified spatial area is in an inconsistent state for multiple consecutive collection times, the abnormal risk level is determined to be continuous inconsistency.
[0023] On the other hand, embodiments of the present invention also provide an anomaly identification system based on multimodal sparse consistency, comprising: The data acquisition module is used to acquire infrared image data and visual image data at the same acquisition time. The spatial mapping module is used to map the infrared image data and the visual image data to a unified spatial region according to the calibration parameters or spatial correspondence between the infrared image data and the visual image data. The sparse feature extraction module is used to extract infrared sparse features and visual sparse features within a unified spatial region and construct a multimodal sparse consistency description. The consistency index calculation module is used to calculate the infrared-visual consistency index based on the multimodal sparse consistency description. The temporal stability calculation module is used to calculate the temporal stability index based on the changes in the infrared-visual consistency index at multiple consecutive acquisition times. The anomaly risk output module is used to determine the anomaly risk level based on the infrared-visual consistency index and the temporal stability index, and output the anomaly identification result.
[0024] Through the above technical solutions, the anomaly identification method and system based on multimodal sparsity consistency provided by the present invention can map data of different modalities to a unified spatial region, and determine the anomaly risk level based on sparse features, consistency indicators and temporal stability indicators, thereby improving the stability of multimodal perception anomaly identification.
[0025] Compared with the prior art, the beneficial effects of the present invention are: First, the present invention can compare sensing data of different modalities within a unified spatial region, thereby reducing the local correspondence instability caused by differences in imaging mechanisms, field of view positions, spatial resolutions and feature representations among different modalities.
[0026] Second, this invention constructs a multimodal sparse consistency description by using sparse features, quality features and difference features of different modalities, which can structurally express the state differences between different modalities in a local region;
[0027] Third, this invention uses cross-modal consistency index to characterize the correspondence between different modalities in salient regions, structural regions, response mutation regions and target contours, avoiding the need to rely solely on the quality of a single modality or the overall fusion result for anomaly judgment;
[0028] Fourth, this invention introduces a temporal stability index for multiple consecutive acquisition times, which can distinguish between single-frame noise, short-term occlusion, instantaneous environmental changes and continuous anomalies, reducing the impact of occasional interference on anomaly identification results.
[0029] Fifth, this invention outputs an anomaly risk level, transforming the anomaly identification result from a single normal or abnormal judgment into a graded result, which facilitates subsequent perception processing, alarm prompts, or security decision-making modules to perform graded processing. Attached Figure Description
[0030] Figure 1 This is a flowchart illustrating an anomaly identification method based on multimodal sparse consistency representation according to the present invention. Figure 2 This is a schematic diagram illustrating the mapping of multimodal sensing data to a unified spatial region in this invention; Figure 3 This is a schematic diagram illustrating the structure of the multimodal sparsity consistency description in this invention; Figure 4 This is a schematic diagram illustrating the determination of anomaly risk levels based on cross-modal consistency indices and temporal stability indices in this invention. Figure 5 This is a schematic diagram of the structure of an anomaly recognition system based on multimodal sparse consistency according to the present invention. Detailed Implementation
[0031] The specific embodiments of the present invention will be further described below with reference to the accompanying drawings. It should be understood that the following embodiments are only used to illustrate and explain the examples of the present invention, and are not intended to limit the scope of protection of the present invention.
[0032] It should be noted that the acquisition, transmission, storage, use, and processing of image data in the embodiments of this invention comply with the provisions of relevant laws and regulations. The image preprocessing, feature extraction, spatial mapping, index calculation, and risk level determination methods involved in the embodiments of this invention are all used to illustrate the feasibility of the technical solution of this invention, and do not limit this invention to the use of a specific image acquisition device, feature extraction algorithm, or data processing platform.
[0033] like Figure 1 As shown, the anomaly identification method of the present invention includes a data input layer, a spatial mapping layer, a sparse consistency description layer, a cross-modal consistency analysis layer, a temporal feature analysis layer, and an anomaly risk output layer. Specifically, the data input layer is used to acquire multimodal sensing data within the same acquisition time moment or the same acquisition time window; the spatial mapping layer is used to map different modal sensing data to a unified spatial region; the sparse consistency description layer is used to extract sparse features of different modalities within the unified spatial region and form a multimodal sparse consistency description; the cross-modal consistency analysis layer is used to calculate the correspondence between different modal sensing data; the temporal feature analysis layer is used to extract temporal features based on consistency changes over multiple consecutive acquisition times; and the anomaly risk output layer is used to determine the anomaly risk level based on the consistency analysis results and the temporal features.
[0034] Example 1:
[0035] This embodiment illustrates the specific execution process of the method of the present invention during data synchronization and spatial mapping, including the following steps: In this embodiment, infrared image data and visual image data are used as two types of perception data for explanation. Infrared image data includes infrared thermal imaging data, near-infrared image data, infrared response intensity maps, or infrared target detection results; visual image data includes visible light image data, depth visual data, target contour maps, edge maps, texture maps, depth maps, or target detection results.
[0036] like Figure 2 As shown, infrared image data and visual image data are first acquired at the same acquisition time. Since the acquisition devices, imaging mechanisms, acquisition frequencies, and field of view of the two types of data may differ, time synchronization and spatial mapping are required before performing consistency calculations. Let the acquisition time of the infrared image data be... The acquisition time of visual image data is When both conditions are met At that time, the infrared image data and visual image data are used as a multimodal sensing data pair at the same acquisition moment. This is the preset time synchronization threshold.
[0037] After time synchronization is completed, the infrared image data and visual image data are mapped to the same reference space according to the calibration parameters or spatial correspondence between the infrared imaging device and the visual imaging device. The calibration parameters may include intrinsic parameters, extrinsic parameters, coordinate transformation matrices, or pixel correspondence between the infrared imaging device and the visual imaging device; the spatial correspondence may include image plane correspondence, forward field of view correspondence, bird's-eye view correspondence, or target candidate region correspondence.
[0038] After mapping, the reference space is divided into multiple unified spatial regions, and a corresponding region number is assigned to each unified spatial region. For region numbered... A unified spatial region, in the first The infrared image data and visual image data corresponding to each acquisition time are respectively represented as follows: and .
[0039] Infrared and visual imaging devices are fixedly mounted on the same miniature UAV platform, and their spatial correspondence is obtained through offline calibration. When the UAV performs inspection or target recognition tasks, the infrared and visual image data are mapped to the same reference plane and divided into multiple local regions of a preset size. For the same region number, the corresponding infrared and visual image data are extracted to form data pairs for subsequent sparsity consistency analysis. This method can reduce erroneous matching caused by sensor installation deviations, field-of-view differences, and sampling delays.
[0040] Example 2:
[0041] like Figure 3 As shown, after spatial mapping is completed, sparse feature extraction is performed on the infrared image data and visual image data within the unified spatial region. For the region numbered... The unified spatial region, in the first At each acquisition time, the infrared sparse features are represented as follows: Visual sparse features are represented as .
[0042] The infrared sparse features include infrared response intensity, infrared salient region area, infrared edge intensity, infrared response gradient, infrared region contrast, infrared target contour integrity, and infrared quality features; the visual sparse features include image edge intensity, texture sharpness, target contour integrity, visual salient region area, effective depth ratio, nearest depth, spatial structure edge, local structural abrupt change, hole ratio, and visual quality features.
[0043] To enable comparison of infrared sparse features and visual sparse features at the same feature scale, normalization and dimension alignment are performed on both to obtain aligned infrared sparse features. and visual sparsity features :
[0044] in, and These represent feature mapping functions for infrared sparse features and visual sparse features, respectively.
[0045] The feature mapping function includes linear normalization, standardization, feature selection, dimension completion, or linear projection. Taking linear normalization as an example, the feature components... It can be handled as follows:
[0046] in, and These represent the maximum and minimum values of the same type of characteristic components, respectively. To prevent constants with a denominator of zero.
[0047] Based on the normalized infrared sparse features, visual sparse features, and modal quality information, a multimodal sparsity consistency description is constructed. :
[0048] in, This represents the regional information of the unified spatial area. This represents the overall quality factor of the multimodal sensing data within the region.
[0049] The comprehensive quality factor The comprehensive quality factor can be determined based on the effective data ratio, edge integrity, target contour integrity, image sharpness, noise level, or hole ratio of infrared and visual image data within the corresponding region. The value range can be [0,1], with higher values indicating higher usability of multimodal sensing data within that region. When the infrared image data within a certain region has high regional contrast and relatively complete edges, while the visual image data shows a clear target contour and a high effective depth ratio, the comprehensive quality factor of that region is... The overall quality factor should be relatively high. When visual image data exhibits motion blur, overexposure in strong light, depth holes, or broken contours, even if the infrared image data is relatively clear, the overall quality factor will be lower. This also reduces the impact of low-quality modes on subsequent consistency judgments.
[0050] Example 3: This embodiment further calculates the cross-modal consistency index, such as Figure 4 As shown, the cross-modal consistency index is used to characterize the degree of correspondence between different modal sensing data within the same unified spatial region. To calculate this index, the infrared salient region and the visual structural region are first determined. The infrared salient region is represented as... It can be determined by infrared response intensity, infrared edges, infrared salient regions, or infrared target detection results; the visual structural region is represented as... It can be determined by visual edges, target contours, depth edges, spatial structure regions, or visual target detection results.
[0051] The degree of overlap between the infrared salient region and the visual structural region is expressed as:
[0052] in, This represents the area of the intersection region between the two. This represents the area of the region where the two are joined.
[0053] The similarity between infrared sparse features and visual sparse features is represented as follows:
[0054] in, For scaling parameters, This represents the L2 norm.
[0055] Calculate the cross-modal consistency index based on regional overlap, feature similarity, and overall quality factor: ,
[0056] in, , These are preset weighting coefficients.
[0057] In the above formula, Characterizes the spatial correspondence of different modes within the same region. Characterize the similarity relationships between sparse features of different modalities. This is used to adjust consistency metrics based on data quality. Therefore, without directly relying on the complete fusion network, it is possible to determine whether different modalities maintain consistency based on spatial overlap, feature similarity, and data quality status in local areas.
[0058] When the target recognition area of a drone is interfered with by infrared patches, local heat sources, or strong heat-reflecting materials, abnormally significant areas may appear in the infrared image data, while the corresponding target outline or structural edge does not appear at the corresponding location in the visual image data. reduce, Reduce, and thus make Decrease. When visual image data is affected by strong light, occlusion, motion blur, or depth measurement failure, the edge intensity, target contour integrity, or effective depth ratio in the visual image data decreases. and The similarity between them decreases, and the cross-modal consistency index decreases. The corresponding decrease is also expected.
[0059] Example 4: After obtaining the cross-modal consistency index, further time-series feature calculations are performed. For example... Figure 4 As shown, the selected area number is A unified spatial region in continuous Cross-modal consistency metrics within each acquisition time point constitute a temporal consistency sequence. The time-consistent sequence includes from the first... From the first data collection moment to the first... The infrared-visual consistency index at each acquisition time. The mean and variance of the temporal consistency sequence are expressed as follows:
[0060]
[0061] in, This indicator is used to characterize the temporal fluctuation of cross-modal consistency metrics within the region. The proportion of times the consistency index remains below a preset consistency threshold is represented as:
[0062] in, To preset a consistency threshold, This is an indicator function that takes the value 1 when the condition within the parentheses is true, and 0 otherwise. Used to characterize the persistence of low consistency states within a region.
[0063] Calculate the anomaly risk value based on the cross-modal consistency index, temporal variance, and the proportion of low consistency persistence:
[0064] in, , , For preset weighting coefficients, This indicates the degree of cross-modal inconsistency at the current acquisition time.
[0065] The abnormal risk level is represented as follows:
[0066] in, The area code is The unified spatial region in the first The level of abnormal risk at each data collection time. Indicates a normal level. Indicates the level of local inconsistency. Indicates a level of persistent inconsistency. This indicates the first risk threshold. This indicates the second risk threshold. This indicates the preset continuous percentage threshold.
[0067] Can be set = 5、 = 0.6、 = 0.3、 = 0.6、 = 0.6. When a region has a high cross-modal consistency index at the current acquisition time, and low consistency states do not persist within consecutive acquisition times, the anomaly risk value is... If the value is below the first risk threshold, the area is classified as normal. When a cross-modal consistency index briefly decreases in an area due to drone attitude jitter, momentary occlusion, short-term lighting changes, or single-frame sensor noise, the risk level drops. The risk level is relatively low, and the anomaly is classified as a local inconsistency level. When a region is affected by continuous heat source interference, continuous obstruction, sensor malfunction, or physical disturbances, causing the cross-modal consistency index to fall below a preset consistency threshold for multiple consecutive data collection times, the risk level is considered low. The risk level has increased, and the abnormal risk level has been determined to be a persistent inconsistency level.
[0068] Example 5: This invention also provides an anomaly detection system based on multimodal sparse consistency. For example... Figure 5 As shown, the system includes a data acquisition module, a spatial mapping module, a sparse feature extraction module, a consistency description construction module, a consistency index calculation module, a time series feature calculation module, and an anomaly risk output module.
[0069] The data acquisition module is used to acquire infrared image data and visual image data at the same acquisition time or within the same acquisition time window, and to perform time synchronization processing on the two.
[0070] The spatial mapping module is connected to the data acquisition module. It is used to map infrared image data and visual image data to a unified spatial region based on the calibration parameters or spatial correspondence between them, and output the corresponding region number. and .
[0071] The sparse feature extraction module is connected to the spatial mapping module for use in extracting sparse features from... Extracting infrared sparse features and from Extracting visual sparse features .
[0072] The consistency description construction module is connected to the sparse feature extraction module to normalize or dimensionally align infrared and visual sparse features and construct a multimodal sparse consistency description.
[0073] The consistency index calculation module is connected to the consistency description construction module, and is used to calculate the consistency index based on the regional overlap. Feature similarity and comprehensive quality factor Calculate the cross-modal consistency index .
[0074] The time series feature calculation module is connected to the consistency index calculation module, and is used to calculate the time series feature based on the continuous time series feature. Calculation of time-series variance of cross-modal consistency index within each acquisition time point and low consistency persistence ratio .
[0075] The anomaly risk output module is connected to the time series feature calculation module and is used to calculate the cross-modal consistency index. Time series variance and low consistency persistence ratio Calculate the abnormal risk value and output the abnormal risk level. .
[0076] The modules described above can be deployed in the same edge computing device, or they can be distributed across front-end sensing devices, edge computing devices, and back-end servers. The data acquisition module and spatial mapping module can be deployed at the front end or the edge, while the consistency index calculation module, time series feature calculation module, and anomaly risk output module can be deployed at the edge or the server.
[0077] In summary, this invention, through multimodal sensing data synchronization, unified spatial region mapping, sparse consistency representation construction, cross-modal consistency index calculation, and temporal feature analysis, forms an anomaly identification method applicable to multimodal sensing scenarios such as drones, mobile robots, intelligent security, and industrial inspection. This method can output corresponding anomaly risk levels when sensor anomalies, environmental occlusion, short-term interference, local response anomalies, or physical adversarial disturbances cause local or persistent inconsistencies in different modal sensing data, providing a basis for subsequent sensing processing, alarm prompts, and security decisions.
[0078] Although the present invention has been described above with reference to specific embodiments, various modifications can be made to it without departing from the principles of the invention, and some technical features can be replaced in an equivalent manner. As long as there is no technical conflict, the various features disclosed in this invention can be combined with each other in any way. This invention is not limited to the specific embodiments disclosed herein, but includes all technical solutions falling within the scope of the claims.
Claims
1. An anomaly identification method based on multimodal sparse consistency, characterized in that, include: Acquire different modal sensing data at the same acquisition time or within the same acquisition time window, and map the different modal sensing data to a unified spatial region; Sparse features are extracted from different modal sensing data within the unified spatial region, and the extracted sparse features are normalized or dimensionally aligned. Based on the unified spatial region, the different modal sparsity features after normalization or dimension alignment, the comprehensive quality factor, and the acquisition time, a multimodal sparsity consistency description is constructed. Based on the multimodal sparse consistency description, calculate the cross-modal consistency index between different modal sensing data; The temporal stability index is calculated based on the cross-modal consistency index at multiple consecutive acquisition times. The anomaly risk level is determined based on the cross-modal consistency index and the temporal stability index, and the anomaly identification result is output.
2. The anomaly identification method based on multimodal sparsity consistency according to claim 1, characterized in that: The different modal sensing data include infrared image data and visual image data; The infrared image data includes infrared thermal imaging data, near-infrared image data, infrared response intensity map, and infrared target detection results; the visual image data includes visible light image data, depth visual data, target contour map, edge map, texture map, depth map, and target detection results.
3. The anomaly identification method based on multimodal sparsity consistency according to claim 1, characterized in that: Mapping the different modal sensing data to a unified spatial region includes: Time synchronization is performed based on the acquisition time of different modal sensing data; Based on the calibration parameters or spatial correspondence between different modal sensing devices, the time-synchronized data of different modal sensing devices are mapped to the same reference space; The same reference space is divided into multiple unified spatial regions, and each unified spatial region is assigned a region number.
4. The anomaly identification method based on multimodal sparsity consistency according to claim 2, characterized in that: The sparse features include infrared sparse features and visual sparse features; The infrared sparse features include infrared response intensity, infrared salient region area, infrared edge intensity, infrared response gradient, infrared region contrast, and infrared target outline integrity. The visual sparsity features include image edge intensity, texture clarity, target outline integrity, visually salient area, effective depth ratio, nearest depth, spatial structure edge, local structural abrupt change, and hole ratio.
5. The anomaly identification method based on multimodal sparsity consistency according to claim 4, characterized in that: Normalizing or dimensional alignment of the sparse features includes: For area code The unified spatial region, in the first At each acquisition time, the infrared sparse features are represented as follows: Visual sparse features are represented as ; To enable comparison of infrared sparse features and visual sparse features at the same feature scale, normalization and dimension alignment are performed on both to obtain aligned infrared sparse features. and visual sparsity features : ; in, and These represent the feature mapping functions for infrared sparse features and visual sparse features, respectively; The feature mapping function includes linear normalization, standardization, feature selection, dimension completion, or linear projection. Taking linear normalization as an example, it applies to the feature components... Normalization is performed: ; in, and These represent the maximum and minimum values of the same type of characteristic components, respectively. To prevent constants with a denominator of zero.
6. The anomaly identification method based on multimodal sparsity consistency according to claim 5, characterized in that: The multimodal sparse consistency description is expressed as: ; in, This represents the regional information of the unified spatial area. This represents the overall quality factor of the multimodal sensing data within this region; The comprehensive quality factor is determined based on at least one of the following: the proportion of effective data of different modal sensing data within the corresponding unified spatial region, edge integrity, target contour integrity, image sharpness, noise level, or hole ratio.
7. The anomaly identification method based on multimodal sparsity consistency according to claim 6, characterized in that: The cross-modal consistency index is determined based on regional overlap, feature similarity, and comprehensive quality factor; The degree of overlap of the regions is expressed as: ; in, Indicates the salient infrared region. Indicates the visual structural region; The feature similarity is expressed as: ; in, For scaling parameters; The cross-modal consistency index is expressed as: ; in, , These are preset weighting coefficients.
8. The anomaly identification method based on multimodal sparsity consistency according to claim 7, characterized in that: The time-series stability index is based on continuous The cross-modal consistency index is determined within each acquisition time point; continuous The mean cross-modal consistency index within each acquisition time point is expressed as: ; continuous The variance of the cross-modal consistency index within each acquisition time point is expressed as: ; The low consistency persistence ratio is expressed as: ; in, This indicates a preset consistency threshold. Indicates an indicator function, and This constitutes the aforementioned time series stability index.
9. The anomaly identification method based on multimodal sparsity consistency according to claim 8, characterized in that: Determining the anomaly risk level based on the cross-modal consistency index and the temporal stability index includes: According to the cross-modal consistency index Cross-modal consistency index variance and low consistency persistence ratio Calculate the anomaly risk value: ; in, , , Preset weighting coefficients; An abnormal risk level is determined based on the abnormal risk value, and the abnormal risk level includes a normal level, a partial inconsistency level, and a persistent inconsistency level.
10. An anomaly identification system based on multimodal sparse consistency, characterized in that, include: The data acquisition module is used to acquire different modal sensing data at the same acquisition time or within the same acquisition time window; A spatial mapping module, connected to the data acquisition module, is used to map the different modal sensing data to a unified spatial region; A sparse feature extraction module, connected to the spatial mapping module, is used to extract sparse features from different modal sensing data within the unified spatial region. The consistency description construction module, connected to the sparse feature extraction module, is used to construct a multimodal sparse consistency description based on a unified spatial region, different modal sparse features after normalization or dimensional alignment, a comprehensive quality factor, and the acquisition time. A consistency index calculation module, connected to the consistency description construction module, is used to calculate a cross-modal consistency index based on the multimodal sparse consistency description. A time series stability index calculation module, connected to the consistency index calculation module, is used to calculate the time series stability index based on the cross-modal consistency index of multiple consecutive acquisition times. An anomaly risk output module, connected to the time series stability index calculation module, is used to determine the anomaly risk level based on the cross-modal consistency index and the time series stability index, and output the anomaly identification result.