A multi-modal based abnormal behavior pattern recognition method and system

By acquiring data from multiple information sources in real time and generating modal health scores, and dynamically adjusting the weight of visual information, the problem of abnormal behavior identification in complex and ever-changing scenarios is solved, and accurate identification of high-risk preparatory behaviors is achieved under conditions where visual information is limited.

CN121637367BActive Publication Date: 2026-04-10SHENYANG HUAANXIN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHENYANG HUAANXIN TECH CO LTD
Filing Date
2026-02-05
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively integrate multi-source information and accurately identify covert abnormal behaviors in complex and variable scenarios, especially when changes in ambient lighting conditions lead to a decline in the quality of core visual information.

Method used

By acquiring data from multiple information sources in real time, a modal health score is generated. The modal health score of visual information is corrected based on ambient lighting information. The weight of each information source in the judgment of abnormal behavior is dynamically adjusted based on a preset weight allocation strategy to achieve adaptive information fusion.

Benefits of technology

Under conditions of limited visual information quality, it can more accurately identify high-risk preparatory behaviors that rely on subtle movements and are highly concealed, thus improving the accuracy and robustness of abnormal behavior identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121637367B_ABST
    Figure CN121637367B_ABST
Patent Text Reader

Abstract

The application provides an abnormal behavior pattern recognition method and system based on multi-modal, relates to the field of abnormal behavior pattern recognition, obtains data of multiple information sources in real time and generates a modal health score, corrects the modal health score of the visual information source according to the ambient light information, obtains a target modal health score, dynamically determines the weight of each information source in the abnormal behavior judgment based on a preset weight distribution strategy according to the target modal health score of the visual information source and the modal health score of the non-visual information source, and effectively identifies abnormal behavior by fusing and analyzing the monitoring data of each information source and the weight of each information source. The application significantly improves the accuracy and robustness of abnormal behavior recognition, especially under the condition that the visual information is severely limited, can more effectively find high-risk preparation behavior, and thus improves the overall performance and safety of the monitoring system.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of abnormal behavior pattern recognition, and in particular, to a multi-modal based abnormal behavior pattern recognition method and system. BACKGROUND

[0002] In environments requiring high security monitoring, such as prisons, it is crucial to continuously observe the behavior of personnel in order to maintain order and ensure safety. Current technologies for discovering abnormal behavior, particularly those relying on a single source of information (e.g., only through capturing actions or recognizing clothing), are difficult to meet the demand for multi-aspect analysis of behavior when faced with complex and variable scenes. For example, when the observed personnel is far away from the camera or is blocked by something, if the system relies on only one source of information, it is easy to miss key features, thus incorrectly judging abnormal behavior. At the same time, information from different sources (such as the movement of body skeleton, the characteristics of clothing, and how long they stay in a certain place) often varies in form, making it difficult to effectively combine them, thus failing to accurately discover abnormal behavior in advance.

[0003] In a prison environment, when the lighting conditions dynamically change due to environmental factors (such as strong light glare or low illumination), resulting in a continuous decline in the quality of core visual information sources (such as key points of the skeleton and clothing features), the information fusion mechanism of existing systems cannot effectively adapt to this complex scene with partial modal degradation. Specifically, the system fails to establish a mechanism that can automatically adjust its information analysis mode when detecting that the key visual modal is continuously degraded due to environmental factors, rather than simply discarding or low-weight processing these low-confidence visual signals, but can perform deeper relevance analysis and meaning reconstruction with reliable but insufficient information granularity of auxiliary modalities (such as stay time). The core problem faced by current technology is: how to design an adaptive information fusion strategy so that the system can recognize and utilize these originally considered as "noise" or "low-confidence" ambiguous visual signals when faced with dynamic decline in visual information quality due to changes in environmental lighting, and effectively combine them with reliable but insufficient information granularity of auxiliary modalities (such as stay time), so as to accurately identify those high-risk preparatory behaviors that rely on subtle actions and have high concealment under conditions of severely limited visual information. SUMMARY

[0004] The present application provides a multi-modal based abnormal behavior pattern recognition method and system, aiming to solve the technical problem that existing technologies are difficult to effectively fuse multi-source information and accurately identify concealed abnormal behavior in complex and variable scenes, particularly when the quality of core visual information declines due to changes in environmental lighting conditions.

[0005] In one aspect, the application provides a multi-modal based abnormal behavior pattern recognition method, comprising:

[0006] real-time acquisition of data of multiple information sources, generation of a modal health score of each information source representing data reliability thereof, the information sources including a visual information source and at least one non-visual information source;

[0007] According to the acquired environmental lighting information, the modal health score of the visual information source is corrected based on a preset correction rule to obtain a target modal health score of the visual information source, and the preset correction rule includes a mapping relationship between visual information quality and lighting conditions;

[0008] According to the target modal health score of the visual information source and the modal health score of the non-visual information source, a weight of the visual information source and the non-visual information source in abnormal behavior judgment is determined based on a preset weight distribution strategy, wherein the weight distribution strategy satisfies that the weight value is positively correlated with the modal health score of the corresponding information source, and the sum of all weights is a fixed value;

[0009] According to the monitoring data of each information source and the weight of each information source, fusion analysis is performed to recognize abnormal behavior.

[0010] In another aspect, the application provides a sealed adaptive soft-pack lithium battery packaging image detection system, which comprises:

[0011] A modal health score generation module is configured to acquire data of multiple information sources in real time, and generate a modal health score of each information source representing data reliability thereof, the information sources including a visual information source and at least one non-visual information source;

[0012] A correction module is configured to correct the modal health score of the visual information source based on a preset correction rule according to acquired environmental lighting information to obtain a target modal health score of the visual information source, and the preset correction rule includes a mapping relationship between visual information quality and lighting conditions;

[0013] A weight determination module is configured to determine a weight of the visual information source and the non-visual information source in abnormal behavior judgment based on a preset weight distribution strategy according to the target modal health score of the visual information source and the modal health score of the non-visual information source, wherein the weight distribution strategy satisfies that the weight value is positively correlated with the modal health score of the corresponding information source, and the sum of all weights is a fixed value;

[0014] An abnormal behavior recognition module is configured to perform fusion analysis according to the monitoring data of each information source and the weight of each information source to recognize abnormal behavior.

[0015] The application relates to a multi-modal based abnormal behavior pattern recognition method and system. By acquiring data from multiple information sources in real time and generating a modal health score, the reliability of each modal data can be objectively evaluated. Especially when the quality of visual information decreases due to changes in environmental light information, the application can correct the modal health score of the visual information source according to the environmental light information to obtain a target modal health score, thereby accurately reflecting the real usability of the visual information. On this basis, the application dynamically determines the weight of each information source in abnormal behavior judgment based on a preset weight distribution strategy according to the target modal health score of the corrected visual information source and the modal health score of the non-visual information source. This adaptive weight distribution mechanism enables the system to no longer simply discard or low-weight process low-confidence visual signals when facing complex scenes with degraded modalities, but can perform deeper relevance analysis and meaning reconstruction on the reliable but insufficient information granularity auxiliary modalities. Finally, by fusing and analyzing the monitoring data of each information source and the weight of each information source, the application can effectively recognize abnormal behavior.

[0016] Through the above technical solution, the application overcomes the limitations of single information source or fixed weight fusion strategy in the prior art, effectively solves the problem that when the quality of core visual information continuously decreases and becomes unreliable due to dynamic changes in lighting conditions in high-security monitoring environments such as prisons, existing systems cannot effectively adapt to such complex scenes with degraded modalities, thereby making it difficult to accurately identify high-risk preparation behaviors that rely on subtle actions and have high concealment. The application dynamically adjusts the information fusion strategy, enabling the system to recognize and utilize fuzzy visual signals that are originally considered as “noise” or “low confidence”, and effectively combines them with reliable but insufficient information granularity auxiliary modalities, significantly improving the accuracy and robustness of abnormal behavior recognition, especially in conditions with severely limited visual information, enabling more effective detection of high-risk preparation behaviors, thereby improving the overall performance and security of the monitoring system. BRIEF DESCRIPTION OF DRAWINGS

[0017] In order to more clearly illustrate the application, the drawings required in the embodiments will be briefly introduced. Obviously, other drawings can be obtained by those skilled in the art without creative labor.

[0018] Figure 1 Fig. 1 exemplarily shows a flow diagram of a multi-modal based abnormal behavior pattern recognition method;

[0019] Figure 2 Fig. 2 exemplarily shows a structural diagram of a multi-modal based abnormal behavior pattern recognition system.

[0020] Reference signs: 100, multi-modal based abnormal behavior pattern recognition system; 10, modality health score generation module; 20, correction module; 30, weight determination module; 40, abnormal behavior recognition module. DETAILED DESCRIPTION

[0021] The technical solutions in the present application will be described clearly and completely below in combination with the drawings in the present application. Obviously, the described embodiments are only some of the embodiments of the present application, rather than all the embodiments. The components of the present application described and shown in the drawings herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the claimed present application, but only represents selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present application.

[0022] It should be noted that: similar reference signs and letters represent similar items in the following drawings, therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. Meanwhile, in the description of the present application, the terms "first", "second" and the like are only used to distinguish the description, and cannot be understood as indicating or implying relative importance.

[0023] Traditional existing abnormal behavior recognition technology, especially those relying on a single source of information (for example, only through the capture of action or recognition of clothing), in the face of complex and variable monitoring scene, it is difficult to meet the demand of multi-aspect analysis of behavior. For example, when the observed person is far away from the camera or is blocked, the system of a single source of information is easy to miss the key features, resulting in the error of abnormal behavior judgment. At the same time, the information from different sources (such as body skeleton action, clothing characteristics and stay time) is different in form, making it difficult to effectively combine these information, so as to accurately discover abnormal behavior in advance. In the prison environment, when the light condition changes dynamically due to environmental factors (such as strong light glare or low illumination), the quality of the core visual information source (such as skeleton key point and clothing feature) continues to decline and is unreliable, the information fusion mechanism of the existing system cannot effectively adapt to this complex scene with partial modal degradation. Specifically, the system fails to establish a mechanism that can automatically adjust its information analysis mode when detecting that the key visual modal is continuously degraded due to environmental factors, instead of simply discarding or low-weight processing these low-confidence visual signals, but can perform deeper relevance analysis and meaning reconstruction with reliable but insufficient information granularity of auxiliary modal (such as stay time). The core problem faced by the current technology is: how to design an adaptive information fusion strategy, so that the system can identify and utilize these originally considered as "noise" or "low confidence" fuzzy visual signals when the quality of visual information is dynamically reduced due to environmental light changes, and effectively combine them with reliable but insufficient information granularity of auxiliary modal (such as stay time), so as to accurately identify those high-risk preparation behaviors that rely on subtle actions and have high concealment under the condition of severely limited visual information.

[0024] As shown in Figure 1 An exemplary flowchart of a multi-modal based abnormal behavior pattern recognition method is shown. The present application proposes a multi-modal based abnormal behavior pattern recognition method, comprising:

[0025] S10, real-time acquisition of data of a plurality of information sources, generating a modal health degree score representing the data reliability of each information source, the information sources including visual information sources and at least one non-visual information source;

[0026] Among them, the modal health degree score refers to the quantitative evaluation of the reliability or quality of the data of each information source. That is, for each information source, the modal health degree score is used to represent the quantitative score of its data reliability. This score aims to evaluate the reliability and quality level of the data provided by the information source at the current time or under certain conditions. The higher the score, the more reliable the data, the more suitable for the judgment of abnormal behavior; on the contrary, the lower the score, the poorer the data reliability, the lower the reference value in judgment.

[0027] For example, visual information sources generally refer to image or video data acquired through cameras, infrared sensors, etc., such as skeletal key points, clothing features, etc. Non-visual information sources include but are not limited to audio information, biometric information (such as heart rate, body temperature), environmental sensor data (such as dwell time, location information), etc. These information sources collectively form the basis of multi-modal data.

[0028] Specifically, the calculation basis, quantification method or evaluation standard of the modality health score can be comprehensively considered by those skilled in the art according to the characteristics of different information sources:

[0029] First intrinsic quality assessment:

[0030] For visual information sources (e.g., video streams, images), the modality health score thereof can be quantified according to intrinsic quality indicators of the image or video. For example, the sharpness of the image (such as through edge sharpness, blur detection), noise level (such as signal-to-noise ratio SNR, peak signal-to-noise ratio PSNR), contrast, brightness uniformity, color distortion degree, etc. These indicators can be calculated in real time through image processing algorithms. For example, a high-definition, low-noise video stream will obtain a higher modality health score, while a fuzzy, snowflake-filled video stream will obtain a lower score.

[0031] For non-visual information sources (which can include at least one of regional dwell time, audio, biometric data), the modality health score thereof can be quantified according to the completeness, accuracy, real-time nature of the data, and the performance indicators of the sensor itself. For example, for regional dwell time data, the signal strength of the data acquisition device (such as an RFID reader, a UWB locator), the packet loss rate of data transmission, the positioning accuracy, etc. can be evaluated. For audio information, background noise size, speech clarity, etc. can be evaluated. The more complete, accurate and real-time the data is, the better the sensor works, and the higher the modality health score is.

[0032] Second, environmental factors impact assessment: For visual information sources, environmental factors have a significant impact on their data reliability. This means that in addition to the intrinsic quality of the data itself, external environmental factors such as lighting conditions are also important factors in calculating or correcting the modal health score. In strong light or low light environments, even if the camera itself has good performance, the quality of the visual information it outputs will decrease sharply, resulting in a lower modal health score. In addition, factors such as occlusion (e.g. people being blocked by obstacles), weather conditions (e.g. rain, snow, fog) can also affect the reliability of visual information. For non-visual information sources, some non-visual information sources may also be affected by environmental factors. For example, sound sensors in noisy environments may have lower modal health scores; wireless positioning systems in environments with a lot of metal reflectors may have lower positioning accuracy, and their modal health scores may also decrease.

[0033] In terms of quantification and evaluation criteria, the modal health score can usually be quantified as a value between 0 and 1, where 1 represents completely reliable data or excellent quality, and 0 represents completely unusable data or extremely poor quality. Specific quantification methods can include at least one of the following methods:

[0034] Threshold-based rule system: A series of rules and thresholds are pre-set. For example, if the video signal-to-noise ratio is below a certain threshold, the modal health score decreases by one level; if the light intensity is below a certain threshold, the score decreases by another level.

[0035] Machine learning model-based: Train a model that takes various data quality indicators and environmental parameters as input and outputs the modal health score. For example, a large amount of visual data of different quality can be collected and manually labeled for reliability, and then a classifier or regression model can be trained to predict the modal health score.

[0036] Weighted sum: Quantify multiple factors (such as clarity, noise, light intensity, occlusion rate) and assign different weights, then perform a weighted sum to obtain the final modal health score.

[0037] S20, according to the acquired environmental lighting information, the modal health score of the visual information source is corrected based on the pre-set correction rule to obtain the target modal health score of the visual information source, and the pre-set correction rule includes the mapping relationship between the visual information quality and the lighting condition;

[0038] S30, determining the weight of the visual information source and the non-visual information source in the abnormal behavior judgment based on the preset weight allocation strategy according to the target modality health score of the visual information source, the modality health score of the non-visual information source, wherein the weight allocation strategy satisfies that the weight value is positively correlated with the modality health score of the corresponding information source, and the sum of all weights is a fixed value;

[0039] S40, performing fusion analysis according to the monitoring data of each information source and the weight of each information source to identify abnormal behavior.

[0040] The present application effectively solves the limitations of the prior art in abnormal behavior recognition in complex and variable environments by introducing modality health score, ambient light correction, adaptive weight allocation and multi-modal fusion analysis. Especially when the quality of visual information is limited, it can more accurately and robustly identify high-risk preparation behavior.

[0041] The multi-modal based abnormal behavior pattern recognition method proposed in the present application aims to integrate and intelligently analyze data from different information sources to improve the accuracy and robustness of abnormal behavior recognition.

[0042] Firstly, the data of multiple information sources need to be obtained in real time. For example, the data of visual information source can be obtained through the camera deployed in the monitoring area, which can be video stream or image sequence. At the same time, the microphone can be deployed to obtain audio information, or the RFID tag, UWB positioning can be used to obtain the position and stay time of personnel as non-visual information source data. While obtaining these data, a modality health score will be generated for each information source. For example, for visual information source, the modality health score can be calculated by analyzing the sharpness, noise level, occlusion condition and other factors of the image. For non-visual information source, such as stay time data, the modality health score can be generated by evaluating the stability of data acquisition device, the integrity of data transmission, etc.

[0043] Secondly, the modality health score of the visual information source will be corrected based on the preset correction rule according to the obtained ambient light information to obtain the target modality health score of the visual information source. The ambient light information can be obtained through the light sensor deployed in the monitoring area, or estimated by analyzing the visual information source itself (such as the brightness histogram of the image).

[0044] In some embodiments, the preset correction rule can be a function or a lookup table that adjusts the visual modality health score based on the current lighting conditions. In combination with real-time acquired environmental lighting information, the process of adjusting the modality health score of the visual information source. The core purpose is to more accurately reflect the actual reliability of visual information under the current environmental conditions, because the change of lighting conditions will directly affect the quality of visual data. For example, in low light or strong light glare conditions, the reliability of the visual information source will decrease, so the correction rule will adjust the original visual modality health score downward, thereby obtaining a lower target modality health score. Conversely, in ideal lighting conditions, the correction rule may not make significant adjustments to the score.

[0045] Specifically, the process of correcting the visual score is as follows:

[0046] First, acquire environmental lighting information. The system will acquire the lighting information of the current monitoring environment in real time. This can be achieved in various ways, for example, by deploying special lighting sensors (such as photoresistors, light meters) to measure the lux value of the environment; or the system can also indirectly estimate the current lighting conditions by analyzing the visual information source itself (for example, extracting the brightness histogram, average pixel intensity or the proportion of overexposure / underexposure areas from the video frames).

[0047] Second, apply the preset correction rule. Once the environmental lighting information is acquired, the system will call the preset correction rule. These rules usually exist in the form of mathematical functions, lookup tables or conditional logic, which define how to adjust the original visual information source modality health score under different lighting conditions.

[0048] Among them, the design of the preset correction rule is based on the understanding of the relationship between the quality of visual information and lighting conditions. For example, in extremely low light or strong glare (overexposure) environment, the quality of visual information (such as skeleton key point extraction, clothing feature recognition) will decrease significantly, and its reliability will decrease. Therefore, the correction rule will instruct the system to reduce the original modality health score under these adverse lighting conditions. In ideal or moderate lighting conditions, the quality of visual information is usually high, and the correction rule may only make minor adjustments or no adjustments.

[0049] The correction method can be linear, nonlinear, or piecewise function. For example, when the light value is below a certain threshold, the modality health score decreases by a certain percentage for every unit decrease in light; when the light value is above a certain threshold and causes overexposure, the score also decreases accordingly.

[0050] Finally, after the preset correction rule processing, the original visual information source modal health score is adjusted to the target modal health score of the visual information source. This target score more truly reflects the availability and reliability of visual information in the current environment, providing a more accurate basis for subsequent weight allocation and abnormal behavior judgment.

[0051] For example, assume that the initial modal health score of the visual information source is 0.8 (full score 1.0), indicating that its data reliability is relatively high under standard conditions.

[0052] Scenario One: Low-light Environment

[0053] The system detects that the current ambient light intensity is 50 Lux, which is much lower than the required 200 Lux for normal operation. The preset correction rule can specify: "When the light intensity is lower than 200 Lux, reduce the visual modal health score by 0.1 for every 50 Lux decrease." According to this rule, the light intensity decreases from 200 Lux to 50 Lux, a total of 150 Lux (3 intervals of 50 Lux). Correction calculation: original score 0.8-(30.1)=0.5. Finally, the target modal health score of the visual information source is corrected to 0.5, reflecting the significant decrease in the reliability of visual information in low light.

[0054] Scenario Two: Strong Light Glare Environment

[0055] The system detects that there is a large area (e.g., more than 30%) of overexposed area in the image, which is usually caused by strong direct sunlight or reflection. The preset correction rule can specify: "When the overexposed area of the image accounts for more than 20%, reduce the visual modal health score by 0.2; if it exceeds 40%, reduce it by 0.4." According to this rule, the overexposed area accounts for 30%, which is between 20% and 40%, the system may use interpolation or directly apply a reduction of 0.2. Correction calculation: original score 0.8-0.2=0.6. Finally, the target modal health score of the visual information source is corrected to 0.6, indicating that the reliability of visual information (such as detailed features) has decreased in strong light glare.

[0056] Scenario Three: Ideal Light Environment

[0057] The system detects that the current ambient light intensity is 300 Lux, and the image quality analysis shows no obvious overexposure or underexposure. The preset correction rule can specify: "When the light intensity is between 200-500 Lux and the image quality is good, do not correct or fine-tune." Correction calculation: original score 0.8-0=0.8. Finally, the target modal health score of the visual information source remains 0.8, reflecting the high reliability of visual information in ideal light.

[0058] Through this dynamic adjustment based on preset correction rules, the system can more accurately assess the actual availability of visual information, thereby more reasonably allocating the weights of different modal information in subsequent abnormal behavior judgment, improving the accuracy and robustness of overall recognition.

[0059] After obtaining the target modality health score, the weight of the visual information source and the non-visual information source in abnormal behavior judgment is determined based on the target modality health score of the visual information source and the modality health score of the non-visual information source, and a preset weight allocation strategy. The weight allocation strategy can be a machine learning model. For example, when the target modality health score of the visual information source is low, the weight of the visual information source in abnormal behavior judgment is reduced, and the weight of the non-visual information source is correspondingly increased. Conversely, when the target modality health score of the visual information source is high, the weight of the visual information source is increased. This dynamic weight allocation mechanism enables adaptive adjustment of the degree of dependence on different modal information to cope with fluctuations in information quality caused by environmental changes.

[0060] Finally, the monitoring data of each information source and the weight of each information source are fused and analyzed to identify abnormal behavior. Fusion analysis can use a variety of techniques, such as feature-level fusion, decision-level fusion, or hybrid fusion. In feature-level fusion, raw data or extracted features from different information sources are integrated into a unified feature vector, which is then input into an abnormal behavior recognition model (such as a support vector machine, neural network, etc.) for classification. In decision-level fusion, each information source independently judges abnormal behavior, and then the weighted voting or weighted average of these judgment results is performed according to their respective weights to obtain the final abnormal behavior recognition result. For example, when the weight of the visual information source is high, its influence on the final abnormal behavior judgment is greater; when the weight of the non-visual information source is high, its influence on the final abnormal behavior judgment is greater. Through this fusion analysis, the advantages of multi-modal information can be utilized to improve the accuracy and robustness of abnormal behavior recognition.

[0061] In some embodiments, the preset weight allocation strategy can dynamically adjust the importance of different information sources (including the visual information source and at least one non-visual information source) in abnormal behavior judgment according to their data reliability, i.e., modality health score. This strategy is predefined and can be determined during system design phase and adjusted according to real-time data during operation. This strategy can generally take the following two main forms:

[0062] One is a rule-based weight allocation strategy. This strategy determines the weight allocation by setting a series of explicit rules and thresholds. The system compares the target modality health score of the visual information source and the modality health score of the non-visual information source with the pre-set thresholds, and then executes the corresponding weight adjustment rules. In principle, the weight allocation strategy satisfies: the weight value is positively correlated with the modality health score of the corresponding information source, and the sum of all weights is a fixed value;

[0063] The specific way to determine the weight includes 1) defining the modality health score interval: dividing the modality health score into different levels, such as high, medium and low. 2) setting weight adjustment rules: preset a set of weight allocation schemes for each modality health score level combination (for example, high visual modality health and medium non-visual modality health). 3) Dynamic adjustment: when the target modality health score of the visual information source is low (for example, due to insufficient environmental light, the quality of visual information decreases), the system will reduce the weight of the visual information source in abnormal behavior judgment according to the pre-set rules, and correspondingly increase the weight of the non-visual information source to make up for the deficiency of visual information. Conversely, when the target modality health score of the visual information source is high, its weight will increase. Generally, the sum of the weights of all information sources will remain a constant (for example, 1) to ensure that the total judgment influence does not change.

[0064] The second is a weight allocation strategy based on a machine learning model: this strategy uses a machine learning model (such as a neural network, support vector machine or decision tree) to learn the relationship between modality health scores and optimal weight allocation in historical data.

[0065] The specific way to determine the weight includes 1) data collection and annotation: collect a large amount of historical monitoring data, including modality health scores under different environmental conditions, actual occurrence of abnormal behavior and corresponding accurate judgment results. These data are used to train the model. 2) Model training: train a machine learning model so that it can output an optimal weight combination according to the input target modality health score of the visual information source and the modality health score of the non-visual information source. The training goal of the model is to maximize the accuracy of abnormal behavior recognition while minimizing false positives and false negatives. 3) Real-time prediction: during system operation, the real-time acquired modality health scores are used as the input of the model, and the model will predict and output the best weight of each information source under the current situation.

[0066] For example, assume that in a prison environment, the system mainly relies on visual information sources (such as camera-captured personnel posture, action) and a non-visual information source (such as area sensor-recorded personnel stay time).

[0067] Scenario one: bright day

[0068] Target modality health score of visual information source: High (e.g., 0.9) because the lighting is good, the camera images are clear, and the skeletal key points are extracted accurately. Modality health score of non-visual information source: High (e.g., 0.95) because the area sensors are working stably.

[0069] Rule-based strategy: The preset rule might state that when the visual modality health score is higher than a certain threshold (e.g., 0.7), the weight of visual information is 0.7 and the weight of non-visual information is 0.3. At this time, the system relies more on visual information to judge abnormal behavior.

[0070] Machine learning-based strategy: The trained model receives (0.9, 0.95) as input, and the output weight might be visual 0.75 and non-visual 0.25, indicating that visual information contributes more under good conditions.

[0071] Scenario Two: Night with dim light

[0072] Target modality health score of visual information source: Low (e.g., 0.3) because the lighting is insufficient, the camera images are blurry, and the skeletal key points are difficult to extract. Modality health score of non-visual information source: High (e.g., 0.95) because the area sensors are not affected by light.

[0073] Rule-based strategy: The preset rule might state that when the visual modality health score is lower than a certain threshold (e.g., 0.5), the weight of visual information is reduced to 0.2 and the weight of non-visual information is increased to 0.8. At this time, the system relies more on non-visual information such as staff stay time to judge abnormal behavior, such as long stay in sensitive areas.

[0074] Machine learning-based strategy: The model receives (0.3, 0.95) as input, and the output weight might be visual 0.15 and non-visual 0.85, indicating that non-visual information becomes the main basis for judgment when visual is limited.

[0075] Scenario Three: Camera failure or obstruction

[0076] Target modality health score of visual information source: Very low (e.g., 0.1), even close to 0. Modality health score of non-visual information source: High (e.g., 0.95).

[0077] Rule-based strategy: The preset rule might state that when the visual modality health score is lower than an extremely low threshold (e.g., 0.2), the weight of visual information is almost 0 (e.g., 0.05) and the weight of non-visual information is close to 1 (e.g., 0.95). The system almost completely relies on non-visual information for judgment.

[0078] Machine learning-based strategy: the model receives (0.1, 0.95) as input, and the output weight may be visual 0.02, non-visual 0.98, and the system will mainly rely on non-visual information for anomaly judgment.

[0079] Through this preset weight distribution strategy, the method can intelligently adjust the relative importance of each information source in anomaly behavior judgment according to the real-time changing modal health score, so as to maintain high anomaly behavior recognition accuracy and robustness in complex and variable environments, especially when the quality of part of the modal information decreases.

[0080] The overall working principle of the present application is to introduce the modal health score and the ambient light correction mechanism, which can evaluate and dynamically adjust the reliability of different information sources in real time. When the visual information source decreases in quality due to changes in ambient light, its modal health score will be corrected, so that in the subsequent weight distribution link, the weight of the visual information source will be correspondingly reduced, and the weight of the non-visual information source will be increased. This adaptive weight distribution strategy ensures that even in the case of partial modal information damage, more judgment basis can be transferred to the more reliable modal. Finally, through weighted fusion analysis of the monitoring data of each information source, the advantages of multi-modal information can be utilized to effectively identify abnormal behavior. For example, in a prison environment, when the camera visual information is blurred due to insufficient night light, the weight of the visual information will be reduced, and more reliance will be placed on non-visual information such as personnel stay time and location information to determine whether there is abnormal behavior, thereby avoiding false positives or false negatives due to the decrease in the quality of a single modal information.

[0081] Compared with the prior art, the core innovation of the present application is its adaptive information fusion strategy. Traditional methods often simply discard or low-weight process low-confidence visual signals, resulting in a significant decrease in recognition ability when visual information is limited. By introducing the modal health score and ambient light correction, the present application can more finely evaluate the usability of visual information. More importantly, by dynamically adjusting the weight of each modal in anomaly behavior judgment, the present application can effectively combine fuzzy visual signals with reliable but information-granularity-lacking auxiliary modalities (such as stay time), so that in the case of severely limited visual information, it can still accurately identify high-risk preparation behaviors that rely on subtle actions and have high concealment. This adaptive and intelligent fusion mechanism significantly improves the accuracy and robustness of anomaly behavior recognition, especially in complex and variable environments, where its advantages are more obvious.

[0082] In some embodiments, after the step of determining the weights of the visual information source and the non-visual information source in the abnormal behavior judgment based on the preset weight allocation strategy according to the target modality health score of the visual information source and the modality health score of the non-visual information source, the method further comprises the steps of:

[0083] determining a judgment threshold of an auxiliary information source in the non-visual information source according to the degree of reduction in visual information availability reflected by the target modality health score of the visual information source.

[0084] Specifically, the degree of reduction in visual information availability can be understood as the degree of reduction in the ability of the visual information source to provide effective information under specific environmental or conditions. This degree is usually directly reflected by the target modality health score of the visual information source, for example, the lower the target modality health score, the higher the degree of reduction in visual information availability. The auxiliary information source in the non-visual information source refers to other modality data sources that can provide additional information to assist in judging abnormal behavior in addition to the visual information source, such as sound sensors, infrared sensors, millimeter wave radars, etc. The judgment threshold refers to the critical value that the auxiliary information source relies on to judge whether there is an abnormal behavior. When the monitoring data of the auxiliary information source exceeds or is lower than the threshold, it is considered that there may be an abnormal behavior. Determining the judgment threshold aims to make the judgment of the auxiliary information source more accurate and adaptive to make up for the impact of the reduction in visual information availability.

[0085] The technical solution of the present application realizes dynamic and adaptive adjustment of the abnormal behavior judgment logic by associating the degree of reduction in visual information availability reflected by the target modality health score of the visual information source with the judgment threshold of the auxiliary information source in the non-visual information source. When the visual information availability is reduced, the judgment threshold of the auxiliary information source can be adjusted accordingly according to the degree of reduction. For example, if the visual information is severely damaged, the judgment threshold of some auxiliary information sources can be reduced to make them more sensitive to potential abnormal signals, so that abnormal behavior clues can still be captured in time in the case of insufficient visual information. Conversely, if the visual information availability is high, the judgment threshold of the auxiliary information source can be appropriately increased to reduce false positives. This dynamic adjustment mechanism enables more flexible and intelligent use of non-visual information in multi-modal information fusion analysis, makes up for the deficiency of visual information, and ensures the robustness of abnormal behavior recognition.

[0086] By the technical solution, the application can effectively solve the problem that abnormal behavior recognition accuracy may be reduced by only relying on weight adjustment when visual information availability is reduced. By dynamically determining the determination threshold of the auxiliary information source according to the degree of reduction of the visual information availability, the sensitivity of the non-visual information source can be more finely controlled, and false negatives or false positives caused by missing or blurred visual information can be avoided. This significantly improves the adaptability and reliability of the abnormal behavior pattern recognition method in complex and variable environments, so that high recognition performance can be maintained under different lighting, shielding and other conditions, thereby improving the overall abnormal behavior detection efficiency and accuracy.

[0087] For example, assuming that in a monitoring scene, the target modality health score of the visual information source (such as a camera) is significantly reduced at night or in dim light, reflecting a higher degree of reduction of visual information availability. At this time, the determination threshold of the non-visual information source (such as a sound sensor or an infrared sensor) will be dynamically adjusted according to this degree of reduction. Specifically, if the visual information is almost unavailable, the determination threshold of the sound sensor for detecting abnormal sound (such as glass breaking sound or violent impact sound) can be appropriately lowered to make it easier to trigger an alarm, even if the sound intensity is slightly lower than the normal threshold. At the same time, the determination threshold of the infrared sensor for detecting abnormal heat source movement can also be adjusted to improve its sensitivity to potential intruders. In this way, even in the case of limited visual information, abnormal behavior such as nighttime intrusion or destruction can be timely and accurately identified by enhancing the reliance on and sensitivity to non-visual auxiliary information.

[0088] In some embodiments, the step of identifying abnormal behavior according to the fusion analysis of the monitoring data of each information source and the weight of each information source comprises:

[0089] extracting potential activity clues from the signals of the visual information source;

[0090] identifying abnormal behavior when the potential activity clues and the auxiliary information source in the non-visual information source with high weight appear at the same time.

[0091] Specifically, through image processing, video analysis and other technologies, specific visual patterns or events that may indicate the occurrence of abnormal behavior are identified and extracted from the data obtained from visual devices such as cameras and infrared sensors, realizing the extraction of potential activity clues from visual information sources. For example, it can include detecting changes in posture, abnormal movement trajectories, the appearance or disappearance of specific objects, and regional intrusions. These clues may not be sufficient to conclusively determine abnormalities, but they are preliminary indications of abnormal behavior. Among them, when the potential activity clues appear at the same time as the auxiliary information sources in the non-visual information sources with high weight, the abnormal behavior is identified, and after the visual potential activity clues are extracted, the non-visual information sources related to the clues are further checked. The non-visual information sources here, for example, can be sound sensors, vibration sensors, access control systems, environmental sensors, etc. Non-visual information sources with high weight refer to non-visual modalities that are given higher credibility or importance in abnormal behavior judgment. Auxiliary information sources specifically refer to non-visual modalities that can supplement or verify visual clues. When visual clues and abnormal signals indicated by these high-weight auxiliary non-visual information sources are synchronized or strongly associated in time or space, it is determined to be abnormal behavior. For example, when visual detection detects someone staying in a sensitive area for a long time (potential activity clues), and the sound sensor detects abnormal impact or alarm sound (high-weight non-visual auxiliary information), it can be more conclusively identified as abnormal behavior.

[0092] In some embodiments, sensitive areas refer to specific spatial ranges in a particular monitoring environment that are considered to have a higher risk of abnormal behavior or require special attention due to their inherent physical characteristics, functional attributes, or historical experience. The characteristics of these areas are that once the activity patterns or stay duration of personnel in the area exceed the norm, it may indicate potential danger or abnormal events.

[0093] The determination of sensitive areas is closely related to the specific application scenario and environmental layout. For example, in high-security monitoring environments, such as warehouses, valuable item storage areas, equipment operation areas, entrances and exits, etc. may also be defined as sensitive areas, because these areas may involve unauthorized operations.

[0094] The determination of sensitive areas also needs to be combined with the assessment of potential risks. By analyzing historical event data, such as the location of past abnormal behavior, areas with high false positives, etc., areas that need to be monitored can be identified. For example, if a certain area has repeatedly occurred where personnel stay for a long time before an abnormal event occurs, that area should be marked as a sensitive area.

[0095] In actual deployment, sensitive areas are usually configured and divided according to preset rules and the experience of domain experts. For example, security experts will manually or semi-automatically demarcate these areas in the monitoring system according to their understanding of the security needs of a particular place. These rules may include:

[0096] Physical boundaries: areas defined by explicit walls, doors, windows, and other physical structures.

[0097] Functional areas: areas with specific uses, such as rest areas, workstations, and areas near storage cabinets.

[0098] Hidden areas: areas with poor camera visibility, corners with insufficient light, etc.

[0099] The technical solution of the present application effectively solves the false positive or false negative problems that may exist in traditional fusion analysis by refining the abnormal behavior recognition process into two stages. First, potential activity clues are extracted from visual information sources, which enables preliminary focus on areas or events that may have abnormalities, avoiding the computational burden of blindly fusing all data. Second, by requiring these potential activity clues to appear simultaneously with high-weight auxiliary information sources in non-visual information sources, the technical solution introduces a key verification mechanism. Because the weights of non-visual information sources have been optimally allocated according to their modal health scores, high-weight non-visual information sources are considered more reliable judgment criteria in the current environment. Therefore, when visual clues are corroborated by these reliable non-visual auxiliary information, the accuracy of abnormal behavior judgment will be significantly improved, thereby avoiding misjudgment that may result from a single modality or simple fusion.

[0100] Through the above technical solution, the present application can significantly improve the accuracy and robustness of abnormal behavior pattern recognition. By introducing the extraction of potential activity clues, it can detect abnormal signs earlier and reduce the processing of irrelevant information. More importantly, by coordinating visual clues with high-weight non-visual auxiliary information, false positives caused by single modality information are effectively avoided. Especially in cases where visual information is limited or ambiguous, non-visual auxiliary information can provide strong evidence, thereby ensuring the reliability of abnormal behavior recognition, reducing the false positive rate, and improving the ability to capture real abnormal events.

[0101] For example, suppose in a warehouse environment, the abnormal lingering behavior of personnel needs to be identified. Visual information sources may include surveillance cameras installed in the warehouse, and non-visual information sources may include sound sensors (for detecting abnormal noises such as glass breaking, lock picking), vibration sensors (for detecting abnormal vibrations of shelves or doors and windows), and RFID readers (for tracking item movement).

[0102] First, the video stream captured by the surveillance camera will be analyzed in real time to extract potential activity clues. For example, if a person is detected to stay in a non-working area for a long time, or there are abnormal rummaging actions near the shelves, these will be identified as potential activity clues.

[0103] At the same time, the modal health score of the visual information source will be corrected according to factors such as environmental lighting, and combined with the modal health scores of the sound sensor and vibration sensor to determine their weights in abnormal behavior judgment. Assuming that in the dark night, the weight of visual information is reduced, while the weights of sound sensor and vibration sensor are relatively increased.

[0104] When the visual system detects potential rummaging actions of a person in a sensitive area (such as next to a shelf storing valuable items), it will immediately check non-visual information sources with high weight. If the sound sensor detects abnormal metal rubbing sound or object falling sound, or the vibration sensor detects strong vibration of the shelf at this time, the potential activity clue of "person rummaging action" will be associated with the high-weight non-visual auxiliary information of "abnormal sound / vibration". Since both occur at the same time and confirm each other, it is identified as an abnormal behavior, and an alarm is triggered. This cooperative judgment mechanism significantly improves the accuracy of abnormal behavior recognition, avoiding false positives due to ambiguous visual images or single abnormal sound.

[0105] In some embodiments, the step of determining the decision threshold of the auxiliary information source in the non-visual information source according to the degree of reduction in visual information availability reflected by the target modal health score of the visual information source comprises:

[0106] Obtaining historical activity duration distribution data of a sensitive area where the person is located, and calculating an expected value of activity duration of the sensitive area according to the historical activity duration distribution data, wherein the sensitive area is a specific spatial range defined based on physical characteristics, functional attributes or historical experience, which has a higher risk of abnormal behavior or needs special attention in a specific monitoring environment;

[0107] According to the degree of reduction in visual information availability reflected by the target modal health score of the visual information source, a vigilance enhancement coefficient is calculated;

[0108] According to the expected value of activity duration of the sensitive area and the vigilance enhancement coefficient, a basic warning duration is generated;

[0109] According to the number of historical false positives of the sensitive area, a false positive suppression factor is adjusted;

[0110] According to the basic warning duration and the false positive suppression factor, the decision threshold of the auxiliary information source is determined.

[0111] Specifically, the stay or activity time records of the personnel in the specific monitoring area at different time periods are collected and analyzed to obtain the historical activity duration distribution data of the sensitive area where the personnel are located. For example, these data can be obtained through historical monitoring video analysis, access control system records, or sensor data in the area. According to the historical activity duration distribution data, the expected value of the activity duration of the sensitive area is calculated, which aims to quantify the typical activity pattern of the sensitive area under normal circumstances and provide an objective benchmark for subsequent anomaly judgment.

[0112] The alertness enhancement coefficient is calculated according to the degree of reduction in visual information availability reflected by the target modality health score of the visual information source. It can be understood that when the quality or availability of visual information decreases, the degree of dependence on non-visual information sources and alertness should be increased accordingly. For example, when the environmental light is insufficient, the visual sensor is blocked or fails, causing a significant reduction in visual information availability, the alertness enhancement coefficient will be calculated as a higher value to encourage more attention to non-visual information to make up for the lack of visual information.

[0113] In practical applications, the base alert duration is generated according to the expected value of the activity duration of the sensitive area and the alertness enhancement coefficient, which is used to combine the inherent activity characteristics of the area and the reliability of the current visual information to preliminarily set a benchmark time length for judging abnormal behavior. For example, if a region usually has a short activity time and the current visual information availability is low, the generated base alert duration may be set shorter to trigger the alert faster.

[0114] Further, the false alarm suppression factor is adjusted according to the number of historical false alarms in the sensitive area, which is used to dynamically adjust the sensitivity by learning from historical false alarms to reduce unnecessary alarms. For example, if false alarms frequently occur in a certain area under certain conditions, the false alarm suppression factor will be adjusted to a larger value, thereby increasing the judgment threshold, reducing the false alarm rate, and improving user experience.

[0115] Finally, the judgment threshold of the auxiliary information source is determined according to the base alert duration and the false alarm suppression factor. This judgment threshold is a dynamic value obtained by considering the activity characteristics of the area, the availability of visual information, and the historical false alarm situation, which is used to determine when the non-visual information source is recognized as an auxiliary information source, thereby triggering further judgment of abnormal behavior.

[0116] The technical solution of the present application solves the problem of insufficient accuracy caused by the simple threshold setting by introducing multi-dimensional consideration. First, by obtaining the historical activity duration distribution data of the sensitive area and calculating the activity duration expectation value, a baseline understanding of the "normal" behavior pattern of the area can be established, avoiding misjudgment of normal activities that meet the characteristics of the area as abnormal. Second, by calculating the alertness enhancement coefficient, the attention to non-visual information sources can be dynamically adjusted according to the real-time changes in visual information availability, ensuring that non-visual information can play a timely and effective role when visual information is limited. Third, by introducing a false alarm suppression factor and adjusting it according to the number of historical false alarms, the environment can be learned and adapted, effectively reducing false alarms caused by environmental noise or incidental events. Thus, the above technical solution considers the inherent characteristics of the environment, real-time sensing ability and historical experience, making the judgment threshold of auxiliary information sources more accurate and adaptive, thereby significantly improving the accuracy and reliability of abnormal behavior recognition.

[0117] Through the above technical solution, the present application can more accurately use non-visual information sources to make abnormal behavior judgments when visual information is limited, effectively reducing the false alarm rate and the missed alarm rate, and improving the overall performance and user experience of abnormal behavior recognition.

[0118] For example, assume that in a specific sensitive area of a warehouse, such as a valuable item storage area, abnormal behavior needs to be identified.

[0119] First, the personnel activity duration distribution data of the valuable item storage area in the past few months will be obtained. For example, by analyzing access control records and regional sensor data, it is found that the average activity duration of the area during weekdays (9:00-17:00) is 15 minutes, and the average activity duration at night (17:00-9:00 the next day) is 2 minutes. Thus, the activity duration expectation value of the area at different time periods can be calculated.

[0120] Second, when the insufficient night ambient light causes the target modality health score of the visual information source to decrease, reflecting that the degree of decrease in visual information availability is moderate, a corresponding alertness enhancement coefficient will be calculated, for example, 1.5.

[0121] Next, combining the activity duration expectation value at night (2 minutes) and the alertness enhancement coefficient (1.5), a basic alert duration will be generated, for example, 2 minutes * 1.5 = 3 minutes. This means that if the non-visual information source (such as an infrared sensor, a sound sensor) continuously detects activity in the area for more than 3 minutes, an alert may be triggered.

[0122] Further, the past one month history of false alarm times of this valuable item storage area will be consulted. If it is found that this area has had 3 false alarms caused by mouse activity at night, the false alarm suppression factor will be adjusted according to these historical false alarm data, for example, from the initial value 1.0 to 1.2.

[0123] Finally, the final auxiliary information source judgment threshold will be determined according to the generated basic alarm duration (3 minutes) and the adjusted false alarm suppression factor (1.2), for example, 3 minutes * 1.2 = 3.6 minutes. This means that only when the non-visual information source continuously detects activity in this area for more than 3.6 minutes will it be judged as an auxiliary information source and further trigger the abnormal behavior recognition process. In this way, the sensitivity and false alarm rate of abnormal detection can be more intelligently balanced, and the accuracy of recognition can be improved.

[0124] In some embodiments, the step of adjusting the false alarm suppression factor according to the historical false alarm times of the sensitive area comprises:

[0125] Obtaining the severity level of the alarm;

[0126] Evaluating the processing cost of the alarm;

[0127] Recording the response behavior of the management personnel to the alarm;

[0128] According to the severity level of the alarm, the processing cost of the alarm, and the response behavior, evaluating the fatigue state of the management personnel, and adjusting the false alarm suppression factor according to the fatigue state of the management personnel.

[0129] Specifically, the severity level of the alarm refers to the degree of harm determined after risk assessment of different types of abnormal behavior alarms, for example, it can be divided into three levels of low, medium and high, which aims to distinguish the urgency and potential impact of different alarms. The processing cost of the alarm can be understood as the resources required to process one alarm, including time, manpower, material resources, etc., for example, it can be quantitatively evaluated by statistical management personnel processing alarm average length of time, the number of security personnel needed to mobilize, etc. The response behavior of the management personnel to the alarm refers to the specific action taken by the management personnel after receiving the alarm, for example, including immediate verification, delayed processing, ignoring, etc., which aims to reflect the importance and processing efficiency of the management personnel to the alarm. Among them, the fatigue state of the management personnel refers to the physiological and psychological fatigue degree shown by the management personnel after a long time of work or frequent processing of alarms, for example, it can be evaluated by analyzing its historical response time, processing time, false alarm verification results, etc. Data is used to reflect the current workload and alertness level of the management personnel. In practical application, the false alarm suppression factor is a parameter used to adjust the alarm trigger threshold or alarm response intensity, and its adjustment aims to balance the false alarm rate and the missed alarm rate, to ensure that the optimal alarm performance can be provided in different situations.

[0130] The technical solution of the present application can more comprehensively evaluate the actual impact of the alarm and the real-time state of the management personnel by introducing the severity level of the alarm, the processing cost and the response behavior of the management personnel. Specifically, the severity level of the alarm and the processing cost are used to quantify the objective impact of the alarm, while the response behavior of the management personnel provides clues to its subjective processing efficiency and alertness. Based on these multi-dimensional information, the fatigue state of the management personnel can be more accurately evaluated. It is because the fatigue state of the management personnel is taken into account that the adjustment of the false alarm suppression factor is no longer dependent on the statistics of the historical false alarm times, but can dynamically adapt to the actual workload and alertness changes of the management personnel. When the management personnel is in a fatigue state, the strength of the false alarm suppression factor can be appropriately reduced to avoid neglecting the real alarm due to excessive suppression; on the contrary, when the management personnel is in good condition, the suppression strength can be appropriately increased to reduce unnecessary interference. This dynamic adjustment mechanism makes the false alarm suppression strategy more intelligent and humanized.

[0131] Through the above technical solution, the present application can overcome the limitations brought by the traditional method of adjusting the false alarm suppression factor only by the historical false alarm times. Specifically, by comprehensively considering the severity level of the alarm, the processing cost and the response behavior of the management personnel, the fatigue state of the management personnel can be more accurately evaluated, so that the adjustment of the false alarm suppression factor is more refined and adaptive. Therefore, not only can the false alarm rate be effectively reduced, but also the alarm fatigue of the management personnel can be reduced. This dynamic adjustment based on multi-dimensional information and personnel state significantly improves the intelligent level and actual application effect of the false alarm suppression strategy.

[0132] For example, assume in a security monitoring system, when an alarm is triggered, first the severity level of the alarm is obtained, for example, if a "person intrusion" is detected, the severity level is set as "high"; if an "object left" is detected, the severity level is set as "medium". At the same time, the cost required to process the alarm is evaluated, for example, a "high" level alarm may require immediate mobilization of two security personnel for on-site verification, and its processing cost is high; while a "medium" level alarm may only need remote video verification, and the processing cost is low. Then, the response behavior of the manager to the alarm is recorded, for example, whether the manager clicks the "verify" button within the specified time, or chooses to "ignore". Based on this information, the current fatigue state of the manager is evaluated. For example, if the manager has handled multiple "high" level alarms in a short period of time, and the response time is significantly prolonged, it can be judged that the manager is in a moderate fatigue state. According to this fatigue state, the false alarm suppression factor is dynamically adjusted. Specifically, if the manager is in a moderate fatigue state, the false alarm suppression factor can be slightly reduced, so that some alarms that may be suppressed originally but have certain potential risks can be triggered to avoid false negatives caused by fatigue; on the contrary, if the manager is in a good mental state, the false alarm suppression factor can be appropriately increased to reduce unnecessary interference. In this way, the false alarm suppression factor can be intelligently adjusted according to the actual impact of the alarm and the real-time state of the manager, so as to optimize the performance.

[0133] In some embodiments, the step of calculating the alertness improvement coefficient according to the degree of reduction in visual information availability reflected by the target modality health score of the visual information source comprises:

[0134] Obtaining a verification result of the alarm;

[0135] Recording the response action of the manager to the alarm;

[0136] Obtaining the degree of reduction in visual information availability when the alarm is triggered;

[0137] Calculating a fatigue influence coefficient according to the verification result of the alarm and the response action of the manager to the alarm;

[0138] Adjusting the calculation logic of the alertness improvement coefficient according to the fatigue influence coefficient and the degree of reduction in visual information availability when the alarm is triggered, to obtain the alertness improvement coefficient.

[0139] wherein, after verifying the triggered alarm, it is determined whether the alarm is a real abnormality (e.g., intrusion, fall) or a false alarm (e.g., pet activity, light change), thereby obtaining a verification result of the alarm. The verification result can serve as an important basis for evaluating accuracy and effectiveness of the manager's response. The specific actions taken by the manager after receiving the alarm are recorded, such as whether to view the monitoring screen, whether to dispatch security personnel, whether to ignore the alarm, etc. These response actions reflect the manager's attention to the alarm and processing efficiency. The degree of reduction in visual information availability refers to the degree of decline in the quality of visual information due to insufficient environmental lighting, obstruction, camera failure, etc. at the time of the alarm. This degree can directly affect the reliability of visual information sources in abnormal behavior judgment.

[0140] Further, according to the verification result of the alarm and the response action of the manager to the alarm, a fatigue influence coefficient is calculated. The fatigue influence coefficient aims to quantify the potential impact of the manager's fatigue state on alertness when handling alarms. For example, if the manager frequently ignores real alarms or overreacts to false alarms, it may indicate that he is in a state of fatigue, in which case the alertness needs to be appropriately increased. In practical applications, the fatigue influence coefficient can be evaluated by analyzing historical data, combined with the manager's identity information, work shift, historical alarm response time, false alarm handling time, and alarm verification result, etc.

[0141] On this basis, according to the calculated fatigue influence coefficient and the degree of reduction in visual information availability at the time of alarm triggering, the calculation logic of the alertness increase coefficient is adjusted. This means that the calculation of the alertness increase coefficient is no longer static, but dynamically adapts to the fatigue state of the manager and the actual availability of visual information. For example, when the fatigue influence coefficient is high, even if the degree of reduction in visual information availability is not high, the alertness increase coefficient may be increased to compensate for the possible negligence of the manager; conversely, when the degree of reduction in visual information availability is high, the alertness is further adjusted according to the fatigue influence coefficient to ensure that, in the case of limited visual information, non-visual information sources can be given a more reasonable weight allocation.

[0142] The technical solution of the present application calculates the fatigue influence coefficient by introducing the alarm verification result, the manager's response action, and the degree of reduction in visual information availability at the time of alarm triggering, and dynamically adjusts the calculation logic of the alertness increase coefficient. This enables the calculation of the alertness increase coefficient to more comprehensively consider dynamic factors in the actual operating environment, particularly the human factors of the manager and the real-time quality of visual information. In this way, the risk in the current situation can be more intelligently assessed, and the sensitivity to abnormal behavior can be adjusted accordingly, thereby avoiding false positives or false negatives due to the fatigue of the manager or fluctuations in the quality of visual information.

[0143] By the above technical solution, the calculation of the alertness enhancement coefficient becomes more refined and adaptive. It can dynamically adjust the alertness level according to the actual working state and historical performance of the manager, as well as the specific environmental conditions at the time of the alarm. This not only improves the accuracy and robustness of abnormal behavior pattern recognition, but also effectively reduces the false positive rate and false negative rate. This technical solution takes into account human factors and environmental factors in depth, enabling it to perform more outstandingly in complex and variable application scenarios.

[0144] For example, assume that in a sensitive area of a factory, insufficient environmental lighting is detected at night, resulting in a medium degree of reduction in visual information availability. At this time, the alertness enhancement coefficient needs to be calculated to adjust the decision threshold of the auxiliary information source.

[0145] First, the verification results of the recently triggered alarms are obtained. For example, in the past week, there have been 5 alarms, of which 3 have been verified as real abnormalities and 2 have been verified as false positives. At the same time, the manager's response actions to these alarms are recorded, for example, for real abnormalities, the manager timely checks and dispatches security personnel; for false positives, the manager labels after verification.

[0146] Next, according to these verification results and response actions, combined with the pre-set evaluation model, the current fatigue influence coefficient is calculated. For example, if the manager shows a longer response time or frequent neglect behavior when handling false positives, the fatigue influence coefficient will be calculated to be higher. Suppose the current calculated fatigue influence coefficient is 0.8 (range 0-1, the higher the value, the greater the fatigue influence).

[0147] Then, this fatigue influence coefficient 0.8 is combined with the degree of reduction in visual information availability (medium) at the time of the current alarm trigger to adjust the calculation logic of the alertness enhancement coefficient. For example, in the basic calculation logic, a medium degree of reduction in visual information availability may correspond to an alertness enhancement coefficient of 1.5. However, due to the high fatigue influence coefficient, the alertness enhancement coefficient will be further increased from 1.5 to 1.8 according to the adjustment logic to compensate for the possible fatigue-induced decrease in alertness of the manager.

[0148] Finally, this adjusted alertness enhancement coefficient 1.8 will be used to generate the subsequent basic alert duration and ultimately determine the judgment threshold of the auxiliary information source, so that in the case of possible fatigue and limited visual information of the manager, potential abnormal behaviors can be more timely and accurately identified.

[0149] In some embodiments, the step of calculating the alertness enhancement coefficient according to the degree of reduction in visual information availability reflected by the target modality health score of the visual information source comprises:

[0150] According to the target modality health score of the visual information source and the modality health score of the non-visual information source, a preliminary judgment of behavior intention is made to obtain a preliminary judgment result of behavior intention;

[0151] According to the preliminary judgment result of behavior intention, a corresponding visual information quality dependence degree is obtained;

[0152] According to the visual information availability reduction degree and the visual information quality dependence degree, the alertness degree promotion coefficient is calculated.

[0153] Specifically, on the basis of the target modality health score of the visual information source and the modality health score of the non-visual information source, a preliminary semantic analysis is performed on the currently observed activity to infer the potential behavior intention, so as to realize the preliminary judgment of behavior intention. For example, a multi-modal fusion model can be used to combine visual features (such as motion trajectory, posture) and non-visual features (such as sound, thermal imaging, environmental sensor data) to identify common behavior patterns such as "normal walking", "stopping", "wandering", "carrying objects", etc., so as to obtain the preliminary judgment result of behavior intention. The purpose is to provide more detailed context information for subsequent alertness adjustment.

[0154] Among them, obtaining the corresponding visual information quality dependence degree can be understood as pre-evaluating or dynamically evaluating the sensitivity of the behavior pattern to the quality of visual information for each behavior intention preliminary judgment result. For example, for "normal walking" behavior, its dependence on the quality of visual information may be low, because its main features can be effectively captured by non-visual information (such as thermal imaging, radar); while for "fine operation" or "object recognition" behavior, its dependence on the quality of visual information may be high. The dependence degree can be a preset numerical value, or a parameter dynamically adjusted according to historical data and expert experience. The purpose is to quantify the demand of different behaviors for visual information, so that the alertness degree promotion is more reasonable.

[0155] In actual application, according to the visual information availability reduction degree and the visual information quality dependence degree, the alertness degree promotion coefficient is calculated through a weighted or functional relationship. For example, the alertness degree promotion coefficient can be calculated as the product or weighted sum of the visual information availability reduction degree and the visual information quality dependence degree. When the visual information availability reduction degree is high and the behavior intention has a high dependence on the quality of visual information, the alertness degree promotion coefficient will be significantly improved; on the contrary, if the behavior intention has a low dependence on the quality of visual information, even if the visual information availability is reduced, the increase of the alertness degree promotion coefficient will be relatively small. The purpose is to realize more intelligent and adaptive alertness adjustment.

[0156] The technical solution of the present application introduces behavior intention preliminary judgment and visual information quality dependence degree, solves the limitation of calculating the alertness promotion coefficient only according to the reduction degree of visual information availability. Specifically, first, by fusing and analyzing the target modality health score of the visual information source and the modality health score of the non-visual information source, the behavior in the current scene can be preliminarily judged. This preliminary judgment provides important context information for subsequent alertness adjustment. Secondly, for different behavior intentions, the inherent dependence degree of visual information quality can be obtained. Because of the difference in the dependence degree of visual information of different behaviors, this factor is included in the calculation of alertness, so that the alertness promotion is no longer a single response to the decline of visual information quality, but combines the characteristics of the behavior itself. Therefore, when the visual information availability is reduced, the alertness promotion coefficient can be adjusted more accurately according to the behavior intention and its dependence on visual information, avoiding excessive or insufficient alertness, thereby improving the accuracy and efficiency of abnormal behavior recognition.

[0157] Through the above technical solution, the present application can realize more refined and intelligent alertness promotion. Specifically, by comprehensively considering the behavior intention and the dependence degree of visual information quality, it can avoid increasing the alertness when the visual information quality is reduced, but differentially process according to the actual situation. This not only helps to reduce false positives caused by environmental factors (such as light changes) and reduce the fatigue of managers, but also ensures that when a key abnormal behavior that really needs visual information to judge occurs, the alertness can be timely and effectively promoted, thereby improving the recognition accuracy and response efficiency of abnormal behavior.

[0158] For example, assume that in a warehouse environment, the target modality health score of the visual information source is reduced, indicating that the reduction degree of visual information availability is high, for example, due to dim light.

[0159] At this time, the behavior intention preliminary judgment will be made according to the target modality health score of the visual information source and the modality health score of the non-visual information source (for example, data from infrared sensors, sound sensors).

[0160] Scenario one: if the preliminary judgment result is "personnel normal patrol", and the preset dependence degree of "personnel normal patrol" behavior on visual information quality is low (because its main characteristics can be captured by infrared or sound sensors), then even if the reduction degree of visual information availability is high, the calculated alertness promotion coefficient will be relatively small.

[0161] Scenario two: if the preliminary judgment result is "person is performing fine operation in the equipment area", and the preset "fine operation" behavior has a high degree of dependence on visual information quality, even if the visual information availability reduction degree is the same as that in scenario one, the calculated alertness promotion coefficient will be significantly improved.

[0162] In this way, the alertness promotion coefficient can be dynamically and intelligently adjusted according to the specific behavior intention and its dependence on visual information, thereby avoiding excessive alertness in non-critical behaviors while ensuring sufficient vigilance in critical behaviors.

[0163] In some embodiments, the step of calculating the fatigue influence coefficient according to the verification result of the alarm and the response action of the management personnel to the alarm comprises:

[0164] According to the verification result of the alarm, the response action is classified to obtain a classification result;

[0165] According to the classification result and a preset fatigue weight, the fatigue influence coefficient is calculated.

[0166] The verification result of the alarm refers to the result obtained after investigating and verifying the triggered alarm, for example, the alarm is confirmed as a real abnormality, the alarm is confirmed as a false alarm, the alarm cannot be verified, etc. The response action of the management personnel to the alarm refers to the specific action taken by the management personnel after receiving the alarm, for example, immediately conducting on-site verification, remotely video confirmation, ignoring the alarm, upgrading the alarm to a higher level for processing, etc. According to the verification result of the alarm, the response action of the management personnel is classified into different preset categories, so as to classify the response action. For example, for the alarm verified as a false alarm, the "immediate on-site verification" action of the management personnel may be classified as the "excessive response" category, and the "ignore" action may be classified as the "reasonable handling" category. The classification result is the classified information. The preset fatigue weight is a numerical value preset for each classification result, which is used to quantify the influence of the response action of this category on the fatigue state of the management personnel. For example, frequent "excessive response" may be assigned a higher fatigue weight, and "reasonable handling" may be assigned a lower fatigue weight. The fatigue influence coefficient is a comprehensive index calculated according to the classification results and the corresponding fatigue weights, which reflects the current fatigue degree of the management personnel and its influence on the subsequent alarm handling.

[0167] The technical solution of the present application can more accurately quantify the fatigue state of the management personnel in the process of handling the alarm by classifying the verification results of the alarm and the response actions of the management personnel in detail and calculating in combination with the preset fatigue weight. Different types of response actions have different psychological and physiological loads on the management personnel under different verification results. For example, the energy input and the fatigue feeling required for handling an alarm verified as a real abnormality are significantly different from those required for handling an alarm verified as a false alarm. By taking these factors into consideration, it can be avoided to simply generalize all response actions, so that the calculation of the fatigue influence coefficient is more refined and contextualized. Thus, the actual working state of the management personnel can be more truly reflected, and a more reliable basis is provided for subsequent adjustment of the alertness degree improvement coefficient.

[0168] Through the above technical solution, refined evaluation of the fatigue state of the management personnel can be realized, and the calculation of the fatigue influence coefficient is more accurate and targeted. This helps to more intelligently identify the potential fatigue of the management personnel and adjust the alarm handling strategy accordingly, for example, appropriately reducing the priority of some non-critical alarms or increasing the auxiliary prompt for critical alarms when the fatigue degree of the management personnel is high, so as to effectively avoid false negatives or misjudgments caused by the fatigue of the management personnel. This deep consideration of human factors significantly improves the robustness and practicality of the multi-modal based abnormal behavior pattern recognition method, ensuring that it still maintains high efficiency and reliability in complex and variable application environments.

[0169] In some embodiments, the step of calculating the fatigue influence coefficient according to the classification result and the preset fatigue weight comprises:

[0170] Obtaining the identity information, the current working shift information and the historical fatigue accumulation data of the management personnel, wherein the historical fatigue accumulation data comprises the historical alarm response time, the historical false alarm handling time and the historical alarm verification result of the management personnel under different working shifts;

[0171] According to the identity information of the management personnel and the current working shift information, obtaining the basic fatigue weight corresponding to the management personnel and the working shift from the preset initial weight configuration;

[0172] According to the historical fatigue accumulation data of the management personnel, calculating a fatigue accumulation correction factor of the management personnel under the current working shift;

[0173] Adjusting the preset fatigue weight according to the basic fatigue weight and the fatigue accumulation correction factor;

[0174] Calculating the fatigue influence coefficient according to the classification result and the adjusted preset fatigue weight.

[0175] Specifically, the identity information of the manager can be any data used to uniquely identify the manager, such as an employee number, a username, etc. The current work shift information refers to a work time period or shift arrangement that the manager is currently performing, such as an early shift, a mid-shift, a night shift, etc. The historical fatigue accumulation data refers to historical data related to the fatigue state of a specific manager, which is recorded for a long time and used to quantify the fatigue degree of the manager. The data specifically includes the historical alarm response time, the historical false alarm processing time, and the historical alarm verification result of the manager under different work shifts. Among them, the historical alarm response time refers to the time interval from when the manager receives an alarm to when the manager takes the first response action; the historical false alarm processing time refers to the time spent by the manager in processing an alarm that is verified as a false alarm; and the historical alarm verification result records the case that the alarm is finally confirmed as a real anomaly or a false alarm. The preset initial weight configuration can be understood as a database or configuration table that stores the initial fatigue weights of different managers under different work shifts. The basic fatigue weight refers to an initial value obtained from the preset initial weight configuration as the starting point for adjusting the fatigue weight, according to the identity information of the manager and the current work shift information. The fatigue accumulation correction factor is a coefficient calculated according to the historical fatigue accumulation data of the manager, which is used to correct the basic fatigue weight, and its purpose is to reflect the fatigue degree accumulated by the manager due to long-term work or work in a specific shift. By combining the basic fatigue weight with the fatigue accumulation correction factor, the fatigue weight used to calculate the fatigue influence coefficient is dynamically updated, the preset fatigue weight is adjusted, and it is more suitable for the actual fatigue state of the manager.

[0176] The technical solution of the present application introduces the identity information of the manager, the current work shift information, and the historical fatigue accumulation data, and realizes the individualization and dynamic adjustment of the fatigue weight. Specifically, a basic fatigue weight is first obtained from the preset initial configuration according to the identity and current shift of the manager, which provides a preliminary distinction for different personnel and different shifts. Subsequently, a fatigue accumulation correction factor is calculated by analyzing the accumulated data such as the historical alarm response time, the historical false alarm processing time, and the historical alarm verification result of the manager. The correction factor can quantify the fatigue degree accumulated by the manager due to long-term work or work in a specific shift. Finally, the basic fatigue weight is combined with the fatigue accumulation correction factor to dynamically adjust the preset fatigue weight used to calculate the fatigue influence coefficient. Thus, the fatigue weight is no longer static and unchangeable, but can be adaptively adjusted according to the individual differences and real-time working state of the manager, so that the fatigue influence coefficient calculated subsequently can more accurately reflect the real fatigue state of the manager, effectively solving the problem of insufficient precision in setting the fatigue weight in the traditional technical solution.

[0177] By the technical solution, the fatigue weight can be dynamically adjusted according to the individual characteristics and historical performance of the manager, so that the calculation of the fatigue influence coefficient is more accurate. This not only improves the accuracy of the evaluation of the fatigue state of the manager, but also makes the adjustment of the alertness improvement coefficient more reasonable, so that when the visual information availability is reduced, non-visual information sources can be more effectively used for abnormal behavior judgment, significantly improving the overall accuracy and robustness of abnormal behavior pattern recognition, and reducing the false negative or false positive situation caused by the fatigue of the manager.

[0178] For example, assume that a monitoring center has a manager A and a manager B. The historical alarm response time of manager A during the night shift period is generally longer, and the false alarm processing time is also longer. The historical verification results show that the false alarm rate of manager A during the night shift period is relatively high. The historical data of manager B during the night shift period is relatively good. When manager A works during the night shift, a basic fatigue weight is first obtained from the initial weight configuration according to the identity information and the night shift information of manager A. Then, the historical fatigue accumulation data of manager A is analyzed to calculate a higher fatigue accumulation correction factor to reflect the higher fatigue degree of manager A during the night shift. The correction factor will be used to adjust the preset fatigue weight, so that the fatigue weight is given a higher value in the night shift situation of manager A. On the contrary, when manager B works during the night shift, the historical data of manager B may result in a lower fatigue accumulation correction factor, so that the adjusted fatigue weight is relatively low. In this way, the fatigue weight can be dynamically adjusted according to the individual differences and historical performance of different managers, so that the actual fatigue state of the manager can be more accurately reflected when calculating the fatigue influence coefficient, and the alertness improvement coefficient can be more reasonably adjusted to optimize the identification effect of abnormal behavior.

[0179] The embodiment of the present application also provides an abnormal behavior pattern recognition system based on multiple modalities, as shown in Figure 2 The abnormal behavior pattern recognition system 100 based on multiple modalities includes:

[0180] A modal health score generation module 10 is configured to acquire data of multiple information sources in real time, and generate a modal health score of each information source representing the data reliability of the information source, wherein the information sources include a visual information source and at least one non-visual information source.

[0181] A correction module 20 is configured to correct the modal health score of the visual information source based on a preset correction rule according to the acquired environmental light information, to obtain a target modal health score of the visual information source, wherein the preset correction rule includes a mapping relationship between the visual information quality and the light condition.

[0182] a weight determination module 30 configured to determine the weight of the visual information source and the non-visual information source in the abnormal behavior judgment according to the target modality health score of the visual information source, the modality health score of the non-visual information source, and a preset weight distribution strategy, wherein the weight distribution strategy satisfies that the weight value is positively correlated with the modality health score of the corresponding information source, and the sum of all weights is a fixed value;

[0183] an abnormal behavior recognition module 40 configured to recognize the abnormal behavior by fusion analysis according to the monitoring data of each information source and the weight of each information source.

[0184] The abnormal behavior pattern recognition system based on multi-modal proposed in the present application aims to effectively solve the problem of insufficient recognition accuracy and robustness of traditional existing abnormal behavior recognition technology in complex and variable environments, especially when the quality of visual information is limited, through modular design and intelligent data processing flow. The system can real-time evaluate and dynamically adjust the reliability of different information sources by introducing core functional modules such as modality health score generation, environment light correction, adaptive weight distribution, and multi-modal fusion analysis, so as to accurately recognize high-risk preparation behaviors that rely on subtle actions and have high concealment when facing dynamic decline of visual information quality caused by changes in environmental light.

[0185] The specific method and working principle of real-time acquisition of data of multiple information sources, generation of modality health score representing data reliability of each information source, correction of the modality health score of the visual information source based on the acquired environment light information and the preset correction rule to obtain the target modality health score of the visual information source, determination of the weight of the visual information source and the non-visual information source in the abnormal behavior judgment according to the target modality health score of the visual information source and the modality health score of the non-visual information source and the preset weight distribution strategy, and fusion analysis according to the monitoring data of each information source and the weight of each information source to recognize the abnormal behavior have been described in the above embodiments, and will not be repeated here. It should be emphasized that the above method steps are realized by system modularization in the present application.

[0186] The multi-modal based abnormal behavior pattern recognition system of the present application, through its modular design, can clearly divide the functional responsibilities, so that the system can recognize and utilize the ambiguous visual signals that are originally considered as "noise" or "low confidence" when the visual information quality is dynamically reduced due to the change of environmental light, and effectively combine them with reliable but insufficient information granularity auxiliary modalities (such as dwell time). Compared with the single modal or fixed weight fusion system in the prior art, the system of the present application can adaptively adjust the information analysis mode, no longer simply discard or low-weight process low-confidence visual signals, but through the modal health score and dynamic weight distribution mechanism, realize deeper relevance analysis and meaning reconstruction. Thus, the system of the present application significantly improves the recognition accuracy and robustness of high-risk preparation behaviors that rely on subtle actions and have high concealment under the condition of severely limited visual information, and provides more reliable abnormal behavior warning capability for high-security monitoring environment.

[0187] The above only describes the embodiments of the present application and is not used to limit the protection scope of the present application. For those skilled in the art, the present application can have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.

Claims

1. A multi-modal based anomalous behavior pattern recognition method, characterized in that, The method comprises: real-time acquisition of data of multiple information sources, generation of a modal health score representing data reliability of each information source, the information sources including visual information sources and at least one non-visual information source; wherein the modal health score is a numerical value between 0 and 1, and for the visual information source, at least one of image clarity, noise level and occlusion rate is quantified; for the non-visual information source, at least one of data integrity, sensor accuracy and environmental interference degree is quantified; based on a preset correction rule, the modal health score of the visual information source is corrected based on the acquired ambient light information to obtain a target modal health score of the visual information source, the preset correction rule including a mapping relationship between visual information quality and light conditions; based on a preset weight allocation strategy, the weight of the visual information source and the non-visual information source in abnormal behavior judgment is determined according to the target modal health score of the visual information source and the modal health score of the non-visual information source, wherein the weight allocation strategy satisfies that the weight value is positively correlated with the modal health score of the corresponding information source, and the sum of all weights is a fixed value; fusing analysis is performed according to the monitoring data of each information source and the weight of each information source to identify abnormal behavior.

2. The multi-modal based abnormal behavior pattern recognition method of claim 1, wherein, The method further comprises, after the step of determining the weight of the visual information source and the non-visual information source in abnormal behavior judgment based on the target modal health score of the visual information source and the modal health score of the non-visual information source according to the preset weight allocation strategy: determining a judgment threshold of an auxiliary information source in the non-visual information source according to the degree of reduction of visual information availability reflected by the target modal health score of the visual information source. 3.The multi-modal based abnormal behavior pattern recognition method of claim 1, wherein, The step of fusing analysis according to the monitoring data of each information source and the weight of each information source to identify abnormal behavior comprises: extracting potential activity clues from the signals of the visual information source; when the potential activity clues and the auxiliary information source in the non-visual information source with high weight appear at the same time, identifying abnormal behavior. 4.The multi-modal based abnormal behavior pattern recognition method of claim 2, wherein, The step of determining a judgment threshold of an auxiliary information source in the non-visual information source according to the degree of reduction of visual information availability reflected by the target modal health score of the visual information source comprises: acquiring historical activity duration distribution data of a sensitive area where a person is located, and calculating an activity duration expectation value of the sensitive area according to the historical activity duration distribution data, wherein the sensitive area is a specific spatial range defined based on physical characteristics, functional attributes or historical experience, and has a higher risk of abnormal behavior or needs special attention in a specific monitoring environment; calculating an alertness improvement coefficient according to the degree of reduction of visual information availability reflected by the target modal health score of the visual information source; generating a basic warning duration according to the activity duration expectation value of the sensitive area and the alertness improvement coefficient; adjusting a false alarm suppression factor according to the historical false alarm times of the sensitive area; determining a judgment threshold of the auxiliary information source according to the base alert duration and the false alarm suppression factor.

5. The multi-modal based anomalous behavior pattern recognition method of claim 4, wherein, The step of adjusting the false alarm suppression factor according to the historical false alarm times of the sensitive area comprises: obtaining a severity level of the alarm; evaluating a processing cost of the alarm; recording a response behavior of the manager to the alarm; evaluating a fatigue state of the manager according to the severity level of the alarm, the processing cost of the alarm and the response behavior, and adjusting the false alarm suppression factor according to the fatigue state of the manager.

6. The multi-modal based abnormal behavior pattern recognition method of claim 4, wherein, The step of calculating the alertness improvement coefficient according to the reduced degree of visual information availability reflected by the target modality health score of the visual information source comprises: obtaining a verification result of the alarm; recording a response action of the manager to the alarm; obtaining a reduced degree of visual information availability at the time when the alarm is triggered; calculating a fatigue influence coefficient according to the verification result of the alarm and the response action of the manager to the alarm; adjusting the calculation logic of the alertness improvement coefficient according to the fatigue influence coefficient and the reduced degree of visual information availability at the time when the alarm is triggered, to obtain the alertness improvement coefficient.

7. The multi-modal based abnormal behavior pattern recognition method of claim 4, wherein, The step of calculating the alertness improvement coefficient according to the reduced degree of visual information availability reflected by the target modality health score of the visual information source comprises: performing a preliminary judgment of the behavior intention according to the target modality health score of the visual information source and the modality health score of the non-visual information source, to obtain a preliminary judgment result of the behavior intention; obtaining a visual information quality dependence degree corresponding to the preliminary judgment result of the behavior intention; calculating the alertness improvement coefficient according to the reduced degree of visual information availability and the visual information quality dependence degree.

8. The multi-modal based abnormal behavior pattern recognition method of claim 6, wherein, The step of calculating the fatigue influence coefficient according to the verification result of the alarm and the response action of the manager to the alarm comprises: classifying the response action according to the verification result of the alarm, to obtain a classification result; calculating the fatigue influence coefficient according to the classification result and a preset fatigue weight.

9. The multi-modal based abnormal behavior pattern recognition method of claim 8, wherein, The step of calculating the fatigue influence coefficient according to the classification result and a preset fatigue weight comprises: obtaining identity information, current work shift information and historical fatigue accumulation data of the manager, wherein the historical fatigue accumulation data comprises historical alarm response times, historical false alarm processing durations and historical alarm verification results of the manager in different work shifts; obtaining a base fatigue weight corresponding to the manager and the work shift from a preset initial weight configuration according to the identity information of the manager and the current work shift information; calculating a fatigue accumulation correction factor of the manager in the current work shift according to the historical fatigue accumulation data of the manager; adjusting the preset fatigue weight according to the base fatigue weight and the fatigue accumulation correction factor; calculating the fatigue influence coefficient according to the classification result and the adjusted preset fatigue weight.

10. A multi-modal based anomalous behavior pattern recognition system, characterized in that, The system comprises: The modal health score generation module is configured to acquire data of multiple information sources in real time, and generate a modal health score of each information source representing the reliability of the data of the information source, the information sources including a visual information source and at least one non-visual information source; The correction module is configured to correct the modal health score of the visual information source based on a preset correction rule according to the acquired ambient light information, to obtain a target modal health score of the visual information source, and the preset correction rule includes a mapping relationship between visual information quality and light conditions; The weight determination module is configured to determine the weights of the visual information source and the non-visual information source in the abnormal behavior judgment based on a preset weight allocation strategy according to the target modal health score of the visual information source and the modal health score of the non-visual information source, wherein the weight allocation strategy satisfies that the weight value is positively correlated with the modal health score of the corresponding information source, and the sum of all weights is a fixed value; The abnormal behavior recognition module is configured to perform fusion analysis according to the monitoring data of each information source and the weights of each information source, and recognize the abnormal behavior.

Citation Information

Patent Citations

  • Personnel abnormal behavior detection method and system based on visual language large model

    CN119992641A

  • Emotion recognition method and system based on visual and auditory collaboration

    CN120852890A