Abnormal behavior pattern recognition method and system based on multiple modes
By acquiring data from multiple information sources in real time and generating modal health scores, correcting the reliability of visual information based on ambient lighting information, and dynamically adjusting weights, the problem of abnormal behavior recognition when visual information quality deteriorates in complex and ever-changing scenarios is solved, achieving higher recognition accuracy and robustness.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-02-05
- Publication Date
- 2026-03-10
AI Technical Summary
Existing technologies struggle to effectively fuse multi-source information and accurately identify covert abnormal behaviors in complex and variable scenarios, especially when changes in ambient lighting conditions lead to a decline in the quality of core visual information.
By acquiring data from multiple information sources in real time, a modal health score is generated. The modal health score of visual information is corrected based on ambient lighting information. The weight of each information source in the judgment of abnormal behavior is dynamically adjusted based on a preset weight allocation strategy to achieve adaptive information fusion.
When the quality of visual information deteriorates, it can effectively combine low-confidence visual signals with reliable but insufficiently granular auxiliary modalities, thereby improving the accuracy and robustness of abnormal behavior recognition, especially under conditions where visual information is severely limited, it can identify high-risk preparatory behaviors.
Smart Images

Figure CN121637367A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of abnormal behavior pattern recognition, and more specifically, to a method and system for abnormal behavior pattern recognition based on multimodality. Background Technology
[0002] In environments requiring high-security monitoring, such as prisons, continuous observation of personnel behavior is crucial for maintaining order and ensuring safety. Current technologies for detecting anomalous behavior, especially those relying on a single source of information (e.g., simply capturing motion or identifying clothing), struggle to meet the demands for multifaceted behavioral analysis in complex and ever-changing scenarios. For example, when the person being observed is far from the camera or obstructed by something, a system relying on only one source of information can easily miss key features, leading to incorrect identification of anomalous behavior. Furthermore, information from different sources (such as skeletal movements, clothing characteristics, and duration of stay in a particular location) often varies in form, making it difficult to effectively combine them and thus hindering accurate early detection of anomalous behavior.
[0003] In prison environments, when lighting conditions dynamically change due to environmental factors (such as strong glare or low illumination), leading to a continuous decline and unreliability in the quality of core visual information sources (such as skeletal key points and clothing features), the information fusion mechanisms of existing systems cannot effectively adapt to such complex scenarios with partial modal degradation. Specifically, the system fails to establish a mechanism that, upon detecting a continuous degradation of key visual modalities due to environmental factors, can automatically adjust its information analysis mode, moving beyond simply discarding or under-weighting these low-confidence visual signals. Instead, it can perform a deeper correlation analysis and meaning reconstruction with reliable but insufficiently granular auxiliary modalities (such as dwell time). The core problem facing current technology is: how to design an adaptive information fusion strategy so that when the system faces the dynamic decline in visual information quality caused by changes in ambient light, it can identify and utilize these blurred visual signals that were originally regarded as "noise" or "low confidence" and effectively combine them with reliable but insufficient information granular auxiliary modalities (such as dwell time), thereby accurately identifying high-risk preparatory behaviors that rely on subtle movements and are highly concealed under conditions of severely limited visual information. Summary of the Invention
[0004] This application provides a method and system for identifying abnormal behavior patterns based on multimodality, aiming to solve the technical problem that existing technologies struggle to effectively fuse multi-source information and accurately identify concealed abnormal behaviors in complex and variable scenarios, especially when changes in ambient lighting conditions lead to a decline in the quality of core visual information.
[0005] On the one hand, this application provides a method for identifying abnormal behavior patterns based on multimodality, including:
[0006] Data from multiple information sources is acquired in real time, and a modal health score is generated for each information source to characterize its data reliability. The information sources include visual information sources and at least one non-visual information source.
[0007] Based on the acquired ambient lighting information, the modal health score of the visual information source is corrected according to a preset correction rule to obtain the target modal health score of the visual information source. The preset correction rule includes the mapping relationship between visual information quality and lighting conditions.
[0008] Based on the target modal health score of the visual information source and the modal health score of the non-visual information source, the weights of the visual information source and the non-visual information source in the abnormal behavior judgment are determined according to a preset weight allocation strategy. The weight allocation strategy satisfies that the weight value is positively correlated with the modal health score of the corresponding information source, and the sum of all weights is a fixed value.
[0009] By integrating and analyzing monitoring data from various information sources and their respective weights, abnormal behaviors can be identified.
[0010] On the other hand, this application provides a sealed adaptive soft-pack lithium battery packaging image inspection system, the system comprising:
[0011] The modal health score generation module is used to acquire data from multiple information sources in real time and generate a modal health score for each information source that represents the reliability of its data. The information sources include visual information sources and at least one non-visual information source.
[0012] The correction module is used to correct the modal health score of the visual information source based on the acquired ambient lighting information and a preset correction rule to obtain the target modal health score of the visual information source. The preset correction rule includes the mapping relationship between visual information quality and lighting conditions.
[0013] The weight determination module is used to determine the weights of the visual information source and the non-visual information source in the abnormal behavior judgment based on the target modal health score of the visual information source and the modal health score of the non-visual information source, according to a preset weight allocation strategy. The weight allocation strategy satisfies that the weight value is positively correlated with the modal health score of the corresponding information source, and the sum of all weights is a fixed value.
[0014] The abnormal behavior identification module is used to identify abnormal behavior by performing fusion analysis based on monitoring data from various information sources and the weights of each information source.
[0015] This application discloses a multimodal abnormal behavior pattern recognition method and system. By acquiring data from multiple information sources in real time and generating modal health scores, it can objectively assess the reliability of each modality's data. Especially when changes in ambient lighting degrade visual information quality, this application can correct the modal health scores of visual information sources based on ambient lighting information to obtain a target modal health score, thus accurately reflecting the true usability of visual information. Based on this, this application dynamically determines the weight of each information source in abnormal behavior judgment according to a preset weight allocation strategy, using the corrected target modal health score of the visual information source and the modal health scores of non-visual information sources. This adaptive weight allocation mechanism allows the system to avoid simply discarding or under-weighting low-confidence visual signals in complex scenarios with partial modal degradation. Instead, it enables deeper correlation analysis and meaning reconstruction with reliable but insufficiently granular auxiliary modalities. Finally, by fusing and analyzing the monitoring data and weights of each information source, this application can effectively identify abnormal behavior.
[0016] Through the above technical solution, this application overcomes the limitations of existing technologies that rely on single information sources or fixed-weight fusion strategies. It effectively addresses the problem that in high-security monitoring environments such as prisons, when dynamic changes in lighting conditions lead to a continuous decline and unreliability in the quality of core visual information, existing systems cannot effectively adapt to complex scenarios with partial modal degradation, making it difficult to accurately identify high-risk preparatory behaviors that rely on subtle movements and are highly concealed. This application, by dynamically adjusting the information fusion strategy, enables the system to identify and utilize previously considered "noise" or "low-confidence" blurred visual signals, effectively combining them with reliable but insufficiently granular auxiliary modalities. This significantly improves the accuracy and robustness of abnormal behavior identification, especially under conditions of severely limited visual information, enabling more effective detection of high-risk preparatory behaviors, thereby enhancing the overall performance and security of the monitoring system. Attached Figure Description
[0017] To illustrate this application more clearly, the accompanying drawings used in the embodiments will be briefly described below. Obviously, those skilled in the art can obtain other drawings based on these drawings without any creative effort.
[0018] Figure 1 The diagram above illustrates a flowchart of a multimodal abnormal behavior pattern recognition method.
[0019] Figure 2 The diagram above illustrates a schematic of a multimodal abnormal behavior pattern recognition system.
[0020] Figure reference numerals: 100, Multimodal Abnormal Behavior Pattern Recognition System; 10, Modal Health Score Generation Module; 20, Correction Module; 30, Weight Determination Module; 40, Abnormal Behavior Recognition Module. Detailed Implementation
[0021] The technical solutions of this application will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this application, and not all embodiments. The components of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.
[0022] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, in the description of this application, terms such as "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0023] Traditional abnormal behavior recognition technologies, especially those relying on a single information source (e.g., capturing motion or identifying clothing), struggle to meet the demands for multifaceted behavioral analysis in complex and ever-changing surveillance scenarios. For instance, when the observed individual is far from the camera or obstructed, a single-source system may miss crucial features, leading to incorrect abnormal behavior identification. Furthermore, information from different sources (such as skeletal movements, clothing characteristics, and duration of contact) varies in form, making effective integration difficult and hindering accurate early detection of abnormal behavior. In prison environments, dynamic changes in lighting conditions due to environmental factors (such as strong glare or low illumination) cause a continuous decline in the quality and reliability of core visual information sources (such as skeletal landmarks and clothing features), rendering existing systems' information fusion mechanisms ineffective in adapting to such complex scenarios with partial modal degradation. Specifically, the system fails to establish a mechanism that, when it detects a continuous degradation of key visual modalities due to environmental factors, can automatically adjust its information analysis mode. Instead of simply discarding or under-weighting these low-confidence visual signals, it should be able to perform deeper correlation analysis and meaning reconstruction with reliable but insufficiently granular auxiliary modalities (such as dwell time). The core problem facing current technology is: how to design an adaptive information fusion strategy that enables the system to identify and utilize these previously considered "noise" or "low-confidence" blurred visual signals when faced with dynamic degradation of visual information quality caused by changes in ambient lighting. This strategy should effectively combine these signals with reliable but insufficiently granular auxiliary modalities (such as dwell time) to accurately identify high-risk preparatory behaviors that rely on subtle movements and are highly concealed under conditions of severely limited visual information.
[0024] like Figure 1 The diagram illustrates a flowchart of a multimodal-based abnormal behavior pattern recognition method. This application proposes a multimodal-based abnormal behavior pattern recognition method, comprising:
[0025] S10, acquire data from multiple information sources in real time, and generate a modal health score for each information source that characterizes the reliability of its data. The information sources include visual information sources and at least one non-visual information source.
[0026] Modal health score refers to a quantitative assessment of the reliability or quality of data from each information source. In other words, for each information source, the modal health score is a quantitative score characterizing its data reliability. This score aims to assess the trustworthiness and quality level of the data provided by that information source at the current moment or under specific conditions. A higher score indicates more reliable data and is more suitable for judging abnormal behavior; conversely, a lower score indicates poorer data reliability and lower reference value in judgment.
[0027] For example, visual information sources typically refer to image or video data acquired through cameras, infrared sensors, etc., such as skeletal key points and clothing features. Non-visual information sources include, but are not limited to, audio information, biometric information (such as heart rate and body temperature), and environmental sensor data (such as dwell time and location information). These information sources collectively form the foundation of multimodal data.
[0028] Specifically, the calculation basis, quantification method, or evaluation criteria for modal health scores will be comprehensively considered by those skilled in the art based on the characteristics of different information sources:
[0029] First intrinsic quality assessment:
[0030] For visual information sources (e.g., video streams, images), modal health scores can be quantified based on intrinsic quality metrics of the image or video. For example, image sharpness (e.g., through edge sharpness, blur detection), noise levels (e.g., signal-to-noise ratio (SNR), peak signal-to-noise ratio (PSNR), contrast, brightness uniformity, and color distortion can be assessed. These metrics can be calculated in real-time using image processing algorithms. For instance, a high-definition, low-noise video stream will receive a higher modal health score, while a blurry, snow-filled video stream will receive a lower score.
[0031] For non-visual information sources (which may include at least one of the following: dwell time, audio, and biometric data), the modal health score can be quantified based on the completeness, accuracy, real-time performance of the data, and the performance indicators of the sensor itself. For example, for dwell time data, the signal strength of the data acquisition device (such as an RFID reader or UWB locator), the packet loss rate of data transmission, and the positioning accuracy can be evaluated. For audio information, the background noise level and speech clarity can be evaluated. The more complete, accurate, and real-time the data, and the better the sensor's operating condition, the higher the modal health score.
[0032] Second, environmental factors impact assessment: For visual information sources, environmental factors significantly affect data reliability. This means that, in addition to the intrinsic quality of the data itself, the external environment (such as lighting conditions) is also an important basis for calculating or correcting modal health scores. In environments with strong light glare or low illumination, even if the camera itself performs well, the quality of its output visual information will drop sharply, causing its modal health score to be corrected to a lower value. Furthermore, occlusion conditions (such as people being obstructed by obstacles) and weather conditions (such as rain, snow, and fog) can also be environmental factors affecting the reliability of visual information. For non-visual information sources, some non-visual information sources may also be affected by environmental factors. For example, the modal health score of a sound sensor may decrease in a noisy environment; the positioning accuracy of a wireless positioning system decreases in an environment with a large number of metallic reflective objects, and its modal health score may also decrease accordingly.
[0033] In terms of quantification methods and evaluation criteria, modal health scores can typically be quantified as a value between 0 and 1, where 1 indicates that the data is completely reliable or of excellent quality, and 0 indicates that the data is completely unusable or of extremely poor quality. Specific quantification methods may include at least one of the following:
[0034] Threshold-based rule system: A set of predefined rules and thresholds. For example, if the video signal-to-noise ratio is lower than a certain threshold, the modal health score is reduced by one level; if the light intensity is lower than a certain threshold, the score is reduced by another level.
[0035] Machine learning model-based approach: Train a model, input various data quality metrics and environmental parameters, and output a modality health score. For example, you can collect a large amount of visual data of varying quality, manually label its reliability, and then train a classifier or regression model to predict the modality health score.
[0036] Weighted summation: Multiple influencing factors (such as sharpness, noise, light intensity, and occlusion rate) are quantified separately and assigned different weights, and then weighted summation is performed to obtain the final modal health score.
[0037] S20, based on the acquired ambient lighting information, the modal health score of the visual information source is corrected according to a preset correction rule to obtain the target modal health score of the visual information source. The preset correction rule includes the mapping relationship between visual information quality and lighting conditions.
[0038] S30, based on the target modal health score of the visual information source and the modal health score of the non-visual information source, and based on a preset weight allocation strategy, determine the weights of the visual information source and the non-visual information source in the abnormal behavior judgment, wherein the weight allocation strategy satisfies that the weight value is positively correlated with the modal health score of the corresponding information source, and the sum of all weights is a fixed value.
[0039] S40 identifies abnormal behavior by integrating and analyzing monitoring data from various information sources and their respective weights.
[0040] This application effectively addresses the limitations of existing technologies in identifying abnormal behavior in complex and variable environments by introducing modal health scores, ambient lighting correction, adaptive weight allocation, and multimodal fusion analysis. In particular, it can more accurately and robustly identify high-risk preparatory behaviors when visual information quality is limited.
[0041] The multimodal abnormal behavior pattern recognition method proposed in this application aims to improve the accuracy and robustness of abnormal behavior recognition by integrating and intelligently analyzing data from different information sources.
[0042] First, it is necessary to acquire data from multiple information sources in real time. For example, visual information data, such as video streams or image sequences, can be acquired through cameras deployed in the monitoring area. Simultaneously, audio information can be acquired through microphones, or non-visual information data, such as personnel location and dwell time, can be obtained through RFID tags and UWB positioning. While acquiring this data, a modal health score is generated for each information source. For example, for visual information sources, the modal health score can be calculated by analyzing factors such as image sharpness, noise level, and occlusion. For non-visual information sources, such as dwell time data, a modal health score can be generated by assessing the stability of the data acquisition equipment and the integrity of data transmission.
[0043] Secondly, based on the acquired ambient lighting information, the modal health score of the visual information source is corrected according to a preset correction rule to obtain the target modal health score of the visual information source. Ambient lighting information can be acquired by lighting sensors deployed in the monitoring area, or estimated by analyzing the visual information source itself (e.g., the brightness histogram of the image).
[0044] In some embodiments, the preset correction rule can be a function or lookup table that adjusts the visual modality health score based on the current lighting conditions. This process adjusts the modality health score of the visual information source by incorporating real-time ambient lighting information. Its core purpose is to more accurately reflect the actual reliability of visual information under current environmental conditions, as changes in lighting conditions directly affect the quality of visual data. For example, under low light or strong glare conditions, the reliability of the visual information source decreases, so the correction rule will adjust the original visual modality health score downwards, resulting in a lower target modality health score. Conversely, under ideal lighting conditions, the correction rule may not significantly adjust the score.
[0045] Specifically, the process of correcting the visual score is as follows:
[0046] First, the system acquires ambient lighting information in real time. This can be achieved in several ways, such as by deploying dedicated light sensors (e.g., photoresistors, illuminometers) to measure the lux value of the environment; or, the system can indirectly estimate the current lighting conditions by analyzing the visual information source itself (e.g., extracting the brightness histogram, average pixel intensity, or proportion of overexposed / underexposed areas from video frames).
[0047] Secondly, preset correction rules are applied. Once ambient lighting information is acquired, the system invokes these preset correction rules. These rules typically exist in the form of mathematical functions, lookup tables, or conditional logic, and they define how to adjust the modal health score of the original visual information source under different lighting conditions.
[0048] The design of the preset correction rules is based on an understanding of the relationship between visual information quality and lighting conditions. For example, in extremely low light or strong glare (overexposure) environments, the quality of visual information (such as skeletal keypoint extraction and clothing feature recognition) will significantly decrease, reducing its reliability. Therefore, the correction rules will instruct the system to reduce the original modality health score under these unfavorable lighting conditions. Under ideal or moderate lighting conditions, the quality of visual information is usually high, and the correction rules may only make minor adjustments or no adjustments at all.
[0049] The correction method can be linear, non-linear, or piecewise. For example, when the illumination value is below a certain threshold, the modal health score decreases proportionally for every unit reduction in illumination; when the illumination value is above a certain threshold and causes overexposure, the score also decreases accordingly.
[0050] Finally, after processing according to preset correction rules, the original visual information source modal health score was adjusted to the target modal health score of the visual information source. This target score more realistically reflects the usability and reliability of visual information in the current environment, providing a more accurate basis for subsequent weight allocation and abnormal behavior judgment.
[0051] For example, suppose the initial modal health score of the visual information source is 0.8 (out of 1.0), which means that its data reliability is high under standard conditions.
[0052] Scene 1: Low-light environment
[0053] The system detected an ambient light intensity of 50 Lux using a light sensor, far below the 200 Lux required for normal operation. A preset correction rule states: "When the light intensity is below 200 Lux, the visual modality health score decreases by 0.1 for every 50 Lux decrease." According to this rule, the light intensity decreased from 200 Lux to 50 Lux, a total reduction of 150 Lux (three intervals of 50 Lux). Correction calculation: Original score 0.8 - (30.1) = 0.5. Ultimately, the target modality health score of the visual information source was corrected to 0.5, reflecting a significant decrease in the reliability of visual information under low light conditions.
[0054] Scenario 2: Strong light glare environment
[0055] The system analyzes visual images and detects large overexposed areas (e.g., exceeding 30%), typically caused by strong direct sunlight or reflections. A preset correction rule specifies: "When the overexposed area exceeds 20%, the visual modality health score is reduced by 0.2; if it exceeds 40%, it is reduced by 0.4." According to this rule, an overexposed area of 30% falls between 20% and 40%, so the system may use interpolation or directly apply a reduction of 0.2. Correction calculation: Original score 0.8 - 0.2 = 0.6. Ultimately, the target modality health score of the visual information source is corrected to 0.6, indicating that the reliability of visual information (such as detailed features) is reduced under strong light glare.
[0056] Scenario 3: Ideal Lighting Environment
[0057] The system detected an ambient light intensity of 300 Lux, and image quality analysis showed no significant overexposure or underexposure. A preset correction rule stipulates: "When the light intensity is between 200-500 Lux and the image quality is good, no correction or fine-tuning is performed." Correction calculation: Original score 0.8 - 0 = 0.8. Ultimately, the target modality health score of the visual information source remains at 0.8, reflecting the high reliability of visual information under ideal lighting conditions.
[0058] Through this dynamic adjustment based on preset correction rules, the system can more accurately assess the actual usability of visual information, thereby more reasonably allocating the weights of different modal information in subsequent abnormal behavior judgments, and improving the overall recognition accuracy and robustness.
[0059] After obtaining the target modal health score, the weights of visual and non-visual information sources in abnormal behavior judgment are determined based on a pre-defined weighting strategy, according to the target modal health scores of visual information sources and non-visual information sources. This weighting strategy can be a machine learning model. For example, when the target modal health score of a visual information source is low, its weight in abnormal behavior judgment is reduced, while the weight of non-visual information sources is increased accordingly. Conversely, when the target modal health score of a visual information source is high, its weight increases. This dynamic weighting mechanism allows for adaptive adjustment of the dependence on different modalities to cope with fluctuations in information quality caused by environmental changes.
[0060] Finally, the monitoring data from each information source and their respective weights are fused for analysis to identify abnormal behavior. Fusion analysis can employ various techniques, such as feature-level fusion, decision-level fusion, or hybrid fusion. In feature-level fusion, raw data or extracted features from different information sources are integrated into a unified feature vector, which is then input into an abnormal behavior recognition model (such as support vector machines or neural networks) for classification. In decision-level fusion, each information source independently judges the abnormal behavior, and these judgments are then weighted and voted on or averaged according to their respective weights to arrive at a comprehensive abnormal behavior recognition result. For example, when the weight of visual information sources is higher, their influence on the final abnormal behavior judgment is greater; when the weight of non-visual information sources is higher, their influence on the final abnormal behavior judgment is greater. Through this fusion analysis, the advantages of multimodal information can be comprehensively utilized to improve the accuracy and robustness of abnormal behavior recognition.
[0061] In some embodiments, a preset weighting strategy can dynamically adjust the importance of different information sources (including visual information sources and at least one non-visual information source) in abnormal behavior judgment based on their data reliability, i.e., modal health scores. This strategy is predefined, can be determined during the system design phase, and can be adjusted based on real-time data during operation. This strategy typically takes two main forms:
[0062] One approach is a rule-based weight allocation strategy. This strategy determines weight allocation by setting a series of explicit rules and thresholds. The system compares the target modal health score of the visual information source with the modal health scores of the non-visual information source against preset thresholds, and then executes the corresponding weight adjustment rules. In principle, the weight allocation strategy satisfies the following conditions: the weight value is positively correlated with the modal health score of the corresponding information source, and the sum of all weights is a fixed value.
[0063] The specific methods for determining weights include: 1) Defining modal health score ranges: Dividing modal health scores into different levels, such as high, medium, and low. 2) Setting weight adjustment rules: Pre-setting a weight allocation scheme for each combination of modal health score levels (e.g., high visual modal health and medium non-visual modal health). 3) Dynamic adjustment: When the target modal health score of the visual information source is low (e.g., due to insufficient ambient light leading to a decrease in visual information quality), the system will reduce the weight of the visual information source in abnormal behavior judgment according to the preset rules, and correspondingly increase the weight of non-visual information sources to compensate for the lack of visual information. Conversely, when the target modal health score of the visual information source is high, its weight will increase. Typically, the sum of the weights of all information sources will remain a constant (e.g., 1) to ensure that the overall judgment influence remains unchanged.
[0064] The second is a weight allocation strategy based on machine learning models: This strategy uses machine learning models (e.g., neural networks, support vector machines, or decision trees) to learn the relationship between modality health scores and optimal weight allocation in historical data.
[0065] The specific methods for determining weights include: 1) Data collection and labeling: Collecting a large amount of historical monitoring data, including modal health scores under different environmental conditions, actual abnormal behaviors, and corresponding accurate judgment results. This data is used to train the model. 2) Model training: Training a machine learning model so that it can output an optimal weight combination based on the target modal health score from the input visual information source and the modal health scores from non-visual information sources. The training objective of the model is to maximize the accuracy of abnormal behavior recognition while minimizing false positives and false negatives. 3) Real-time prediction: During system operation, the modal health scores acquired in real time are used as input to the model, which predicts and outputs the optimal weights for each information source in the current context.
[0066] For example, suppose in a prison environment, the system mainly relies on visual information sources (such as personnel posture and movement captured by cameras) and a non-visual information source (such as the time personnel stay recorded by area sensors).
[0067] Scenario 1: Daytime with plenty of sunlight
[0068] Target modal health score from visual information sources: High (e.g., 0.9) due to good lighting, clear camera images, and accurate skeletal keypoint extraction. Modal health score from non-visual information sources: High (e.g., 0.95) due to stable operation of the region sensors.
[0069] Rule-based strategy: Pre-defined rules may stipulate that when the visual modality health score is higher than a certain threshold (e.g., 0.7), the weight of visual information is 0.7, and the weight of non-visual information is 0.3. In this case, the system will rely more on visual information to judge abnormal behavior.
[0070] The machine learning-based strategy is as follows: the trained model receives (0.9, 0.95) as input, and the output weights may be 0.75 for visual information and 0.25 for non-visual information, indicating that visual information contributes more under good conditions.
[0071] Scenario 2: A dimly lit night
[0072] Target modal health score from visual information sources: low (e.g., 0.3) due to insufficient lighting, blurry camera images, and difficulty in extracting skeletal key points. Modal health score from non-visual information sources: high (e.g., 0.95) because the area sensor is unaffected by lighting conditions.
[0073] Rule-based strategy: Pre-defined rules may stipulate that when the visual modality health score falls below a certain threshold (e.g., 0.5), the weight of visual information decreases to 0.2, while the weight of non-visual information increases to 0.8. In this case, the system will rely more on non-visual information such as the duration of human stay to determine abnormal behavior, such as prolonged stay in sensitive areas.
[0074] The machine learning-based strategy is as follows: the model receives (0.3, 0.95) as input, and the output weights may be 0.15 for visual and 0.85 for non-visual, indicating that when vision is limited, non-visual information becomes the main basis for judgment.
[0075] Scenario 3: Camera malfunction or obstruction
[0076] Target modality health score for visual information sources: very low (e.g., 0.1), even close to 0. Modality health score for non-visual information sources: relatively high (e.g., 0.95).
[0077] Rule-based strategy: Pre-defined rules may stipulate that when the visual modality health score is below a very low threshold (e.g., 0.2), the weight of visual information is almost 0 (e.g., 0.05), and the weight of non-visual information is close to 1 (e.g., 0.95). The system relies almost entirely on non-visual information for judgment.
[0078] The machine learning-based strategy is as follows: the model receives (0.1, 0.95) as input, and the output weights may be 0.02 for visual and 0.98 for non-visual. The system will mainly rely on non-visual information to make anomaly judgments.
[0079] Through this pre-defined weighting strategy, the method can intelligently adjust the relative importance of each information source in the judgment of abnormal behavior based on the real-time changing modal health score. Thus, even in complex and ever-changing environments, especially when the quality of some modal information declines, it can still maintain high accuracy and robustness in the identification of abnormal behavior.
[0080] The overall working principle of this application lies in the ability to assess and dynamically adjust the reliability of different information sources in real time by introducing modal health scores and ambient lighting correction mechanisms. When the quality of visual information sources deteriorates due to changes in ambient lighting, their modal health scores are corrected. Consequently, in subsequent weighting processes, the weight of visual information sources is reduced, while the weight of non-visual information sources is increased. This adaptive weighting strategy ensures that even when some modal information is compromised, more judgment criteria can be transferred to more reliable modalities. Ultimately, by performing weighted fusion analysis on the monitoring data from various information sources, the advantages of multimodal information can be comprehensively utilized to effectively identify abnormal behavior. For example, in a prison environment, when insufficient light at night causes blurred visual information from cameras, the weight of visual information is reduced, and more reliance is placed on non-visual information such as the duration and location of individuals to determine the presence of abnormal behavior. This avoids false alarms or missed alarms caused by a decline in the quality of a single modal information.
[0081] Compared to existing technologies, the core innovation of this application lies in its adaptive information fusion strategy. Traditional methods often simply discard or under-weight low-confidence visual signals, leading to a significant decrease in recognition ability when visual information is limited. This application, by introducing modality health scores and ambient lighting correction, can more finely assess the availability of visual information. More importantly, by dynamically adjusting the weight of each modality in abnormal behavior judgment, this application can effectively combine ambiguous visual signals with reliable but insufficiently granular auxiliary modalities (such as dwell time), thereby accurately identifying high-risk preparatory behaviors that rely on subtle movements and are highly concealed, even under conditions of severely limited visual information. This adaptive and intelligent fusion mechanism significantly improves the accuracy and robustness of abnormal behavior recognition, especially in complex and ever-changing environments, where its advantages are even more pronounced.
[0082] In some embodiments, after the step of determining the weights of the visual information source and the non-visual information source in abnormal behavior judgment based on the target modal health score of the visual information source and the modal health score of the non-visual information source, according to a preset weight allocation strategy, the method further includes:
[0083] Based on the degree of reduction in visual information availability reflected by the target modality health score of the visual information source, the determination threshold for auxiliary information sources among the non-visual information sources is determined.
[0084] Specifically, the degree of reduction in visual information availability can be understood as the extent to which a visual information source's ability to provide effective information decreases under specific environments or conditions. This degree is typically directly reflected by the target modality health score of the visual information source; for example, the lower the target modality health score, the greater the degree of reduction in visual information availability. Auxiliary information sources among non-visual information sources refer to other modal data sources besides visual information sources that can provide additional information to assist in judging abnormal behavior, such as sound sensors, infrared sensors, and millimeter-wave radar. The judgment threshold is the critical value that auxiliary information sources use to determine whether abnormal behavior exists. When the monitoring data of the auxiliary information source exceeds or falls below this threshold, it is considered that abnormal behavior may exist. Determining this judgment threshold aims to make the judgment of auxiliary information sources more accurate and adaptive, thereby compensating for the impact of reduced visual information availability.
[0085] The technical solution of this application achieves dynamic adaptive adjustment of the abnormal behavior judgment logic by associating the degree of reduction in visual information availability reflected by the target modality health score of visual information sources with the judgment threshold of auxiliary information sources among non-visual information sources. When the availability of visual information decreases, the judgment threshold of auxiliary information sources can be adjusted accordingly based on the degree of reduction. For example, if visual information is severely impaired, the judgment threshold of some auxiliary information sources may be lowered, making them more sensitive to potential abnormal signals, thus enabling timely capture of abnormal behavior clues even when visual information is insufficient. Conversely, if the availability of visual information is high, the judgment threshold of auxiliary information sources can be appropriately increased to reduce false alarms. This dynamic adjustment mechanism enables more flexible and intelligent utilization of non-visual information to compensate for the deficiencies of visual information during multimodal information fusion analysis, ensuring the robustness of abnormal behavior recognition.
[0086] Through the above technical solution, this application effectively addresses the problem that relying solely on weight adjustments may lead to a decrease in the accuracy of abnormal behavior recognition when the availability of visual information is reduced. By dynamically determining the judgment threshold for auxiliary information sources based on the degree of reduction in visual information availability, the sensitivity of non-visual information sources can be controlled more precisely, avoiding missed or false alarms caused by missing or blurred visual information. This significantly improves the adaptability and reliability of abnormal behavior pattern recognition methods in complex and changing environments, enabling high recognition performance under different lighting and occlusion conditions, thereby improving the overall efficiency and accuracy of abnormal behavior detection.
[0087] For example, suppose in a surveillance scenario, at night or in low light, the target modal health score of visual information sources (such as cameras) decreases significantly, reflecting a high degree of reduction in visual information availability. In this case, the judgment thresholds of non-visual information sources (such as sound sensors or infrared sensors) are dynamically adjusted based on this reduction. Specifically, if visual information is almost unavailable, the judgment threshold for sound sensors detecting abnormal sounds (such as breaking glass or violent impacts) might be appropriately lowered, making it easier to trigger alarms, even if the sound intensity is slightly below the normal threshold. Simultaneously, the judgment threshold for infrared sensors detecting abnormal heat source movement might also be adjusted to improve their sensitivity to potential intruders. In this way, even when visual information is limited, abnormal behavior, such as nighttime intrusion or sabotage, can be identified promptly and accurately by enhancing reliance on and sensitivity to non-visual auxiliary information.
[0088] In some embodiments, the step of identifying abnormal behavior by performing fusion analysis based on monitoring data from each information source and the weights of each information source includes:
[0089] Extract potential activity cues from the signals of the visual information source;
[0090] Abnormal behavior is identified when the potential activity cues appear simultaneously with auxiliary information sources from non-visual information sources that have high weighting.
[0091] Specifically, through image processing, video analysis, and other technologies, specific visual patterns or events that may indicate abnormal behavior are identified and extracted from data acquired by visual devices such as cameras and infrared sensors. This allows for the extraction of potential activity cues from signals from visual information sources. For example, this may include detecting changes in posture, abnormal movement trajectories, the appearance or disappearance of specific objects, and intrusion into an area. These cues themselves may not be sufficient to definitively determine an anomaly, but they are preliminary indications of abnormal behavior. When these potential activity cues coexist with auxiliary information sources among high-weighted non-visual information sources, the abnormal behavior is identified. After extracting the visual potential activity cues, the non-visual information sources related to those cues are further examined. These non-visual information sources can include, for example, sound sensors, vibration sensors, access control systems, and environmental sensors. High-weighted non-visual information sources refer to non-visual modalities that are given high credibility or importance in the judgment of abnormal behavior. Auxiliary information sources specifically refer to non-visual modalities that can supplement or verify the visual cues. When visual cues are synchronized or strongly correlated in time or space with anomalous signals indicated by these high-weighted auxiliary non-visual information sources, they are identified as anomalous behavior. For example, when vision detects someone lingering in a sensitive area for an extended period (potential activity cue), and sound sensors detect unusual impact sounds or alarm sounds (high-weighted non-visual auxiliary information), anomalous behavior can be identified more definitively.
[0092] In some embodiments, a sensitive area refers to a specific spatial range within a particular monitoring environment that, due to its inherent physical characteristics, functional attributes, or historical experience, is considered to pose a higher risk of abnormal behavior or requires special attention. These areas are characterized by the fact that any unusual activity patterns or duration of stay by personnel within these areas may indicate potential dangers or abnormal events.
[0093] The identification of sensitive areas is primarily related to the specific application scenario and environmental layout. For example, in high-security monitoring environments such as warehouses, areas where valuables are stored, equipment operation areas, and entrances / exits may be defined as sensitive areas because these areas may involve unauthorized operations.
[0094] The identification of sensitive areas also requires an assessment of potential risks. By analyzing historical event data, such as locations where unusual behavior has occurred in the past and areas with a high incidence of false alarms, areas requiring close monitoring can be identified. For example, if an area has repeatedly experienced unusual events following prolonged periods of human presence, then that area should be marked as a sensitive area.
[0095] In actual deployment, sensitive areas are typically configured and demarcated based on pre-defined rules and the experience of domain experts. For example, security experts may manually or semi-automatically demarcate these areas in the monitoring system based on their understanding of the security needs of a specific location. These rules may include:
[0096] Physical boundary: The area defined by physical structures such as walls, doors, and windows.
[0097] Functional areas: Areas with specific purposes, such as rest areas, workbenches, and areas next to lockers.
[0098] Concealed areas: blind spots of cameras, poorly lit corners, etc.
[0099] The technical solution of this application effectively solves the problem of false positives or false negatives that may exist in traditional fusion analysis by refining the process of identifying abnormal behavior into two stages. First, potential activity cues are extracted from visual information sources, which allows for initial focusing on areas or events that may be abnormal, avoiding the computational burden of blindly fusing all data. Second, by requiring that these potential activity cues appear simultaneously with auxiliary information sources among non-visual information sources with high weights in order to identify abnormal behavior, this technical solution introduces a crucial verification mechanism. Because the weights of non-visual information sources have been optimized and allocated according to their modality health scores, high-weighted non-visual information sources are considered more reliable judgment criteria in the current environment. Therefore, when visual cues corroborate these reliable non-visual auxiliary information sources, the accuracy of abnormal behavior judgment will be significantly improved, thereby avoiding misjudgments that may result from relying solely on a single modality or simple fusion.
[0100] Through the above technical solutions, this application can significantly improve the accuracy and robustness of abnormal behavior pattern recognition. By introducing the extraction of potential activity cues, abnormal signs can be detected earlier, and the processing of irrelevant information is reduced. More importantly, by co-judging visual cues with high-weight non-visual auxiliary information, false alarms that may be caused by single-modal information are effectively avoided. Especially when visual information is limited or ambiguous, non-visual auxiliary information can provide strong corroboration, thereby ensuring the reliability of abnormal behavior recognition, reducing the false alarm rate, and improving the ability to capture real abnormal events.
[0101] For example, suppose we are in a warehouse environment and need to identify unusual loitering behavior by people. Visual information sources may include surveillance cameras installed in the warehouse, while non-visual information sources may include sound sensors (for detecting unusual noises, such as the sound of breaking glass or picking locks), vibration sensors (for detecting unusual vibrations of shelves or doors and windows), and RFID readers (for tracking the movement of items).
[0102] First, the system analyzes the video streams captured by surveillance cameras in real time to extract potential activity clues. For example, if a person is detected lingering in a non-work area for an extended period of time, or engaging in unusual rummaging near shelves, these will be identified as potential activity clues.
[0103] Simultaneously, the modal health score of the visual information source will be adjusted based on factors such as ambient lighting, and combined with the modal health scores of the sound sensor and vibration sensor to determine their weights in abnormal behavior judgment. It is assumed that in low light conditions at night, the weight of visual information decreases, while the weights of the sound sensor and vibration sensor relatively increase.
[0104] When the vision system detects potential rummaging in a sensitive area (such as near a shelf storing valuables), it immediately checks for high-weight non-visual information. If the sound sensor detects an unusual metallic scraping sound or the sound of items falling, or the vibration sensor detects severe vibration of the shelf, then the potential activity cue of "person rummaging" is associated with the high-weight non-visual auxiliary information of "abnormal sound / vibration." Because both occur simultaneously and corroborate each other, this is identified as abnormal behavior, triggering an alarm. This collaborative judgment mechanism significantly improves the accuracy of abnormal behavior identification, avoiding false alarms based solely on blurry visual images or a single abnormal sound.
[0105] In some embodiments, the step of determining the threshold for determining auxiliary information sources among the non-visual information sources based on the degree of reduction in visual information availability reflected by the target modality health score of the visual information source includes:
[0106] The system acquires historical activity duration distribution data of personnel in sensitive areas, and calculates the expected activity duration of the sensitive areas based on the historical activity duration distribution data. The sensitive areas are specific spatial ranges that are defined based on physical characteristics, functional attributes, or historical experience and that have a high risk of abnormal behavior or require special attention in a specific monitoring environment.
[0107] The alertness enhancement coefficient is calculated based on the degree of visual information availability reduction reflected by the target modality health score of the visual information source.
[0108] A basic warning duration is generated based on the expected activity duration of the sensitive area and the alertness enhancement coefficient.
[0109] Adjust the false alarm suppression factor based on the historical false alarm count of the sensitive area;
[0110] The threshold for determining the source of the auxiliary information is determined based on the basic warning duration and the false alarm suppression factor.
[0111] Specifically, the system collects and analyzes records of the time individuals spend or are active within a specific monitored area at different times, thereby obtaining historical activity duration distribution data for the sensitive areas in which these individuals are located. For example, this data can be obtained through historical surveillance video analysis, access control system records, or sensor data within the area. Based on this historical activity duration distribution data, the expected activity duration of the sensitive area is calculated. The purpose is to quantify the typical activity patterns of this sensitive area under normal circumstances, providing an objective benchmark for subsequent anomaly detection.
[0112] Specifically, based on the degree of decrease in visual information availability reflected by the target modality health score of the visual information source, an alertness enhancement coefficient is calculated. This can be understood as follows: when the quality or availability of visual information declines, the reliance on non-visual information sources and alertness should increase accordingly. For example, when insufficient ambient light, obstruction of visual sensors, or malfunctions significantly reduce the availability of visual information, the alertness enhancement coefficient will be calculated as a higher value to encourage greater attention to non-visual information and compensate for the lack of visual information.
[0113] In practical applications, a base alert duration is generated based on the expected activity duration of the sensitive area and the alertness enhancement coefficient. This base duration is used to initially set a baseline time length for judging abnormal behavior, taking into account the inherent activity characteristics of the area and the reliability of current visual information. For example, if an area typically has a short activity duration and the availability of current visual information is low, the generated base alert duration may be set shorter to trigger an alert more quickly.
[0114] Furthermore, based on the historical false alarm count of the sensitive area, a false alarm suppression factor is adjusted. This allows for dynamic adjustment of sensitivity by learning from historical false alarm patterns, thereby reducing unnecessary alarms. For example, if a certain area experiences frequent false alarms under specific conditions, the false alarm suppression factor will be adjusted to a larger value, thereby increasing the judgment threshold, reducing the false alarm rate, and improving the user experience.
[0115] Finally, based on the basic warning duration and the false alarm suppression factor, a threshold for determining the source of auxiliary information is determined. This threshold is a dynamic value derived by comprehensively considering regional activity characteristics, visual information availability, and historical false alarms. It is used to determine when to identify non-visual information sources as auxiliary information sources, thereby triggering further judgment of abnormal behavior.
[0116] The technical solution of this application addresses the inaccuracy issues that may arise from simple threshold settings by introducing multi-dimensional considerations. First, by acquiring historical activity duration distribution data for sensitive areas and calculating expected activity durations, a baseline understanding of "normal" behavioral patterns in these areas can be established, avoiding misjudging normal activities consistent with regional characteristics as abnormal. Second, by calculating an alertness enhancement coefficient, the focus on non-visual information sources can be dynamically adjusted based on real-time changes in visual information availability, ensuring that non-visual information can play a timely and effective role when visual information is limited. Third, by introducing a false alarm suppression factor and adjusting it based on historical false alarm counts, the solution can learn and adapt to the environment, effectively reducing false alarms caused by environmental noise or unforeseen events. Therefore, the above technical solution comprehensively considers the inherent characteristics of the environment, real-time perception capabilities, and historical experience, making the judgment threshold for auxiliary information sources more accurate and adaptive, thereby significantly improving the accuracy and reliability of abnormal behavior identification.
[0117] Through the above technical solution, this application can more accurately use non-visual information sources to judge abnormal behavior when visual information is limited, effectively reducing the false alarm rate and false negative rate, and improving the overall performance and user experience of abnormal behavior recognition.
[0118] For example, suppose there is a specific sensitive area in a warehouse, such as a valuables storage area, where it is necessary to identify abnormal behavior.
[0119] First, data on the distribution of activity duration in the valuables storage area over the past few months will be obtained. For example, by analyzing access control records and sensor data within the area, it is found that the average activity duration in this area during weekday daytime (9:00-17:00) is 15 minutes, and the average activity duration at night (17:00-9:00 the next day) is 2 minutes. From this, the expected activity duration in this area during different time periods can be calculated.
[0120] Secondly, when insufficient ambient light at night leads to a decrease in the target modality health score of the visual information source, reflecting a moderate decrease in the availability of visual information, a corresponding alertness enhancement coefficient, such as 1.5, will be calculated.
[0121] Next, by combining the expected duration of nighttime activity (2 minutes) and the alertness enhancement factor (1.5), a base alert duration is generated, for example, 2 minutes * 1.5 = 3 minutes. This means that if non-visual information sources (such as infrared sensors or sound sensors) continuously detect activity in the area for more than 3 minutes, an alert may be triggered.
[0122] Furthermore, the system will query the historical false alarm count for the valuables storage area over the past month. If it finds that there have been 3 false alarms caused by rat activity at night, the false alarm suppression factor will be adjusted based on this historical false alarm data, for example, from the initial value of 1.0 to 1.2.
[0123] Finally, based on the generated base alert duration (3 minutes) and the adjusted false alarm suppression factor (1.2), the final threshold for determining auxiliary information sources is determined, for example, 3 minutes * 1.2 = 3.6 minutes. This means that only when non-visual information sources continuously detect activity in the area for more than 3.6 minutes will they be identified as auxiliary information sources, further triggering the abnormal behavior recognition process. In this way, the sensitivity and false alarm rate of anomaly detection can be more intelligently balanced, improving the accuracy of recognition.
[0124] In some embodiments, the step of adjusting the false alarm suppression factor based on the historical false alarm count of the sensitive region includes:
[0125] Get the severity level of the alert;
[0126] Assess the processing costs of receiving alerts;
[0127] Record the management personnel's response to the alarm;
[0128] The fatigue level of management personnel is assessed based on the severity level of the alarm, the processing cost of the alarm, and the response behavior, and the false alarm suppression factor is adjusted accordingly.
[0129] Specifically, the severity level of an alarm refers to the degree of harm determined after risk assessment of different types of abnormal behavior alarms. For example, it can be divided into three levels: low, medium, and high, with the aim of distinguishing the urgency and potential impact of different alarms. The processing cost of an alarm can be understood as the resources required to handle one alarm, including time, manpower, and material resources. For example, it can be quantitatively assessed by statistically analyzing the average time managers take to handle alarms and the number of security personnel required. Managerial response behavior to alarms refers to the specific actions taken by managers after receiving an alarm, such as immediate verification, delayed processing, or ignoring it. Its purpose is to reflect the level of importance managers attach to alarms and their processing efficiency. Managerial fatigue refers to the degree of physiological and psychological fatigue exhibited by managers after long hours of work or frequent alarm handling. For example, it can be assessed by analyzing historical response times, processing durations, and false alarm verification results to reflect the current workload and alertness level of managers. In practical applications, the false alarm suppression factor is a parameter used to adjust the alarm trigger threshold or alarm response intensity. Its adjustment aims to balance the false alarm rate and the missed alarm rate, ensuring optimal alarm performance in different situations.
[0130] The technical solution of this application, by introducing alarm severity levels, processing costs, and manager response behavior, enables a more comprehensive assessment of the actual impact of alarms and the real-time status of managers. Specifically, alarm severity levels and processing costs are used to quantify the objective impact of alarms, while manager response behavior provides clues to their subjective processing efficiency and alertness. Based on this multi-dimensional information, manager fatigue can be assessed more accurately. Because manager fatigue is taken into account, the adjustment of the false alarm suppression factor no longer relies solely on the statistics of historical false alarms, but can dynamically adapt to changes in the actual workload and alertness of managers. When managers are fatigued, the intensity of the false alarm suppression factor can be appropriately reduced to avoid ignoring genuine alarms due to over-suppression; conversely, when managers are in good condition, the suppression intensity can be appropriately increased to reduce unnecessary interference. This dynamic adjustment mechanism makes the false alarm suppression strategy more intelligent and human-centered.
[0131] Through the above technical solution, this application overcomes the limitations of traditional methods that rely solely on historical false alarm counts to adjust the false alarm suppression factor. Specifically, by comprehensively considering the severity level of the alarm, processing costs, and the response behavior of management personnel, the fatigue state of management personnel can be more accurately assessed, thereby making the adjustment of the false alarm suppression factor more refined and adaptive. This not only effectively reduces the false alarm rate but also alleviates alarm fatigue among management personnel. This dynamic adjustment based on multi-dimensional information and personnel status significantly improves the intelligence level and practical application effect of the false alarm suppression strategy.
[0132] For example, in a security monitoring system, when an alarm is triggered, the severity level of the alarm is first determined. For instance, if "personnel trespassing into a restricted area" is detected, the severity level is set to "high"; if "item left behind" is detected, the severity level is set to "medium". Simultaneously, the cost of handling the alarm is assessed. For example, a "high" alarm might require immediately dispatching two security personnel for on-site verification, which is costly; while a "medium" alarm might only require remote video verification, which is less costly. Next, the manager's response to the alarm is recorded, such as whether the manager clicked the "verify" button within a specified time or chose "ignore". Based on this information, the manager's current fatigue level is assessed. For example, if a manager handles multiple "high" alarms consecutively within a short period and their response time is significantly prolonged, they might be judged to be in a state of moderate fatigue. Based on this fatigue level, the false alarm suppression factor is dynamically adjusted. Specifically, if managers are moderately fatigued, the false alarm suppression factor may be slightly reduced, allowing some potentially risky alarms that could otherwise be suppressed to be triggered, thus avoiding missed alarms due to fatigue. Conversely, if managers are in good spirits, the false alarm suppression factor may be appropriately increased to reduce unnecessary interference. In this way, the false alarm suppression factor can be intelligently adjusted based on the actual impact of alarms and the real-time status of managers, thereby optimizing performance.
[0133] In some embodiments, the step of calculating the alertness enhancement coefficient based on the degree of visual information availability reduction reflected by the target modality health score of the visual information source includes:
[0134] Obtain the verification results of the alarm;
[0135] Record the actions taken by management personnel in response to the alarm;
[0136] The degree of reduction in the availability of visual information when an alarm is triggered;
[0137] Calculate the fatigue impact coefficient based on the verification results of the alarm and the management personnel's response to the alarm;
[0138] Based on the fatigue impact coefficient and the degree of reduction in visual information availability when the alarm is triggered, the calculation logic of the alertness enhancement coefficient is adjusted to obtain the alertness enhancement coefficient.
[0139] The process involves verifying triggered alarms to determine whether they represent a genuine anomaly (e.g., intrusion, fall) or a false alarm (e.g., pet activity, light changes), thus obtaining the verification results. These results serve as crucial evidence for assessing accuracy and the effectiveness of management responses. The specific actions taken by management upon receiving an alarm are recorded, such as whether they reviewed surveillance footage, dispatched security personnel, or ignored the alarm. These actions reflect the level of importance management places on the alarm and their processing efficiency. The degree of visual information availability degradation refers to the extent to which the quality of visual information deteriorates due to insufficient ambient light, obstructions, camera malfunctions, etc., when an alarm occurs. This directly impacts the reliability of visual information sources in identifying abnormal behavior.
[0140] Furthermore, based on the verification results of the alarms and the management personnel's response actions, a fatigue impact coefficient is calculated. The fatigue impact coefficient aims to quantify the potential impact of managerial fatigue on alertness when handling alarms. For example, if managers frequently ignore real alarms or overreact to false alarms, it may indicate that they are fatigued, in which case increased alertness is necessary. In practical applications, the fatigue impact coefficient can be comprehensively evaluated by analyzing historical data, combined with managerial identity information, work shifts, historical alarm response times, false alarm handling durations, and alarm verification results.
[0141] Based on this, the calculation logic of the alertness enhancement coefficient is adjusted according to the calculated fatigue impact coefficient and the degree of reduction in visual information availability when the alarm is triggered. This means that the calculation of the alertness enhancement coefficient is no longer static, but dynamically adapts to the fatigue state of managers and the actual availability of visual information. For example, when the fatigue impact coefficient is high, even if the degree of reduction in visual information availability is not high, the alertness enhancement coefficient may be increased to compensate for possible negligence by managers; conversely, when the degree of reduction in visual information availability is high, the alertness will be further adjusted according to the fatigue impact coefficient to ensure that non-visual information sources receive a more reasonable weight allocation when visual information is limited.
[0142] The technical solution of this application calculates the fatigue impact coefficient by incorporating alarm verification results, manager response actions, and the degree of reduction in visual information availability when the alarm is triggered, and dynamically adjusts the calculation logic of the alertness enhancement coefficient accordingly. This allows the calculation of the alertness enhancement coefficient to more comprehensively consider dynamic factors in the actual operating environment, especially the human factors of managers and the real-time quality of visual information. In this way, the risks in the current situation can be assessed more intelligently, and the sensitivity to abnormal behavior can be adjusted accordingly, thereby avoiding misjudgments or missed reports caused by manager fatigue or fluctuations in the quality of visual information.
[0143] Through the aforementioned technical solution, the calculation of the alertness enhancement coefficient becomes more precise and adaptive. It can dynamically adjust the alertness level based on the manager's actual work status and historical performance, as well as the specific environmental conditions at the time of the alarm. This not only improves the accuracy and robustness of abnormal behavior pattern recognition but also effectively reduces false alarm and false negative rates. This technical solution, through in-depth consideration of human and environmental factors, demonstrates superior performance in complex and ever-changing application scenarios.
[0144] For example, suppose in a sensitive area of a factory, insufficient ambient light at night is detected, resulting in a moderate reduction in the availability of visual information. In this case, it is necessary to calculate an alertness enhancement factor to adjust the threshold for determining the source of auxiliary information.
[0145] First, obtain the verification results of recently triggered alerts. For example, there were 5 alerts in the past week, of which 3 were verified as genuine anomalies and 2 were verified as false alarms. Simultaneously, the management's response actions to these alerts were recorded; for genuine anomalies, management promptly reviewed and dispatched security personnel; for false alarms, management marked them after verification.
[0146] Next, based on these verification results and response actions, and in conjunction with a pre-set evaluation model, the current fatigue impact coefficient is calculated. For example, if managers exhibit longer response times or frequent ignoring of false alarms, the fatigue impact coefficient will be calculated as higher. Assume the currently calculated fatigue impact coefficient is 0.8 (range 0-1, higher values indicate greater fatigue impact).
[0147] Then, this fatigue impact coefficient of 0.8 is combined with the degree of visual information availability reduction (moderate) at the time of the current alarm trigger to adjust the calculation logic of the alertness enhancement coefficient. For example, in the basic calculation logic, a moderate degree of visual information availability reduction may correspond to an alertness enhancement coefficient of 1.5. However, due to the high fatigue impact coefficient, the alertness enhancement coefficient will be further increased from 1.5 to 1.8 according to the adjustment logic to compensate for the possible decrease in alertness caused by fatigue among management personnel.
[0148] Ultimately, this adjusted alertness enhancement coefficient of 1.8 will be used to generate the basic alert duration and ultimately determine the judgment threshold for auxiliary information sources, thereby enabling more timely and accurate identification of potential abnormal behavior when managers may be fatigued and have limited visual information.
[0149] In some embodiments, the step of calculating the alertness enhancement coefficient based on the degree of visual information availability reduction reflected by the target modality health score of the visual information source includes:
[0150] Based on the target modal health score from the visual information source and the modal health score from the non-visual information source, a preliminary judgment of behavioral intent is made to obtain a preliminary judgment result of behavioral intent.
[0151] Based on the preliminary judgment of behavioral intent, obtain the corresponding visual information quality dependence degree;
[0152] The alertness enhancement coefficient is calculated based on the degree of reduction in the availability of the visual information and the degree of dependence on the quality of the visual information.
[0153] Specifically, based on the target modal health scores from visual information sources and the modal health scores from non-visual information sources, preliminary semantic analysis is performed on the currently observed activity to infer potential behavioral intentions and achieve a preliminary judgment of behavioral intentions. For example, a multimodal fusion model can be used to combine visual features (such as motion trajectory and posture) and non-visual features (such as sound, thermal imaging, and environmental sensor data) to identify common behavioral patterns, such as "normal walking," "standing still," "loitering," and "carrying objects," thereby obtaining a preliminary judgment of behavioral intentions. The purpose is to provide more refined contextual information for subsequent alertness adjustments.
[0154] The dependence on visual information quality can be understood as, for each behavioral intention, pre-evaluating or dynamically assessing the sensitivity of that behavioral pattern to visual information quality based on the initial judgment result. For example, for behaviors like "normal walking," the dependence on visual information quality may be low because its main features can be effectively captured by non-visual information (such as thermal imaging or radar); while for behaviors like "fine manipulation" or "object recognition," the dependence on visual information quality may be higher. This dependence can be a preset value or a parameter dynamically adjusted based on historical data and expert experience. The purpose is to quantify the visual information requirements of different behaviors, making the increase in alertness more reasonable.
[0155] In practical applications, the alertness enhancement coefficient is calculated using a weighted or functional relationship based on the degree of reduction in visual information availability and the degree of dependence on visual information quality. For example, the alertness enhancement coefficient can be calculated as the product or weighted sum of the degree of reduction in visual information availability and the degree of dependence on visual information quality. When the degree of reduction in visual information availability is high and the behavioral intention is also highly dependent on visual information quality, the alertness enhancement coefficient will be significantly increased; conversely, if the behavioral intention is less dependent on visual information quality, even if visual information availability is reduced, the increase in the alertness enhancement coefficient will be relatively small. The aim is to achieve more intelligent and adaptive alertness adjustment.
[0156] The technical solution of this application overcomes the limitations of calculating the alertness enhancement coefficient solely based on the degree of decrease in visual information availability by introducing a preliminary judgment of behavioral intent and the degree of dependence on visual information quality. Specifically, firstly, by fusing and analyzing the modal health scores of the visual information source and the modal health scores of non-visual information sources, a preliminary intent judgment can be made on the behavior in the current scene. This preliminary judgment provides important contextual information for subsequent alertness adjustment. Secondly, for different behavioral intents, their inherent dependence on visual information quality can be obtained. Because different behaviors have different degrees of dependence on visual information, incorporating this factor into the alertness calculation means that alertness enhancement is no longer a simple response to a decrease in visual information quality, but rather combines the characteristics of the behavior itself. Therefore, when the availability of visual information decreases, the alertness enhancement coefficient can be adjusted more accurately according to the behavioral intent and its dependence on visual information, avoiding excessive or insufficient alertness, thereby improving the accuracy and efficiency of abnormal behavior recognition.
[0157] Through the above technical solution, this application can achieve a more refined and intelligent enhancement of alertness. Specifically, by comprehensively considering the dependence of behavioral intent and visual information quality, it can avoid indiscriminately increasing alertness when visual information quality declines, and instead differentiate treatment based on the actual situation. This not only helps reduce false alarms caused by environmental factors (such as changes in lighting) and reduce the fatigue of management personnel, but also ensures that when critical abnormal behaviors that truly require visual information for judgment occur, alertness can be enhanced in a timely and effective manner, thereby improving the accuracy of abnormal behavior identification and response efficiency.
[0158] For example, suppose in a warehouse environment, a decrease in the target modality health score of a visual information source is detected, indicating a high degree of reduction in the availability of visual information, such as due to dim lighting.
[0159] At this point, a preliminary judgment of behavioral intent will be made based on the target modal health score from visual information sources and the modal health scores from non-visual information sources (e.g., data from infrared sensors and sound sensors).
[0160] Scenario 1: If the initial judgment is "normal personnel patrol", and it is assumed that the "normal personnel patrol" behavior has a low dependence on the quality of visual information (because its main features can be captured by infrared or sound sensors), then even if the availability of visual information is reduced to a high degree, the calculated alertness enhancement coefficient will be relatively small.
[0161] Scenario 2: If the initial judgment is that "personnel are performing fine operations in the equipment area", and it is assumed that the "fine operation" behavior is highly dependent on the quality of visual information, then even if the degree of reduction in visual information availability is the same as in Scenario 1, the calculated alertness enhancement coefficient will be significantly increased.
[0162] In this way, the alertness enhancement coefficient can be dynamically and intelligently adjusted according to the specific behavioral intention and the degree of reliance on visual information, thereby avoiding over-vigilance in non-critical behaviors while ensuring sufficient vigilance in critical behaviors.
[0163] In some embodiments, the step of calculating the fatigue impact coefficient based on the verification result of the alarm and the manager's response to the alarm includes:
[0164] Based on the verification results of the alarm, the response actions are classified to obtain the classification results;
[0165] The fatigue influence coefficient is calculated based on the classification results and the preset fatigue weights.
[0166] The verification result of an alarm refers to the outcome obtained after investigating and verifying the triggered alarm, such as the alarm being confirmed as a real anomaly, the alarm being confirmed as a false alarm, or the alarm being unverifiable. The manager's response to an alarm refers to the specific actions taken by the manager after receiving the alarm, such as immediately conducting on-site verification, remote video confirmation, ignoring the alarm, or escalating the alarm to a higher level of handling. Based on the alarm verification result, the manager's response actions are categorized into different preset categories, thus classifying the response actions. For example, for an alarm confirmed as a false alarm, the manager's action of "immediately conducting on-site verification" might be classified as "overreaction," while the action of "ignoring" might be classified as "appropriate handling." The classification result is this categorized information. The preset fatigue weight is a pre-set value for each classification result, used to quantify the degree of impact of that type of response action on the manager's fatigue state. For example, frequent "overreactions" might be assigned a higher fatigue weight, while "appropriate handling" might be assigned a lower fatigue weight. The fatigue impact coefficient is a comprehensive index calculated based on these classification results and their corresponding fatigue weights. It is used to reflect the current level of fatigue of managers and its impact on subsequent alarm handling.
[0167] The technical solution of this application, by meticulously classifying the verification results of alarms and the response actions of management personnel, and combining this with preset fatigue weights for calculation, can more accurately quantify the fatigue state of management personnel during alarm handling. Different types of response actions, under different verification results, impose different psychological and physiological burdens on management personnel. For example, the effort required and the resulting fatigue from handling an alarm verified as a genuine anomaly differ significantly from that from handling an alarm verified as a false alarm. By taking these factors into account, it avoids simply generalizing all response actions, thus making the calculation of the fatigue impact coefficient more refined and contextualized. This allows for a more realistic reflection of the actual working state of management personnel, providing a more reliable basis for subsequent adjustments to the alertness enhancement coefficient.
[0168] The aforementioned technical solution enables a refined assessment of managerial fatigue, making the calculation of the fatigue impact coefficient more accurate and targeted. This facilitates the intelligent identification of potential managerial fatigue and allows for adjustments to alarm handling strategies accordingly. For example, when managerial fatigue is high, the priority of certain non-critical alarms can be appropriately reduced, or supplementary prompts for critical alarms can be added, effectively preventing missed alarms or misjudgments due to managerial fatigue. This deep consideration of human factors significantly improves the robustness and practicality of multimodal abnormal behavior pattern recognition methods, ensuring high efficiency and reliability even in complex and ever-changing application environments.
[0169] In some embodiments, the step of calculating the fatigue influence coefficient based on the classification result and a preset fatigue weight includes:
[0170] Obtain the manager's identity information, current work shift information, and historical fatigue accumulation data. The historical fatigue accumulation data includes the manager's historical alarm response time, historical false alarm handling time, and historical alarm verification results under different work shifts.
[0171] Based on the manager's identity information and the current work shift information, the basic fatigue weight corresponding to the manager and the work shift is obtained from the preset initial weight configuration;
[0172] Based on the manager's historical fatigue accumulation data, calculate the fatigue accumulation correction factor for the manager in the current work shift;
[0173] The preset fatigue weight is adjusted based on the basic fatigue weight and the fatigue accumulation correction factor.
[0174] The fatigue influence coefficient is calculated based on the classification results and the adjusted preset fatigue weights.
[0175] Specifically, manager identity information can refer to any data used to uniquely identify a manager, such as employee number, username, etc. Current work shift information refers to the manager's current work period or shift schedule, such as morning shift, afternoon shift, night shift, etc. Historical fatigue accumulation data refers to long-term recorded historical data related to the fatigue state of a specific manager, used to quantify the manager's fatigue level. This data specifically includes the manager's historical alarm response time, historical false alarm handling time, and historical alarm verification results under different work shifts. Historical alarm response time refers to the time interval between receiving an alarm and taking the first response action; historical false alarm handling time refers to the time spent by the manager processing alarms verified as false alarms; historical alarm verification results record whether the alarm was ultimately confirmed as a real anomaly or a false alarm. The preset initial weight configuration can be understood as a database or configuration table storing the initial fatigue weights of different managers under different work shifts. The base fatigue weight refers to the initial value obtained from the preset initial weight configuration based on the manager's identity information and current work shift information, serving as the starting point for fatigue weight adjustment. The fatigue accumulation correction factor is a coefficient calculated based on the historical fatigue accumulation data of managers and used to correct the basic fatigue weight. Its purpose is to reflect the degree of fatigue accumulated by managers due to long working hours or specific shifts. By combining the basic fatigue weight with the fatigue accumulation correction factor, the fatigue weight used to calculate the fatigue impact coefficient is dynamically updated, thereby adjusting the preset fatigue weight to better reflect the actual fatigue state of managers.
[0176] The technical solution of this application achieves personalized and dynamic adjustment of fatigue weights by introducing the manager's identity information, current work shift information, and historical fatigue accumulation data. Specifically, firstly, a basic fatigue weight is obtained from a preset initial configuration based on the manager's identity and current shift, providing an initial distinction between different personnel and shifts. Subsequently, by analyzing the manager's historical alarm response time, historical false alarm handling time, and historical alarm verification results, a fatigue accumulation correction factor is calculated. This correction factor can quantify the degree of fatigue accumulated by the manager under long-term work or specific shifts. Finally, the basic fatigue weight and the fatigue accumulation correction factor are combined to dynamically adjust the preset fatigue weight used to calculate the fatigue impact coefficient. Thus, the fatigue weight is no longer static but can adaptively adjust according to the individual differences and real-time work status of the manager, enabling the subsequently calculated fatigue impact coefficient to more accurately reflect the manager's true fatigue state, effectively solving the problem of insufficient precision in fatigue weight settings in traditional technical solutions.
[0177] The above technical solution allows for dynamic adjustment of fatigue weights based on individual manager characteristics and historical performance, resulting in more accurate calculation of the fatigue impact coefficient. This not only improves the accuracy of assessing manager fatigue but also makes the adjustment of the alertness enhancement coefficient more reasonable. Consequently, when the availability of visual information decreases, it enables more effective use of non-visual information sources for abnormal behavior judgment, significantly improving the overall accuracy and robustness of abnormal behavior pattern recognition and reducing missed or false alarms caused by manager fatigue.
[0178] For example, suppose a monitoring center has manager A and manager B. Manager A's historical alarm response time during night shifts is generally longer, and false alarm handling time is also longer, with historical verification results showing a relatively high false alarm rate during night shifts. Manager B, on the other hand, has relatively better historical data performance during night shifts. When manager A works the night shift, a basic fatigue weight is first obtained from the initial weight configuration based on their identity information and night shift information. Subsequently, manager A's historical cumulative fatigue data is analyzed to calculate a higher fatigue accumulation correction factor to reflect the potentially higher level of fatigue during the night shift. This correction factor is used to adjust the preset fatigue weight, assigning it a higher value in manager A's night shift scenario. Conversely, when manager B works the night shift, their historical data may lead to a lower calculated fatigue accumulation correction factor, resulting in a relatively lower adjusted fatigue weight. In this way, the fatigue weight can be dynamically adjusted according to the individual differences and historical performance of different managers, so that the fatigue impact coefficient can more accurately reflect the actual fatigue state of managers, thereby more reasonably adjusting the alertness enhancement coefficient and optimizing the identification effect of abnormal behavior.
[0179] This application also proposes a multimodal abnormal behavior pattern recognition system, such as... Figure 2 As shown, a multimodal abnormal behavior pattern recognition system 100 is provided, the system comprising:
[0180] The modal health score generation module 10 is used to acquire data from multiple information sources in real time and generate a modal health score for each information source that represents the reliability of its data. The information sources include visual information sources and at least one non-visual information source.
[0181] The correction module 20 is used to correct the modal health score of the visual information source based on the acquired ambient lighting information and a preset correction rule to obtain the target modal health score of the visual information source. The preset correction rule includes the mapping relationship between visual information quality and lighting conditions.
[0182] The weight determination module 30 is used to determine the weights of the visual information source and the non-visual information source in the abnormal behavior judgment based on the target modal health score of the visual information source and the modal health score of the non-visual information source, according to a preset weight allocation strategy. The weight allocation strategy satisfies that the weight value is positively correlated with the modal health score of the corresponding information source, and the sum of all weights is a fixed value.
[0183] The abnormal behavior identification module 40 is used to identify abnormal behavior by performing fusion analysis based on the monitoring data from various information sources and the weights of each information source.
[0184] The multimodal abnormal behavior pattern recognition system proposed in this application aims to effectively address the shortcomings of traditional abnormal behavior recognition technologies in terms of accuracy and robustness under complex and variable environments, especially when visual information quality is limited, through modular design and intelligent data processing. By introducing core functional modules such as modal health score generation, ambient lighting correction, adaptive weight allocation, and multimodal fusion analysis, this system can evaluate and dynamically adjust the reliability of different information sources in real time. Therefore, even when faced with dynamic degradation of visual information quality due to changes in ambient lighting, it can still accurately identify high-risk preparatory behaviors that rely on subtle movements and are highly concealed.
[0185] The above embodiments have described the real-time acquisition of data from multiple information sources, the generation of a modal health score for each information source representing its data reliability, the correction of the modal health score of the visual information source based on the acquired ambient lighting information and a preset correction rule to obtain the target modal health score of the visual information source, the determination of the weights of the visual and non-visual information sources in abnormal behavior judgment based on the target modal health score of the visual information source and the modal health scores of the non-visual information sources and a preset weight allocation strategy, and the fusion analysis based on the monitoring data of each information source and the weights of each information source to identify abnormal behavior. The specific methods and working principles are not elaborated here. It should be emphasized that this application implements the above method steps in a modular system manner.
[0186] This application's multimodal abnormal behavior pattern recognition system, through its modular design, clearly defines functional responsibilities. This allows the system to identify and utilize previously considered "noise" or "low-confidence" blurred visual signals when faced with dynamic degradation in visual information quality due to changes in ambient lighting. These signals are then effectively combined with reliable but insufficiently granular auxiliary modalities (such as dwell time). Compared to existing single-modal or fixed-weight fusion systems, this application's system adaptively adjusts the information analysis mode. It no longer simply discards or under-weights low-confidence visual signals, but instead achieves deeper correlation analysis and meaning reconstruction through modal health scores and dynamic weight allocation mechanisms. Therefore, this application's system significantly improves the accuracy and robustness of identifying high-risk preparatory behaviors that rely on subtle movements and are highly concealed, even under conditions of severely limited visual information, providing a more reliable abnormal behavior early warning capability for high-security monitoring environments.
[0187] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A multi-modal based anomalous behavior pattern recognition method, characterized in that, The method comprises: real-time acquisition of data of multiple information sources, generation of a modal health score representing data reliability of each information source, the information sources including visual information sources and at least one non-visual information source; wherein the modal health score is a numerical value between 0 and 1, and for the visual information source, at least one of image clarity, noise level and occlusion rate is quantified; for the non-visual information source, at least one of data integrity, sensor accuracy and environmental interference degree is quantified; based on a preset correction rule, the modal health score of the visual information source is corrected based on the acquired ambient light information to obtain a target modal health score of the visual information source, the preset correction rule including a mapping relationship between visual information quality and light conditions; based on a preset weight allocation strategy, the weight of the visual information source and the non-visual information source in abnormal behavior judgment is determined according to the target modal health score of the visual information source and the modal health score of the non-visual information source, wherein the weight allocation strategy satisfies that the weight value is positively correlated with the modal health score of the corresponding information source, and the sum of all weights is a fixed value; fusing analysis is performed according to the monitoring data of each information source and the weight of each information source to identify abnormal behavior.
2. The multi-modal based abnormal behavior pattern recognition method of claim 1, wherein, The method further comprises, after the step of determining the weight of the visual information source and the non-visual information source in abnormal behavior judgment based on the target modal health score of the visual information source and the modal health score of the non-visual information source according to the preset weight allocation strategy: determining a judgment threshold of an auxiliary information source in the non-visual information source according to the degree of reduction of visual information availability reflected by the target modal health score of the visual information source. 3.The multi-modal based abnormal behavior pattern recognition method of claim 1, wherein, The step of fusing analysis according to the monitoring data of each information source and the weight of each information source to identify abnormal behavior comprises: extracting potential activity clues from the signals of the visual information source; when the potential activity clues and the auxiliary information source in the non-visual information source with high weight appear at the same time, identifying abnormal behavior. 4.The multi-modal based abnormal behavior pattern recognition method of claim 2, wherein, The step of determining a judgment threshold of an auxiliary information source in the non-visual information source according to the degree of reduction of visual information availability reflected by the target modal health score of the visual information source comprises: acquiring historical activity duration distribution data of a sensitive area where a person is located, and calculating an activity duration expectation value of the sensitive area according to the historical activity duration distribution data, wherein the sensitive area is a specific spatial range defined based on physical characteristics, functional attributes or historical experience, and has a higher risk of abnormal behavior or needs special attention in a specific monitoring environment; calculating an alertness improvement coefficient according to the degree of reduction of visual information availability reflected by the target modal health score of the visual information source; generating a basic warning duration according to the activity duration expectation value of the sensitive area and the alertness improvement coefficient; adjusting a false alarm suppression factor according to the historical false alarm times of the sensitive area; determining a judgment threshold of the auxiliary information source according to the base alert duration and the false alarm suppression factor.
5. The multi-modal based anomalous behavior pattern recognition method of claim 4, wherein, The step of adjusting the false alarm suppression factor according to the historical false alarm times of the sensitive area comprises: obtaining a severity level of the alarm; evaluating a processing cost of the alarm; recording a response behavior of the manager to the alarm; evaluating a fatigue state of the manager according to the severity level of the alarm, the processing cost of the alarm and the response behavior, and adjusting the false alarm suppression factor according to the fatigue state of the manager.
6. The multi-modal based abnormal behavior pattern recognition method of claim 4, wherein, The step of calculating the alertness improvement coefficient according to the reduced degree of visual information availability reflected by the target modality health score of the visual information source comprises: obtaining a verification result of the alarm; recording a response action of the manager to the alarm; obtaining a reduced degree of visual information availability at the time when the alarm is triggered; calculating a fatigue influence coefficient according to the verification result of the alarm and the response action of the manager to the alarm; adjusting the calculation logic of the alertness improvement coefficient according to the fatigue influence coefficient and the reduced degree of visual information availability at the time when the alarm is triggered, to obtain the alertness improvement coefficient.
7. The multi-modal based abnormal behavior pattern recognition method of claim 4, wherein, The step of calculating the alertness improvement coefficient according to the reduced degree of visual information availability reflected by the target modality health score of the visual information source comprises: performing a preliminary judgment of the behavior intention according to the target modality health score of the visual information source and the modality health score of the non-visual information source, to obtain a preliminary judgment result of the behavior intention; obtaining a visual information quality dependence degree corresponding to the preliminary judgment result of the behavior intention; calculating the alertness improvement coefficient according to the reduced degree of visual information availability and the visual information quality dependence degree.
8. The multi-modal based abnormal behavior pattern recognition method of claim 6, wherein, The step of calculating the fatigue influence coefficient according to the verification result of the alarm and the response action of the manager to the alarm comprises: classifying the response action according to the verification result of the alarm, to obtain a classification result; calculating the fatigue influence coefficient according to the classification result and a preset fatigue weight.
9. The multi-modal based abnormal behavior pattern recognition method of claim 8, wherein, The step of calculating the fatigue influence coefficient according to the classification result and a preset fatigue weight comprises: obtaining identity information, current work shift information and historical fatigue accumulation data of the manager, wherein the historical fatigue accumulation data comprises historical alarm response times, historical false alarm processing durations and historical alarm verification results of the manager in different work shifts; obtaining a base fatigue weight corresponding to the manager and the work shift from a preset initial weight configuration according to the identity information of the manager and the current work shift information; calculating a fatigue accumulation correction factor of the manager in the current work shift according to the historical fatigue accumulation data of the manager; adjusting the preset fatigue weight according to the base fatigue weight and the fatigue accumulation correction factor; calculating the fatigue influence coefficient according to the classification result and the adjusted preset fatigue weight.
10. A multi-modal based anomalous behavior pattern recognition system, characterized in that, The system comprises: The modal health score generation module is configured to acquire data of multiple information sources in real time, and generate a modal health score of each information source representing the reliability of the data of the information source, the information sources including a visual information source and at least one non-visual information source; The correction module is configured to correct the modal health score of the visual information source based on a preset correction rule according to the acquired ambient light information, to obtain a target modal health score of the visual information source, and the preset correction rule includes a mapping relationship between visual information quality and light conditions; The weight determination module is configured to determine the weights of the visual information source and the non-visual information source in the abnormal behavior judgment based on a preset weight allocation strategy according to the target modal health score of the visual information source and the modal health score of the non-visual information source, wherein the weight allocation strategy satisfies that the weight value is positively correlated with the modal health score of the corresponding information source, and the sum of all weights is a fixed value; The abnormal behavior recognition module is configured to perform fusion analysis according to the monitoring data of each information source and the weights of each information source, and recognize the abnormal behavior.
Citation Information
Patent Citations
Systems and methods for sensory and cognitive profiling
CN104871160A
Personnel abnormal behavior detection method and system based on visual language large model
CN119992641A
Emotion recognition method and system based on visual and auditory collaboration
CN120852890A
Machine room robot inspection system based on industrial vision
CN120932163A
KR20240105623A