Unsafe behavior risk profile generation method based on multimodal data fusion and dynamic visualization
By real-time perception and dynamic regulation of the information value of data sources, identifying and weakening high-frequency, low-value redundant data, and dynamically adjusting the fusion contribution, the problem of redundant data interference in multimodal data fusion is solved, accurate risk portraits are generated, and the response capability and early warning level of the safety monitoring system are improved.
Patent Information
- Application Number
- CN202510472639.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-04-16
AI Technical Summary
In the existing technology of multimodal data fusion, high-frequency, low-value redundant data interferes with key signals, resulting in the system's inability to identify high-risk behaviors and reducing the accuracy of safety monitoring and accident warnings.
By real-time perception and dynamic regulation of the information value of the data source, identifying high-frequency, low-value redundant data, dynamically adjusting the fusion contribution, weakening the impact of redundant data, and restoring its contribution when the value of the data source recovers, and using machine learning models to evaluate the information value density, adaptive recovery is achieved.
It improves the accuracy and stability of multimodal data fusion, generates more accurate and dynamic risk portraits of unsafe behaviors, and improves the response capability and accident warning level of the safety monitoring system.
Smart Images

Figure CN120354359B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of safety management and risk assessment, and in particular to a method for generating risk profiles of unsafe behaviors based on multimodal data fusion and dynamic visualization. Background Art
[0002] Unsafe behavior risk profiles generated through multimodal data fusion and dynamic visualization integrate heterogeneous information from multiple data sources (such as video surveillance, sensor data, behavior logs, and environmental information). Using multimodal data fusion techniques, the behavioral data of individuals or equipment is comprehensively processed, analyzed, and deeply mined to identify potential unsafe behaviors and their associated risks. Subsequently, dynamic visualization techniques are used to present the analysis results in intuitive and easy-to-understand formats, such as risk heat maps, time series analysis charts, 3D scene reconstruction, and behavior trajectory maps, thereby constructing a dynamically evolving risk profile. This profile not only reflects the current safety status in real time but can also be dynamically updated based on new data, assisting safety managers in real-time monitoring, intelligent early warning, and informed decision-making, thereby optimizing management strategies, improving safety levels, and reducing accident rates.
[0003] Existing technologies have the following shortcomings: They typically assess the credibility of different data sources by analyzing their historical performance (e.g., recognition accuracy, data loss rate, and consistency with other modalities) and assign different fusion contributions accordingly. However, when a data source exhibits high-frequency, low-value redundancy—that is, frequently uploaded information containing little valid content (e.g., repeated background images, low-difference logs)—the data fusion process will continuously dilute the high-value information contained in other modalities, obscuring key risk signals and reducing overall recognition accuracy. If the system cannot adaptively adjust the fusion contribution of this data source, the high-frequency, redundant data will occupy a significant portion of its attention resources, marginalizing or even completely ignoring truly abnormal but less frequent critical signals (e.g., sudden equipment vibrations, dangerous movements, and environmental changes). Ultimately, the system may fail to identify high-risk behaviors (e.g., personnel straying into hazardous areas or equipment overload) at critical moments, resulting in no warning or response when accidents occur, seriously threatening personnel safety and system stability.
[0004] The above information disclosed in this Background section is only for enhancement of understanding of the background of the present disclosure and therefore it may contain information that does not form the prior art that is already known to a person of ordinary skill in the art. Summary of the Invention
[0005] The purpose of the present invention is to provide a method for generating unsafe behavior risk profiles based on multimodal data fusion and dynamic visualization. By real-time sensing and dynamically regulating the value of data source information, it effectively identifies high-frequency, low-value redundant data, avoids redundant modal interference in fusion results, and enhances the dominant position of key modal information in risk profiles. The dynamic adjustment of fusion contribution and the adaptive recovery mechanism enhance the system's ability to respond to data changes, making the generated unsafe behavior risk profiles more accurate, dynamic, and targeted, thereby improving the accuracy and reliability of overall safety monitoring and accident warnings, and solving the problems in the above-mentioned background technology.
[0006] To achieve the above objectives, the present invention provides the following technical solution: a method for generating an unsafe behavior risk profile based on multimodal data fusion and dynamic visualization, comprising the following steps:
[0007] Through network communication, raw data from different modal data sources is continuously and in real time collected to ensure the freshness and completeness of the information relied upon in the overall fusion process;
[0008] After preprocessing the acquired raw data, the normalized data source data is organized into a data set;
[0009] Extract key indicators that reflect the high-frequency, low-value redundancy of data sources from the data set, conduct a comprehensive analysis of the extracted key indicators, and characterize the information value density of the data obtained by the data source;
[0010] The key indicators after comprehensive analysis are input into a pre-trained machine learning model. The machine learning model then performs an intelligent assessment of the information value density of the current data source to determine whether the current data source is in a high-frequency, low-value redundant state.
[0011] When the machine learning model determines that the current data source has entered a high-frequency, low-value redundant state, the fusion contribution of the current data source in the multimodal fusion process is reduced based on the machine learning model evaluation results, weakening the interference of the current data source on the overall judgment results. At the same time, the data changes of the current data source are continuously monitored, and the fusion contribution is adaptively restored when the "value of the data source recovers".
[0012] Preferably, continuously and in real time collecting raw data from different modality data sources through network communication includes the following specific steps:
[0013] First, uniformly configure access to various data sources to ensure identification of their communication protocols, data formats, and sampling frequencies;
[0014] Secondly, establish a stable data transmission channel and use wired and wireless communication methods to achieve real-time reception of data from various modalities;
[0015] Subsequently, reasonable sampling strategies and bandwidth allocation mechanisms are set according to the characteristics of different data sources to avoid data congestion or delays. For resource-constrained terminal devices, edge computing modules are deployed to perform preliminary data filtering and compression in advance, reducing the burden on the network and central servers.
[0016] Finally, through timestamp alignment and data caching mechanisms, the real-time collected data streams are fed into the fusion process in an orderly and synchronous manner, ensuring that subsequent analysis and processing steps are based on multimodal information with freshness, continuity and integrity.
[0017] Preferably, key indicators reflecting the high-frequency and low-value redundancy of the data source are extracted from the data set, including the difference between modes at the information layer and the rate of change of information entropy per unit time. After comprehensively analyzing the change in the difference between modes at the information layer and the rate of change of information entropy per unit time under the detection window, an inter-modal difference convergence reference value and an information density entropy reference value are generated respectively. The inter-modal difference convergence reference value and the information density entropy reference value are used to jointly characterize the information value density of the data obtained by the current data source.
[0018] Preferably, the specific steps of comprehensively analyzing the differences between modes on the information layer within the detection window to generate the inter-modal difference convergence reference value are as follows:
[0019] In order to characterize the similarity of different modes at the information trend level, we first extract the local extreme points of each mode in the current period, construct a trend structure feature set, and generate a trend similarity index. The calculation expression is as follows: , where is the trend similarity index, is modal The set of local extreme points of is modal The set of local extreme points of is modal and The number of coincidences at local extreme points, and They are modal and modal The total number of local extreme points, is modal and modal Minimum number of local extreme points;
[0020] In order to further analyze the differences between modalities in event recognition and response, an index to measure the independence of modal responses, namely the event response overlap index, is constructed. and modal The respective event response sets are identified, and some events are identified by both modalities. The event response overlap index is generated as follows: , where is the incident response overlap index, is modal and modal The number of responses to the same type of events at similar time points, is modal and modal the number of parts of the set of events identified by each party that are not identified by the other party;
[0021] Based on trend similarity index and Incident Response Overlap Index , generate the inter-modal difference convergence reference value, the calculation expression is as follows: , where is the intermodal difference convergence reference value, A very small positive number added to prevent the denominator from being zero.
[0022] Preferably, the specific steps of analyzing the rate of change of information entropy per unit time in the detection window to generate the information density entropy reference value are as follows:
[0023] The original data in a continuous time period is sliced according to a fixed amount of data, and the Shannon information entropy of each data segment is calculated. Then, the ratio of the information entropy change between each two adjacent data segments is calculated to capture the intensity and direction of the entropy fluctuation. The calculation expression is as follows: , where is the rate of change of information entropy, i.e. Segment data and The rate of change of information entropy between segment data, It is Shannon information entropy of segment data, It is Shannon information entropy of segment data, It is a normalization factor to avoid the scale inconsistency problem caused by different entropy values. A very small positive number added to prevent the denominator from being zero;
[0024] All information entropy change rates Perform nonlinear transformation and compression to generate an information density entropy reference value, which is used to comprehensively characterize the information activity of the data in the current monitoring window. The calculation expression is as follows: , where is the information density entropy reference value, is the natural base, is the sensitivity coefficient, which controls the sensitivity of the system to changes in entropy.
[0025] Preferably, the inter-modal difference convergence reference value and information density entropy reference value after comprehensive analysis are input into a pre-trained machine learning model, and the information value density coefficient is generated by the machine learning model. The information value density coefficient is used to determine whether the current data source is in a high-frequency, low-value redundant state.
[0026] Preferably, the information value density variation coefficient generated when the information value density of the current data source is intelligently evaluated by a pre-trained machine learning model is compared with a preset information value density variation coefficient reference threshold to determine whether the current data source is in a high-frequency, low-value redundant state. The determination steps are as follows:
[0027] If the information value density change coefficient is less than the pre-set information value density change coefficient reference threshold, it is judged that the current data source data is in a high-frequency, low-value redundant state; if the information value density change coefficient is greater than or equal to the pre-set information value density change coefficient reference threshold, it is judged that the current data source data is not in a high-frequency, low-value redundant state.
[0028] Preferably, when the machine learning model determines that the current data source has entered a high-frequency, low-value redundant state, based on the machine learning model evaluation results, the fusion contribution of the current data source in the multimodal fusion process is reduced, while continuously monitoring the data changes of the current data source. When the "value of the data source recovers", the specific steps for adaptively restoring the fusion contribution are as follows:
[0029] When the machine learning model determines that the current data source has entered a high-frequency, low-value redundant state, the information value density change coefficient output by the model is Compare with the set reference threshold to trigger the weight adjustment mechanism, dynamically attenuate the fusion contribution of the current data source in the multimodal fusion process. The calculation expression is as follows: , where is the effective fusion contribution after dynamic adjustment, is the original fusion contribution, is the decay rate factor, is the reference threshold of the information value density change coefficient, is the decay curve control factor;
[0030] After continuously monitoring the data source, once the machine learning model re-determines the change coefficient of its information value density Greater than or equal to the threshold , indicating that the data source content changes actively and the information quality is restored. At this time, the fusion contribution of the current data source should be gradually improved based on the adaptive recovery strategy. The formula is as follows: , where is the fusion weight after adaptive recovery, is the minimum weight after being weakened, is the recovery rate adjustment factor, is the restoration sensitivity control parameter.
[0031] In the above technical solution, the technical effects and advantages provided by the present invention are:
[0032] The present invention realizes the real-time perception and dynamic regulation of the information value of the data source, which significantly improves the quality of the fused data and the accuracy of the decision-making results. Specifically, this method avoids the problem of redundant modalities "taking the lead" in data fusion by introducing information value density assessment and high-frequency low-value redundancy identification mechanism, and ensures the dominant position of truly critical and identification-significant modal information in the generation of risk portraits. At the same time, the dynamic adjustment of the fusion contribution and the adaptive recovery mechanism enable the system to respond in a timely manner according to changes in the state of the data source, maintaining the stability and robustness of the fusion process. The unsafe behavior risk portrait finally generated will be more accurate, dynamic and targeted, and can more effectively identify high-risk behaviors, highlight risk priorities, and provide safety managers with more insightful auxiliary decision-making basis, thereby improving the overall safety monitoring system's response capability and accident warning level. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, a brief introduction to the drawings required for use in the embodiments will be given below. Obviously, the drawings described below are only some embodiments recorded in the present invention. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.
[0034] Figure 1 This is a flow chart of the method for generating unsafe behavior risk profiles based on multimodal data fusion and dynamic visualization of the present invention. DETAILED DESCRIPTION
[0035] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, these example embodiments are provided so that the description of this disclosure will be thorough and complete and will fully convey the concepts of the example embodiments to those skilled in the art.
[0036] The present invention provides Figure 1 The method for generating unsafe behavior risk profiles based on multimodal data fusion and dynamic visualization includes the following steps:
[0037] Through network communication, it continuously and in real time collects raw data from different modal data sources (such as cameras, sensors, and log systems), ensuring the freshness and completeness of the information relied upon in the overall fusion process.
[0038] Through network communication, raw data from various modal data sources (such as cameras, sensors, and logging systems) is continuously and in real time collected to ensure that the system can quickly capture the latest environmental or operational status and respond to potential changes immediately. To achieve efficient collection, it is necessary to properly configure the sampling frequency and bandwidth to ensure data integrity while controlling transmission costs. For some front-end devices, edge computing strategies can also be adopted to filter redundant or invalid data and upload only valuable information to the central system, thereby reducing network burden. The core significance of real-time collection is to provide the most authentic and timely basic data support for subsequent data processing, analysis, and judgment, ensuring the efficiency and accuracy of the entire system's perception chain.
[0039] The continuous and real-time collection of raw data from different modal data sources (such as cameras, sensors, log systems, etc.) through network communication includes the following specific steps: First, the system needs to uniformly configure access to various data sources to ensure that its communication protocol, data format and sampling frequency can be identified; second, a stable data transmission channel is established, and real-time reception of data from various modalities is achieved through wired networks, wireless communications or field buses; then, reasonable sampling strategies and bandwidth allocation mechanisms are set according to the characteristics of different data sources to avoid data congestion or delays; for resource-constrained terminal devices, edge computing modules can be deployed to perform preliminary data filtering and compression in advance to reduce the burden on the network and central server; finally, through timestamp alignment and data caching mechanisms, the real-time collected data streams are sent to the fusion process in an orderly and synchronous manner to ensure that subsequent analysis and processing links are based on multimodal information with freshness, continuity and integrity.
[0040] After preprocessing the acquired raw data, the normalized data source data is organized into a data set;
[0041] After preprocessing the acquired raw data, the normalized data source data is organized into a data set. This means that after completing preprocessing operations such as data cleaning, format conversion, and anomaly removal, the system will structure and integrate various types of data (such as video frames, sensor readings, behavior logs, etc.) according to unified standards to build an orderly and manageable data set for subsequent indicator extraction and model analysis. The specific steps include: first, time-aligning the data of each modality to ensure the comparability of multi-source data; second, slicing and classifying the data according to a certain time window or event label; then converting the data into a unified format (such as tensor, table, sequence, etc.) and unifying the encoding method; finally, organizing and storing the processed data according to metadata tags such as source, time, and location to form a structured data set for subsequent unified analysis and call. The core purpose of this process is to lay a stable and standardized foundation for subsequent feature extraction and model evaluation.
[0042] Extract key indicators that reflect the high-frequency, low-value redundancy of data sources from the data set, conduct a comprehensive analysis of the extracted key indicators, and characterize the information value density of the data obtained by the data source;
[0043] Key indicators reflecting the high-frequency, low-value redundancy of the data source are extracted from the data set, including the difference between modes at the information layer (such as event triggering, trends) and the rate of change of information entropy per unit time. After comprehensively analyzing the changes in the difference between modes at the information layer (such as event triggering, trends) and the rate of change of information entropy per unit time under the detection window, the inter-modal difference convergence reference value and the information density entropy reference value are generated respectively. The inter-modal difference convergence reference value and the information density entropy reference value are used to jointly characterize the information value density of the data obtained by the current data source.
[0044] Information value density refers to the ratio between the amount of effective information contained in a unit of time or a unit of data and the total amount of information. It is used to measure the richness of "useful content" contained in the data. It not only reflects dimensions such as data variability, information entropy, and event triggering frequency, but also focuses on whether these changes have practical significance and whether they can provide support for decision-making. In multimodal data processing or fusion scenarios, data sources with high information value density tend to have stronger recognition, differentiation, and risk perception capabilities; on the contrary, if a data source continuously uploads repetitive, single, and unchanging content, even if it is frequently updated, its information value density may be very low, which can easily lead to "high-frequency, low-value redundancy" and interfere with the overall analysis results. Therefore, accurately assessing information value density helps the system to rationally allocate resources, dynamically adjust data weights, and improve the quality of fusion and the accuracy of risk identification.
[0045] When the intermodal differences in information (such as event trigger frequency and trend changes) gradually approach zero, and the data within the current modality itself varies minimally, it can generally be considered that the data source is in a state of high-frequency, low-value redundancy. This phenomenon indicates that the modality lacks independence or complementarity with other modalities, and the information it provides no longer offers differentiated value in the fusion process. Furthermore, low internal variation indicates a high degree of repetitiveness or a lack of dynamic data content. Despite frequent data uploads, no new information is actually introduced, resulting in low information entropy and reduced value density. In multimodal fusion scenarios, the system relies on complementary information from each modality to improve overall recognition accuracy and judgment. If a modality exhibits neither internal information variation nor converges with other modalities, its data is likely to become redundant input, diluting the signals of other high-value modalities and reducing fusion efficiency and risk identification accuracy. Therefore, intermodal differences approaching zero and minimal internal variation are important signals for identifying high-frequency, low-value redundancy, and should prompt the system to dynamically adjust its fusion contribution.
[0046] The specific steps for comprehensively analyzing the differences between modes at the information layer within the detection window to generate the inter-modal difference convergence reference value are as follows:
[0047] In order to characterize the similarity of different modes at the information trend level, we first extract the local extreme points (such as inflection points, mutation points, and trend reversal points) of each mode in the current period, construct a trend structure feature set, and generate a trend similarity index. The calculation expression is as follows: , where is the trend similarity index, which is used to measure the synchronization of trend changes between two modes. If the index is close to 1, it means that the two modes have the same trend changes at similar time points, indicating that they are highly synchronized and there may be information redundancy. If the index is low (close to 0), it means that the trend changes of the two modes are independent and they may provide complementary information rather than redundant information. is modal The set of local extreme points of is modal The set of local extreme points of is modal and The number of coincidences at local extreme points, and They are modal and modal The total number of local extreme points, is modal and modal Minimum number of local extreme points;
[0048] The trend similarity index reflects the similarity between the two modes in the trend change structure. The closer the index is to 1, the more consistent the trend changes of the two modes are and the higher the trend redundancy is.
[0049] In order to further analyze the differences between modalities in event recognition and response, an index to measure the independence of modal responses, namely the event response overlap index, is constructed. and modal The respective event response sets are identified, and some events are identified by both modalities. The event response overlap index is generated as follows: , where is the incident response overlap index, is modal and modal The number of responses to the same type of events at similar time points, is modal and modal the number of parts of the set of events identified by each party that are not identified by the other party;
[0050] The event response overlap index is used to measure whether the responses of two modes to risk events are highly overlapped. It approaches 1, indicating that the two modes are similar in terms of behavioral response and have strong information redundancy.
[0051] Based on trend similarity index and Incident Response Overlap Index , generate the inter-modal difference convergence reference value, the calculation expression is as follows: , where is the intermodal difference convergence reference value, To prevent the denominator from being zero, a very small positive number (such as 1e-6) is added to enhance stability.
[0052] The inter-modal difference convergence reference value comprehensively measures the redundancy of the two modes at the trend and event response levels. The smaller the value, the higher the degree of information convergence between the two modes, and the more likely the current mode is in a high-frequency, low-value redundant state; the larger the value, the more different and complementary it is, and the higher the information value density.
[0053] The smaller the intermodal difference convergence reference value, generated by comprehensively analyzing the information-level differences between modalities within the detection window, the more likely the current data source is in a high-frequency, low-value redundancy state. This reference value reflects the degree of information-level differences (e.g., event trigger frequency, trend changes) between the current modality and other modalities. If this reference value approaches 0, it indicates that the information provided by this modality is highly similar or duplicated in content, trends, and structure, lacking differentiation and independence, and thus failing to contribute significant information increments to the fusion system. In this case, while this data source may be frequently uploading data, the content is highly duplicated or overlaps with other modalities, a typical example of high-frequency, low-value redundancy. Conversely, a higher intermodal difference convergence reference value indicates that the information provided by this modality is semantically unique and dynamically changing, providing complementary value, thereby enhancing the multidimensional perception capabilities of the fusion system. Therefore, it is not considered a redundant modality.
[0054] A persistently low level of information entropy per unit time typically indicates that the current data source is in a state of high-frequency, low-value redundancy. Information entropy is a key indicator of data uncertainty and information diversity. Higher entropy values indicate richer data content and greater information inclusion; conversely, lower entropy values indicate high data repetitiveness, less variation, and a homogenous content. When a data source continuously uploads information but its information entropy remains low, it indicates that while it generates a large amount of data per unit time, it is not providing new, valid information, demonstrating a state of "active form but poor content." In the multimodal fusion process, this redundant signal from the data source may mask critical, emergent, and high-value information from other modalities, leading to misjudgments or omissions in the system's identification of risk events. Therefore, a persistently low entropy value per unit time is a key indicator of a data source entering a state of high-frequency, low-value redundancy, and has direct implications for information value density assessment and fusion weight regulation.
[0055] The specific steps for analyzing the rate of change of information entropy per unit time within the detection window to generate the information density entropy reference value are as follows:
[0056] The original data in a continuous time period is sliced according to a fixed amount of data (for example, every 100 data points are a segment), and the Shannon information entropy of each segment is calculated. Then, the information entropy change between each two adjacent data segments is calculated using a ratio formula to capture the intensity and direction of the entropy fluctuation. The calculation expression is as follows: , where is the rate of change of information entropy, i.e. Segment data and The rate of change of information entropy between segment data, It is Shannon information entropy of segment data, It is Shannon information entropy of segment data, It is a normalization factor to avoid the scale inconsistency problem caused by different entropy values. To prevent the denominator from being zero, a very small positive number (such as 1e-6) is added to enhance stability;
[0057] The above steps measure the local entropy dynamics through proportional changes, which can more sensitively identify the difference between "static redundancy" and "sudden high values".
[0058] All information entropy change rates Perform nonlinear transformation and compression to generate an information density entropy reference value, which is used to comprehensively characterize the information activity of the data in the current monitoring window. The calculation expression is as follows: , where is the information density entropy reference value, is the natural base, It is the sensitivity coefficient, which controls the sensitivity of the control system to changes in entropy, and is equivalent to the adjustment knob of a "magnifying glass".
[0059] By performing exponential compression and weighted aggregation on the information entropy change rate, an information density entropy reference value is generated to comprehensively characterize the information activity and value density of the current data source within the monitoring window.
[0060] The smaller the information density entropy reference value generated after analyzing the rate of change of information entropy per unit time under the detection window, the more limited the change in information entropy within a certain period of time, the data content is stable, highly repetitive, lacks novelty and dynamic changes, and is a typical "high-frequency, low-value redundancy" feature, that is, although the data is uploaded frequently, it lacks actual useful information; on the contrary, when the information density entropy reference value is large, it means that the information entropy fluctuates significantly per unit time, indicating that the data source continues to produce valuable content or has undergone a state change, and is not a redundant state.
[0061] The key indicators after comprehensive analysis are input into a pre-trained machine learning model. The machine learning model then performs an intelligent assessment of the information value density of the current data source to determine whether the current data source is in a high-frequency, low-value redundant state.
[0062] The inter-modal difference convergence reference value and information density entropy reference value after comprehensive analysis are input into the pre-trained machine learning model, and the information value density coefficient is generated by the machine learning model. The information value density coefficient is used to determine whether the current data source is in a high-frequency and low-value redundant state.
[0063] A "pre-trained machine learning model" refers to a model that is pre-trained using a machine learning algorithm, based on a large amount of existing sample data and expert annotations, before the system is officially launched. This allows the model to intelligently judge unknown data. In this scenario, the model training process includes the following key steps: First, historical data is collected from multiple data sources, and key feature indicators such as inter-modal difference convergence reference values and information density entropy reference values are extracted. Second, a supervised training sample set is constructed based on the data source status (whether it is high-frequency, low-value, and redundant) annotated manually or by a rule-based system. Next, a suitable machine learning algorithm (such as random forest, support vector machine, XGBoost, or lightweight neural network) is selected and fed into the model for training, allowing it to learn the correspondence between different feature patterns and data states. After sufficient training and validation, the model can identify which feature combinations are most likely to indicate a data source is in a "high-frequency, low-value, and redundant" state, thus achieving a certain level of generalization ability.
[0064] In actual application, the model is not relearned but directly invoked as an intelligent decision-making tool. During operation, the system calculates the intermodal difference convergence reference value and information density entropy reference value for each data source in real time. These features are then fed into the trained model, which rapidly infers and outputs a quantitative coefficient representing the current "information value density" of the data source (i.e., the information value density coefficient). This coefficient reflects the effective information content of the current data per unit time. If its value falls below a preset threshold, the data source is considered to be in a high-frequency, low-value redundant state, and the system dynamically adjusts its fusion weight accordingly. Because the model is built based on prior knowledge and real-world cases, it avoids the subjectivity and limitations of manually set fixed thresholds. It adapts more flexibly and accurately to data changes in different scenarios, improving the efficiency and robustness of the entire system in identifying redundant data sources. This approach embodies the integrated "data-driven + model-driven" approach and is a key means of achieving intelligent and automated data quality management.
[0065] The machine learning model is not limited here and can achieve the convergence reference value of the difference between modalities and information density entropy reference value Conduct comprehensive analysis to generate information value density change coefficient The machine learning model can be used. In order to realize the technical solution of the present invention, the present invention provides a specific implementation method;
[0066] Information value density change coefficient The generation formula is as follows: Where, and are the convergence reference values of inter-modal differences and information density entropy reference value The preset scaling factor of and Both are greater than 0.
[0067] Preset scale factor and Refers to the calculation of the information value density change coefficient When the intermodal difference converges to the reference value and information density entropy reference value The weight factors assigned are used to adjust their relative contributions to the final result. Due to the different convergence reference values between modalities in different application scenarios and information density entropy reference value The impact on information value may be different, so it is necessary to set appropriate weights so that the information value density change coefficient The calculation can more accurately reflect the information value density of the data source. and Determined through experience, data analysis, or machine learning model training to ensure the information value density variation coefficient Maintain reasonable calculation results in different environments. Both are greater than 0, which means that the difference between modes converges to the reference value. and information density entropy reference value In calculating the information value density variation coefficient It always plays a positive role and will not be completely ignored.
[0068] It can be seen from the information value density variation coefficient that the smaller the inter-modal difference convergence reference value generated after comprehensive analysis of the differences between modes at the information layer under the detection window, and the smaller the information density entropy reference value generated after analyzing the rate of change of information entropy per unit time under the detection window, the smaller the information value density variation coefficient generated when the information value density of the current data source is intelligently evaluated by the pre-trained machine learning model, indicating that the probability that the current data source data is in a high-frequency, low-value redundant state is greater, and vice versa, the probability that the current data source data is in a high-frequency, low-value redundant state is smaller.
[0069] The information value density change coefficient generated by the pre-trained machine learning model when performing intelligent evaluation of the information value density of the current data source is compared with the pre-set information value density change coefficient reference threshold to determine whether the current data source is in a high-frequency, low-value redundant state. The judgment steps are as follows:
[0070] If the information value density change coefficient is less than the pre-set information value density change coefficient reference threshold, it is judged that the current data source data is in a high-frequency, low-value redundant state; if the information value density change coefficient is greater than or equal to the pre-set information value density change coefficient reference threshold, it is judged that the current data source data is not in a high-frequency, low-value redundant state.
[0071] When the machine learning model determines that the current data source has entered a high-frequency, low-value redundant state, the model reduces the current data source's contribution to multimodal fusion (weight distribution, attention mechanism, weighted averaging, etc.) based on the machine learning model's evaluation results, weakening the current data source's interference with the overall judgment result. At the same time, the model continuously monitors data changes in the current data source and adaptively restores the fusion contribution when the data source's value recovers.
[0072] The core role of the above steps is to achieve intelligent dynamic regulation of high-frequency, low-value, redundant data sources, ensuring the data quality and accuracy of the fusion results during the multimodal fusion process. When the machine learning model determines that a data source is in a high-frequency, low-value, redundant state, it means that the data uploaded by the data source in the current period is frequent but the information value density is low, that is, its content lacks variability, the event triggering rate is low, or it is highly repetitive. At this time, if a high fusion contribution is still maintained, it will cause the system to give it unreasonable "trust" during the fusion stage, thereby diluting the contribution of other high-value data sources, covering up key risk signals, and even causing wrong judgments. Therefore, based on the model evaluation results, the system dynamically lowers the proportion of the data source in fusion strategies such as weight allocation, attention mechanism, or feature weighting, which essentially weakens its influence on the final decision and enhances the robustness and recognition accuracy of the system.
[0073] At the same time, this step also assumes the responsibility of maintaining adaptive capabilities. The system does not simply "block" or "discard" redundant modalities. Instead, it continuously monitors changes in the information value density of their data streams, assessing whether there are signs of "value recovery" (such as increased information entropy, increased event frequency, and increased feature differentiation). Once the data source returns to an active and high-value state, the system will gradually restore its contribution to the fusion process, implementing a closed-loop control mechanism of "dynamic suppression - continuous monitoring - intelligent recovery." This mechanism ensures that the system can both quickly respond to redundant data and retain room for future integration when it regains value, avoiding "one-size-fits-all" data elimination and achieving multimodal fusion capabilities that combine high flexibility, high precision, and high fault tolerance.
[0074] When the machine learning model determines that the current data source has entered a high-frequency, low-value redundant state, the fusion contribution of the current data source in the multimodal fusion process (weight distribution, attention mechanism, weighted averaging, etc.) is reduced based on the machine learning model evaluation results. At the same time, the current data source data changes are continuously monitored. When the data source "value recovers", the fusion contribution is adaptively restored in the following specific steps:
[0075] When the machine learning model determines that the current data source has entered a high-frequency, low-value redundant state, the information value density change coefficient output by the model is Compare with the set reference threshold to trigger the weight adjustment mechanism, dynamically attenuate the fusion contribution of the current data source in the multimodal fusion process. The calculation expression is as follows: , where is the effective fusion contribution after dynamic adjustment, is the original fusion contribution, It is the decay rate factor, which controls the sensitivity of the fusion contribution to the reduction. The larger the value, the faster the fusion contribution decreases. is the reference threshold of the information value density change coefficient, is the attenuation curve control factor, which is used to control the degree of nonlinearity. Indicates accelerated downward adjustment. Indicates slow adjustment;
[0076] This step uses the exponential decay mechanism to The difference from the threshold is converted into a weight decay factor, allowing the system to quickly reduce the interference of redundant modalities on overall judgment when redundant trends are detected. This method is more flexible than linear reduction, preventing over-adjustment caused by small fluctuations, while significantly weakening severely redundant modalities, effectively improving the fusion model's ability to perceive real key information.
[0077] After continuously monitoring the data source, once the machine learning model re-determines the change coefficient of its information value density Greater than or equal to the threshold , indicating that the data source content changes actively and the information quality is restored. At this time, the fusion contribution of the current data source should be gradually improved based on the adaptive recovery strategy. The formula is as follows: , where is the fusion weight after adaptive recovery, is the minimum weight after being weakened, It is the recovery rate adjustment factor. The larger the value, the faster the recovery. It is the recovery sensitivity control parameter, which determines the coefficient of change of information value density Nonlinear effects on the recovery amplitude, Information value density change coefficient Time to accelerate recovery.
[0078] This step introduces a nonlinear recovery mechanism, ensuring that the fusion contribution is gradually restored only after a data source demonstrates a genuine "substantial recovery." Compared to directly restoring the original weights, this approach is more robust, avoiding system shocks caused by short-term data fluctuations. It also ensures the reintegration of valuable data sources, improving the comprehensiveness and dynamic adaptability of fusion judgments.
[0079] The present invention realizes the real-time perception and dynamic regulation of the information value of the data source, which significantly improves the quality of the fused data and the accuracy of the decision-making results. Specifically, this method avoids the problem of redundant modalities "taking the lead" in data fusion by introducing information value density assessment and high-frequency low-value redundancy identification mechanism, and ensures the dominant position of truly critical and identification-significant modal information in the generation of risk portraits. At the same time, the dynamic adjustment of the fusion contribution and the adaptive recovery mechanism enable the system to respond in a timely manner according to changes in the state of the data source, maintaining the stability and robustness of the fusion process. The unsafe behavior risk portrait finally generated will be more accurate, dynamic and targeted, and can more effectively identify high-risk behaviors, highlight risk priorities, and provide safety managers with more insightful auxiliary decision-making basis, thereby improving the overall safety monitoring system's response capability and accident warning level.
[0080] The above formulas are all dimensionless and numerical calculations. The formulas are obtained by collecting a large amount of data and performing software simulation to obtain the most recent real situation. The preset parameters in the formulas are set by technicians in this field according to actual conditions.
[0081] The above description is merely illustrative of certain exemplary embodiments of the present invention. It goes without saying that those skilled in the art will be able to modify the described embodiments in various ways without departing from the spirit and scope of the present invention. Therefore, the above drawings and description are illustrative in nature and should not be construed as limiting the scope of protection of the claims.
[0082] It should be noted that, in this document, if there are relational terms such as first and second, etc., they are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprises", "comprising" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device that includes a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprising a ..." does not exclude the presence of other identical elements in the process, method, article or device that includes the element.
[0083] It should be understood that in the various embodiments of the present application, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0084] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0085] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0086] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0087] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0088] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
[0089] The above description is merely illustrative of certain exemplary embodiments of the present invention. It goes without saying that those skilled in the art will be able to modify the described embodiments in various ways without departing from the spirit and scope of the present invention. Therefore, the above drawings and description are illustrative in nature and should not be construed as limiting the scope of protection of the claims.
Claims
1. A method for generating risk profiles of unsafe behaviors based on multimodal data fusion and dynamic visualization, characterized by: The following steps are involved: Through network communication, raw data from different modal data sources is continuously and in real time collected to ensure the freshness and completeness of the information relied upon in the overall fusion process; After preprocessing the acquired raw data, the normalized data source data is organized into a data set; Extract key indicators that reflect the high-frequency, low-value redundancy of data sources from the data set, conduct a comprehensive analysis of the extracted key indicators, and characterize the information value density of the data obtained by the data source; The key indicators after comprehensive analysis are input into a pre-trained machine learning model. The machine learning model then performs an intelligent assessment of the information value density of the current data source to determine whether the current data source is in a high-frequency, low-value redundant state. When the machine learning model determines that the current data source has entered a high-frequency, low-value redundant state, the model reduces the current data source's contribution to the multimodal fusion process based on the model's evaluation results, weakening the current data source's interference with the overall judgment result. At the same time, the model continuously monitors data changes in the current data source and adaptively restores its contribution to the fusion process when the data source's value recovers. When the machine learning model determines that the current data source has entered a high-frequency, low-value redundant state, the fusion contribution of the current data source in the multimodal fusion process is reduced based on the machine learning model evaluation results. At the same time, the data changes of the current data source are continuously monitored. When the data source "value recovers", the fusion contribution is adaptively restored. The specific steps are as follows: When the machine learning model determines that the current data source has entered a high-frequency, low-value redundant state, the information value density change coefficient output by the model is Compare with the set reference threshold to trigger the weight adjustment mechanism, dynamically attenuate the fusion contribution of the current data source in the multimodal fusion process. The calculation expression is as follows: , where is the effective fusion contribution after dynamic adjustment, is the original fusion contribution, is the decay rate factor, is the reference threshold of the information value density change coefficient, is the decay curve control factor; After continuously monitoring the data source, once the machine learning model re-determines the change coefficient of its information value density Greater than or equal to the threshold , indicating that the data source content changes actively and the information quality is restored. At this time, the fusion contribution of the current data source should be gradually improved based on the adaptive recovery strategy. The formula is as follows: , where is the fusion weight after adaptive recovery, is the minimum weight after being weakened, is the recovery rate adjustment factor, is the restoration sensitivity control parameter.
2. The method for generating unsafe behavior risk profiles based on multimodal data fusion and dynamic visualization according to claim 1 is characterized in that: Continuously and in real time, raw data from different modal data sources is collected through network communication, including the following specific steps: First, uniformly configure access to various data sources to ensure identification of their communication protocols, data formats, and sampling frequencies; Secondly, establish a stable data transmission channel and use wired and wireless communication methods to achieve real-time reception of data from various modalities; Subsequently, reasonable sampling strategies and bandwidth allocation mechanisms are set according to the characteristics of different data sources to avoid data congestion or delays. For resource-constrained terminal devices, edge computing modules are deployed to perform preliminary data filtering and compression in advance, reducing the burden on the network and central servers. Finally, through timestamp alignment and data caching mechanisms, the real-time collected data streams are fed into the fusion process in an orderly and synchronous manner, ensuring that subsequent analysis and processing steps are based on multimodal information with freshness, continuity and integrity.
3. The method for generating unsafe behavior risk profiles based on multimodal data fusion and dynamic visualization according to claim 1 is characterized in that: Key indicators reflecting the high-frequency, low-value redundancy of the data source are extracted from the data set, including the difference between modes at the information layer and the rate of change of information entropy per unit time. After a comprehensive analysis of the changes in the difference between modes at the information layer and the rate of change of information entropy per unit time under the detection window, the inter-modal difference convergence reference value and the information density entropy reference value are generated respectively. The inter-modal difference convergence reference value and the information density entropy reference value are used to jointly characterize the information value density of the data obtained by the current data source.
4. The method for generating unsafe behavior risk profiles based on multimodal data fusion and dynamic visualization according to claim 3 is characterized in that: The specific steps for comprehensively analyzing the differences between modes at the information layer within the detection window to generate the inter-modal difference convergence reference value are as follows: In order to characterize the similarity of different modes at the information trend level, we first extract the local extreme points of each mode in the current period, construct a trend structure feature set, and generate a trend similarity index. The calculation expression is as follows: , where is the trend similarity index, is modal The set of local extreme points of is modal The set of local extreme points of is modal and The number of coincidences at local extreme points, and They are modal and modal The total number of local extreme points, is modal and modal Minimum number of local extreme points; In order to further analyze the differences between modalities in event recognition and response, an index to measure the independence of modal responses, namely the event response overlap index, is constructed. and modal The respective event response sets are identified, and some events are identified by both modalities. The event response overlap index is generated as follows: , where is the incident response overlap index, is modal and modal The number of responses to the same type of events at similar time points, is modal and modal the number of parts of the set of events identified by each party that are not identified by the other party; Based on trend similarity index and Incident Response Overlap Index , generate the inter-modal difference convergence reference value, the calculation expression is as follows: , where is the intermodal difference convergence reference value, A very small positive number added to prevent the denominator from being zero.
5. The method for generating unsafe behavior risk profiles based on multimodal data fusion and dynamic visualization according to claim 3 is characterized in that: The specific steps for analyzing the rate of change of information entropy per unit time within the detection window to generate the information density entropy reference value are as follows: The original data in a continuous time period is sliced according to a fixed amount of data, and the Shannon information entropy of each data segment is calculated. Then, the ratio of the information entropy change between each two adjacent data segments is calculated to capture the intensity and direction of the entropy fluctuation. The calculation expression is as follows: , where is the rate of change of information entropy, i.e. Segment data and The rate of change of information entropy between segment data, It is Shannon information entropy of segment data, It is Shannon information entropy of segment data, It is a normalization factor to avoid the scale inconsistency problem caused by different entropy values. A very small positive number added to prevent the denominator from being zero; All information entropy change rates Perform nonlinear transformation and compression to generate an information density entropy reference value, which is used to comprehensively characterize the information activity of the data in the current monitoring window. The calculation expression is as follows: , where is the information density entropy reference value, is the natural base, is the sensitivity coefficient, which controls the sensitivity of the system to changes in entropy.
6. The method for generating unsafe behavior risk profiles based on multimodal data fusion and dynamic visualization according to claim 3 is characterized in that: The inter-modal difference convergence reference value and information density entropy reference value after comprehensive analysis are input into the pre-trained machine learning model, and the information value density coefficient is generated by the machine learning model. The information value density coefficient is used to determine whether the current data source is in a high-frequency and low-value redundant state.
7. The method for generating unsafe behavior risk profiles based on multimodal data fusion and dynamic visualization according to claim 6 is characterized in that: The information value density change coefficient generated by the pre-trained machine learning model when performing intelligent evaluation of the information value density of the current data source is compared with the pre-set information value density change coefficient reference threshold to determine whether the current data source is in a high-frequency, low-value redundant state. The judgment steps are as follows: If the information value density change coefficient is less than the pre-set information value density change coefficient reference threshold, it is judged that the current data source data is in a high-frequency, low-value redundant state; if the information value density change coefficient is greater than or equal to the pre-set information value density change coefficient reference threshold, it is judged that the current data source data is not in a high-frequency, low-value redundant state.
Citation Information
Patent Citations
Safety early warning system based on multi-modal data fusion
CN116881850A
Multi-modal data fusion control optimization system and method based on information bottleneck
CN117390432A