Unsafe behavior risk portrait generation method based on multi-modal data fusion and dynamic visualization
By real-time perception and dynamically regulating the value of data source information, identifying and weakening high-frequency and low-value redundant data, dynamically adjusting the fusion contribution, solving the problem of redundant data diluting key signals in multimodal data fusion, and achieving more accurate risk image generation and security monitoring.
Patent Information
- Application Number
- CN202510472639.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-04-16
AI Technical Summary
In the multimodal data fusion, high-frequency and low-value redundant data dilutes key risk signals, resulting in the system being unable to identify high-risk behaviors, reducing the accuracy of safety monitoring and accident warning.
Through real-time perception and dynamic regulation of the data source information value, identify high-frequency and low-value redundant data, dynamically adjust the fusion contribution, use machine learning models to evaluate the data source status, weaken redundant data interference and continuously monitor data changes, and achieve adaptive recovery.
It improves the accuracy and dynamic nature of risk portraits, can more effectively identify high-risk behaviors, and improves the response ability and accident warning level of the safety monitoring system.
Smart Images

Figure CN120354359A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of safety management and risk assessment, and particularly to a method for generating a risk portrait of unsafe behaviors based on multimodal data fusion and dynamic visualization. Background Art
[0002] The generation of a risk portrait of unsafe behaviors based on multimodal data fusion and dynamic visualization refers to integrating heterogeneous information from multiple data sources (such as video surveillance, sensor data, behavior logs, environmental information, etc.), applying multimodal data fusion technology to comprehensively process, correlate, and deeply mine the behavior data of personnel or equipment to identify potential unsafe behaviors and their accompanying risk hazards. Subsequently, with the help of dynamic visualization technology, the analysis results are presented in an intuitive and easy-to-understand form, such as a risk heat map, time series analysis chart, three-dimensional scene restoration, behavior trajectory chart, etc., and then a dynamically evolving risk portrait is constructed. This portrait can not only reflect the current safety status in real time but also be dynamically updated according to new data, assisting safety management personnel in achieving real-time monitoring, intelligent early warning, and scientific decision-making, thereby optimizing management strategies, improving safety levels, and reducing the accident incidence rate.
[0003] The prior art has the following deficiencies: The prior art usually evaluates the credibility of different data sources by analyzing their historical performance (such as recognition accuracy, data loss rate, consistency with other modalities, etc.) and assigns different fusion contribution degrees accordingly. However, when a certain data source shows high-frequency and low-value redundancy, that is, it frequently uploads information but contains almost no effective content (such as repeated background images, low-difference logs), it will continuously dilute the high-value information contained in other modalities during the data fusion process, thus masking key risk signals and reducing the overall recognition accuracy. If the system cannot adaptively adjust the fusion contribution degree of this data source, the high-frequency redundant data will occupy a large amount of "attention" resources, resulting in marginalization or even complete neglect of those truly abnormal but low-frequency key signals (such as sudden equipment vibration, dangerous actions, environmental mutations). Ultimately, the system may fail to identify high-risk behaviors (such as personnel straying into dangerous areas or equipment operating overloaded) at critical moments, resulting in no early warning and response during accidents, seriously threatening personnel safety and the stable operation of the system.
[0004] The above information disclosed in the background art section is only used to enhance the understanding of the background of the present disclosure, and thus it may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0005] The object of the present invention is to provide a method for generating an unsafe behavior risk portrait based on multi-modal data fusion and dynamic visualization. By real-time sensing and dynamically regulating the information value of data sources, high-frequency low-value redundant data can be effectively identified, redundant modal interference with the fusion result can be avoided, and the dominant position of key modal information in the risk portrait can be enhanced. The dynamic adjustment and adaptive recovery mechanism of the fusion contribution degree enhance the system's response ability to data changes, making the generated unsafe behavior risk portrait more accurate, dynamic, and targeted, thereby improving the accuracy and reliability of overall safety monitoring and accident warning to solve the problems in the above background technology.
[0006] To achieve the above object, the present invention provides the following technical solution: A method for generating an unsafe behavior risk portrait based on multi-modal data fusion and dynamic visualization, comprising the following steps: Through network communication, continuously and real-time collect the original data from different modal data sources to ensure the freshness and integrity of the information relied on by the overall fusion process; After preprocessing the obtained original data, organize the standardized data source data into a data set; Extract the key indicators reflecting the high-frequency low-value redundancy of the data source from the data set, and comprehensively analyze the extracted key indicators to characterize the information value density of the data obtained by the data source; Input the key indicators after comprehensive analysis into a pre-trained machine learning model, and use the machine learning model to intelligently evaluate the information value density of the current data source to determine whether the current data source is in a high-frequency low-value redundant state; When the machine learning model determines that the current data source enters a high-frequency low-value redundant state, based on the evaluation result of the machine learning model, reduce the fusion contribution degree of the current data source in the multi-modal fusion process, weaken the interference of the current data source on the overall judgment result, and continuously monitor the data change of the current data source. When the data source "value rebounds", perform adaptive recovery on the fusion contribution degree.
[0007] Preferably, continuously and real-time collect the original data from different modal data sources through network communication, including the following specific steps: First, perform unified access configuration on various data sources to ensure the identification of their communication protocols, data formats, and sampling frequencies; Second, establish a stable data transmission channel, and use wired networks and wireless communication methods to achieve real-time reception of data from each modality; Subsequently, set reasonable sampling strategies and bandwidth allocation mechanisms according to the characteristics of different data sources to avoid data congestion or delay; for resource-constrained terminal devices, deploy edge computing modules to perform preliminary data filtering and compression in advance to reduce the burden on the network and the central server; Finally, through the timestamp alignment and data caching mechanism, the real-time collected data stream is sent into the fusion process orderly and synchronously, ensuring that the subsequent analysis and processing are carried out based on the multi-modal information with freshness, continuity, and integrity.
[0008] Preferably, key indicators reflecting the high-frequency low-value redundancy of the data source are extracted from the data set, including the difference degree between modalities at the information layer and the change rate of information entropy per unit time. After comprehensively analyzing the change of the difference degree between modalities at the information layer and the change rate of information entropy per unit time under the detection window, a reference value for the convergence of differences between modalities and a reference value for information density entropy are generated respectively. The information value density of the data obtained from the current data source is characterized jointly by the reference value for the convergence of differences between modalities and the reference value for information density entropy.
[0009] Preferably, the specific steps for comprehensively analyzing the difference degree between modalities at the information layer under the detection window to generate a reference value for the convergence of differences between modalities are as follows: To characterize the similarity degree of different modalities at the information trend level, first, local extreme points within the current time period are extracted for each modality to construct a trend structure feature set and generate a trend similarity index. The calculation expression is as follows: , where is the trend similarity index, is the set of local extreme points of modality , is the set of local extreme points of modality , is the number of overlaps of modality and at the local extreme points, and are the total numbers of local extreme points of modality and modality respectively, is the minimum value of the number of local extreme points of modality and modality ; To further analyze the differences between modalities at the event recognition and response levels, an index measuring the independence of modal responses, i.e., the event response overlap index, is constructed. Suppose within the detection window, modality and modality respectively identify their own event response sets, and some events are jointly recognized by the two modalities. The generation formula for the event response overlap index is as follows: , where is the event response overlap index, is the number of times that modality and modality respond to the same type of event at similar time points, is a modality and the modality The number of parts in the set of events recognized by each that are not recognized by the other; Based on the trend similarity index and the event response overlap index generate a reference value for the convergence of differences between modalities. The calculation expression is as follows: , where is the reference value for the convergence of differences between modalities, is a very small positive number added to prevent the denominator from being zero.
[0010] Preferably, the specific steps for analyzing the change rate of information entropy per unit time under the detection window to generate the reference value of information density entropy are as follows: Slice the original data within a continuous time period according to a fixed data volume, calculate the Shannon information entropy for each segment of data, and then calculate the ratio of the change in information entropy between every two adjacent segments of data to capture the intensity and direction of the entropy value fluctuation. The calculation expression is as follows: , where is the change rate of information entropy, that is, the change rate of information entropy between the th segment of data and the th segment of data, is the Shannon information entropy of the th segment of data, is the Shannon information entropy of the th segment of data, is a normalization factor to avoid scale inconsistency problems caused by different entropy values, is a very small positive number added to prevent the denominator from being zero; Perform non-linear transformation and compression on all information entropy change rates to generate a reference value of information density entropy, which is used to overall characterize the information activity of the data within the current monitoring window. The calculation expression is as follows: , where is the reference value of information density entropy, is the natural base, is the sensitivity coefficient, which controls the sensitivity of the system to entropy changes.
[0011] Preferably, input the reference value for the convergence of differences between modalities and the reference value of information density entropy after comprehensive analysis into a pre-trained machine learning model, generate an information value density coefficient through the machine learning model, and determine whether the current data source is in a high-frequency low-value redundancy state through the information value density coefficient.
[0012] Preferably, the information value density change coefficient generated when the information value density of the current data source is intelligently evaluated by a pre-trained machine learning model is compared with a pre-set reference threshold of the information value density change coefficient to determine whether the current data source is in a high-frequency low-value redundant state. The determination steps are as follows: If the information value density change coefficient is less than the pre-set reference threshold of the information value density change coefficient, it is determined that the data of the current data source is in a high-frequency low-value redundant state; if the information value density change coefficient is greater than or equal to the pre-set reference threshold of the information value density change coefficient, it is determined that the data of the current data source is not in a high-frequency low-value redundant state.
[0013] Preferably, when the machine learning model determines that the current data source enters a high-frequency low-value redundant state, based on the evaluation result of the machine learning model, the fusion contribution degree of the current data source in the multi-modal fusion process is reduced, and at the same time, the data change of the current data source is continuously monitored. When the data source "value rebounds", the specific steps for adaptive recovery of the fusion contribution degree are as follows: When the machine learning model determines that the current data source enters a high-frequency low-value redundant state, according to the information value density change coefficient output by the model Compare with the set reference threshold, trigger the weight adjustment mechanism, and dynamically attenuate the fusion contribution degree of the current data source in the multi-modal fusion process. The calculation expression is as follows: , In the formula, is the effective fusion contribution degree after dynamic adjustment, is the original fusion contribution degree, is the attenuation rate factor, is the reference threshold of the information value density change coefficient, is the attenuation curve control factor; After continuously monitoring the data source, once the machine learning model re-determines that its information value density change coefficient is greater than or equal to the threshold , it indicates that the data source content changes actively and the information quality recovers. At this time, based on the adaptive recovery strategy, the fusion contribution degree of the current data source should be gradually increased. The formula is as follows: , In the formula, is the fusion weight after adaptive recovery, is the minimum weight after being weakened, is the recovery rate adjustment factor, is the recovery sensitivity control parameter.
[0014] In the above technical solution, the technical effects and advantages provided by the present invention are: The present invention realizes the real-time perception and dynamic regulation of the information value of data sources, significantly improving the quality of fused data and the accuracy of decision-making results. Specifically, by introducing an information value density evaluation and high-frequency low-value redundancy identification mechanism, this method avoids the problem of redundant modalities "dominating" in data fusion, ensuring the dominant position of truly critical and discriminative modality information in the generation of risk portraits. At the same time, the dynamic adjustment and adaptive recovery mechanism of the fusion contribution degree enable the system to respond in a timely manner according to changes in the data source state, maintaining the stability and robustness of the fusion process. The finally generated risk portrait of unsafe behaviors will be more accurate, dynamic, and targeted, capable of more effectively identifying high-risk behaviors, highlighting key risks, and providing more insightful auxiliary decision-making basis for safety management personnel, thereby enhancing the response ability of the overall safety monitoring system and the accident early warning level. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments recorded in the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings.
[0016] Figure 1 It is a flowchart of the method for generating a risk portrait of unsafe behaviors based on multi-modal data fusion and dynamic visualization according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0017] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these example embodiments are provided so that this disclosure will be more complete and comprehensive, and will fully convey the concept of the example embodiments to those skilled in the art.
[0018] The present invention provides a method for generating a risk portrait of unsafe behaviors based on multi-modal data fusion and dynamic visualization as shown in Figure 1 the following, including the following steps: Continuously and real-time collect raw data from different modal data sources (such as cameras, sensors, log systems, etc.) through network communication to ensure the freshness and integrity of the information relied on by the overall fusion process; Continuously and real-time collect the raw data from different modality data sources (such as cameras, sensors, log systems, etc.) through network communication, so as to ensure that the system can quickly capture the latest environmental or operation status and respond to potential changes in the first time. To achieve efficient collection, it is necessary to reasonably configure the sampling frequency and bandwidth to control the transmission cost while ensuring data integrity; for some front-end devices, edge computing strategies can also be adopted to filter redundant or invalid data in advance and only upload valuable information to the central system, thus reducing the network burden. The core significance of real-time collection lies in providing the most authentic and timely basic data support for subsequent data processing, analysis and judgment, and ensuring the efficiency and accuracy of the entire system's perception link.
[0019] Continuously and real-time collect the raw data from different modality data sources (such as cameras, sensors, log systems, etc.) through network communication, including the following specific steps: First, the system needs to perform unified access configuration on various data sources to ensure that it can identify their communication protocols, data formats and sampling frequencies; second, establish a stable data transmission channel and use wired networks, wireless communications or field buses, etc. to achieve real-time reception of data from each modality; then, set reasonable sampling strategies and bandwidth allocation mechanisms according to the characteristics of different data sources to avoid data congestion or delay; for resource-constrained terminal devices, edge computing modules can be deployed to perform preliminary data filtering and compression in advance to reduce the burden on the network and the central server; finally, through timestamp alignment and data caching mechanisms, the real-time collected data stream is sent into the fusion process orderly and synchronously to ensure that the subsequent analysis and processing links are carried out based on multi-modal information with freshness, continuity and integrity.
[0020] After preprocessing the obtained raw data, organize the normalized data source data into a data set; After preprocessing the obtained raw data, organizing the normalized data source data into a data set means that after completing preprocessing operations such as data cleaning, format conversion, and anomaly elimination, the system will structurally integrate various types of data (such as video frames, sensor readings, behavior logs, etc.) according to a unified standard to construct an ordered and manageable data set for subsequent metric extraction and model analysis. The specific steps include: First, perform time alignment on the data from each modality to ensure the comparability of multi-source data; second, slice and classify the data according to a certain time window or event label; then convert the data into a unified format (such as tensors, tables, sequences, etc.) and unify the encoding method; finally, organize and store the processed data according to meta-information labels such as source, time, and location to form a structured data set for subsequent unified analysis and invocation. The core purpose of this process is to lay a stable and standardized foundation for subsequent feature extraction and model evaluation.
[0021] Extract key indicators reflecting the high-frequency low-value redundancy of the data source from the data set, comprehensively analyze the extracted key indicators, and characterize the information value density of the data obtained by the data source; Extract key indicators reflecting the high-frequency low-value redundancy of the data source from the data set, including the difference degree between modalities at the information layer (such as event triggering, trend) and the change rate of information entropy per unit time. After comprehensively analyzing the change of the difference degree between modalities at the information layer (such as event triggering, trend) and the change rate of information entropy per unit time under the detection window, generate the reference value of inter-modal difference convergence and the reference value of information density entropy respectively, and jointly characterize the information value density of the data obtained by the current data source through the reference value of inter-modal difference convergence and the reference value of information density entropy.
[0022] Information value density refers to the ratio between the amount of effective information and the total amount of information contained within a unit of time or a unit of data volume, and is used to measure the richness of "useful content" contained in the data. It not only reflects dimensions such as data variability, information entropy, and event triggering frequency, but also focuses on whether these changes have practical significance and can provide support for decision-making. In multi-modal data processing or fusion scenarios, data sources with high information value density often possess stronger recognition capabilities, discrimination capabilities, and risk perception capabilities; conversely, if a data source continuously uploads repetitive, single, and unchanging content, even if it is updated frequently, its information value density may be very low, easily forming "high-frequency low-value redundancy" and interfering with the overall analysis results. Therefore, accurately evaluating information value density helps the system reasonably allocate resources, dynamically adjust data weights, and improve the accuracy of fusion quality and risk recognition.
[0023] When the difference degree between modalities at the information layer (such as event triggering frequency, trend change, etc.) gradually approaches 0, and the data change amplitude of the current modality itself is also extremely small, it is usually considered that the data source is in a state of high-frequency low-value redundancy. This phenomenon indicates that there is a lack of independence or complementarity between this modality and other modalities, and the information it provides no longer has differential value during the fusion process. At the same time, small self-change means that the data content is highly repetitive or lacks dynamics. Although the uploaded data is frequent, it actually does not introduce new information content, resulting in low information entropy and decreased value density. In multi-modal fusion scenarios, the system relies on each modality to provide complementary information to improve the overall recognition accuracy and judgment ability. Once a modality neither has internal information changes nor "converges" with other modalities, its data is very likely to become redundant input, instead diluting the signals of other high-value modalities and reducing the fusion efficiency and risk recognition accuracy. Therefore, the inter-modal difference degree approaching 0 and extremely small self-change are important signals for identifying high-frequency low-value redundancy, and the system should dynamically adjust its contribution to fusion.
[0024] The specific steps for comprehensively analyzing the difference degree between modalities at the information layer under the detection window to generate the inter-modal difference convergence reference value are as follows: To characterize the similarity degree of different modalities at the information trend level, first extract the local extreme points (such as inflection points, mutation points, trend reversal points) of each modality within the current time period, construct a trend structure feature set, generate a trend similarity index, and the calculation expression is as follows: , where is the trend similarity index, which is used to measure the synchronicity of the trend changes between two modalities. If is close to 1, it indicates that the two modalities have the same trend changes at similar time points, showing that they are highly synchronous and there may be information redundancy; if this index is low (close to 0), it means that the trend changes of the two modalities are independent, and they may provide complementary information rather than redundant information; is the set of local extreme points of modality , is the set of local extreme points of modality , is the number of coincidences of modality and at the local extreme points, and are the total numbers of local extreme points of modality and modality respectively, is the minimum value of the numbers of local extreme points of modality and modality ; The trend similarity index reflects the similarity degree of the trend change structures between two modalities. The closer the index is to 1, the more consistent the trend changes of the two are, and the higher the degree of trend redundancy.
[0025] To further analyze the differences between modalities at the event recognition and response levels, construct an index to measure the independence of modal responses, that is, the event response overlap index. Suppose within the detection window, modality and modality respectively identify their own event response sets, and some events are jointly recognized by the two modalities. The generation formula of the event response overlap index is as follows: , where is the event response overlap index, is the number of times that modality and modality both respond to the same type of event at similar time points, is the number of parts in the event sets respectively identified by modality and modality that are not recognized by the other party; The event response overlap index is used to measure whether the responses of two modalities to risk events are highly coincident. If the event response overlap index approaches 1, it indicates that the two modalities converge at the behavioral response level and have strong information redundancy.
[0026] Based on the trend similarity index and the event response overlap index , a reference value for the convergence of differences between modalities is generated, and the calculation expression is as follows: , where, is the reference value for the convergence of differences between modalities, is a very small positive number (such as 1e-6) added to prevent the denominator from being zero, enhancing stability.
[0027] The reference value for the convergence of differences between modalities comprehensively measures the redundancy of the two modalities at the trend and event response levels. The smaller the value, the higher the degree of information convergence between the two, and the more likely the current modality is in a high-frequency and low-value redundancy state; the larger the value, the more it indicates their difference and complementarity, and the higher the information value density.
[0028] The smaller the reference value for the convergence of differences between modalities generated by comprehensively analyzing the difference degree between modalities at the information layer under the detection window, the more likely the current data source is in a high-frequency and low-value redundancy state. This reference value reflects the difference degree between the current modality and other modalities at the information layer (such as event trigger frequency, trend change, etc.). If this reference value approaches 0, it means that the information provided by this modality is highly similar or repetitive to other modalities in terms of content, trend, and structure, lacking difference and independence, and unable to bring effective information increment to the fusion system. At this time, although this data source may upload data frequently, the content is highly repetitive or redundant and coincident with other modalities, which is a typical manifestation of high-frequency and low-value redundancy. On the contrary, if the reference value for the convergence of differences between modalities is relatively high, it means that the information provided by this modality has uniqueness and supplementary value in semantics and dynamic changes, which helps to enhance the multi-dimensional perception ability of the fusion system, so it does not belong to a redundant modality.
[0029] If the information entropy remains at a low level over a long period of time per unit time, it usually indicates that the data of the current data source is in a state of high-frequency low-value redundancy. Information entropy is an important indicator for measuring data uncertainty and information diversity. The higher the entropy value, the more diverse the data content and the more information it contains; conversely, the lower the entropy value, the higher the data repetition rate, the less change, and the more single the content. When a data source continuously uploads information but the information entropy always remains at a low level, it means that although a large amount of data is generated per unit time, no new and effective information is provided, showing "active in form but poor in content". During the multi-modal fusion process, the redundant signals of such data sources may mask key and sudden high-value information in other modalities, resulting in misjudgment or missed judgment when the system identifies risk events. Therefore, the continuous low entropy value per unit time is an important signal for identifying that the data source has entered the state of high-frequency low-value redundancy, which has direct significance for information value density evaluation and fusion weight regulation.
[0030] The specific steps for analyzing the change rate of information entropy per unit time under the detection window to generate the information density entropy reference value are as follows: Slice the original data within a continuous time period according to a fixed data volume (such as every 100 data as a segment), calculate the Shannon information entropy for each segment of data, and then calculate the ratio of the information entropy change between every two adjacent segments of data to capture the intensity and direction of the entropy value fluctuation. The calculation expression is as follows: , where is the information entropy change rate, that is, the information entropy change rate between the th segment of data and the th segment of data, is the Shannon information entropy of the th segment of data, is the Shannon information entropy of the th segment of data, is the normalization factor to avoid the scale inconsistency problem caused by different entropy values, is a very small positive number (such as 1e-6) added to prevent the denominator from being zero, enhancing stability; The above steps measure the local entropy dynamics through proportional changes, and can more sensitively identify the differences between "static redundancy" and "mutant high values".
[0031] Perform non-linear transformation and compression on all information entropy change rates to generate the information density entropy reference value, which is used to overall characterize the information activity of the data within the current monitoring window. The calculation expression is as follows: , where is the information density entropy reference value, is the natural base, is the sensitivity coefficient, which controls the sensitivity of the system to entropy changes, equivalent to the "magnifying glass" adjustment knob.
[0032] By performing exponential compression and weighted aggregation on the change rate of information entropy, a reference value of information density entropy is generated to comprehensively characterize the information activity and value density of the current data source within the monitoring window.
[0033] The smaller the reference value of information density entropy generated by analyzing the change rate of information entropy per unit time under the detection window, the more limited the change range of information entropy within a certain period of time, indicating that the data content is stable, repetitive, lacking novelty and dynamic changes, belonging to the typical "high-frequency low-value redundancy" feature, that is, although the data is uploaded frequently, it lacks practical useful information; conversely, when the reference value of information density entropy is larger, it represents that there are significant fluctuations in information entropy per unit time, indicating that the data source continuously produces valuable content or undergoes state changes, and does not belong to the redundant state.
[0034] Input the key indicators after comprehensive analysis into a pre-trained machine learning model, and use the machine learning model to intelligently evaluate the information value density of the current data source to determine whether the current data source is in the high-frequency low-value redundancy state; Input the reference value of inter-modal difference convergence and the reference value of information density entropy after comprehensive analysis into a pre-trained machine learning model, generate an information value density coefficient through the machine learning model, and determine whether the current data source is in the high-frequency low-value redundancy state through the information value density coefficient.
[0035] The "pre-trained machine learning model" refers to the process of completing the modeling in advance using machine learning algorithms based on a large amount of existing sample data and expert annotation results before the formal operation of the system, so that the model has the ability to make intelligent judgments on unknown data. In this scenario, the training process of the model includes the following key steps: First, collect historical data from multiple data sources and extract key feature indicators such as the reference value of inter-modal difference convergence and the reference value of information density entropy; Second, construct a supervised training sample set according to the data source status (whether it is high-frequency low-value redundancy) annotated by humans or rule systems; Then, select appropriate machine learning algorithms (such as random forest, support vector machine, XGBoost, lightweight neural network, etc.), input these sample data into the model for training, and enable it to learn the corresponding relationship between different feature patterns and data states. After sufficient training and verification, the model can identify under what combination of features the data source is more likely to be in the "high-frequency low-value redundancy" state, thus having a certain generalization ability.
[0036] In the actual application stage, the model is no longer retrained but directly invoked and used as an intelligent judgment tool. During the current operation of the system, the system will calculate in real time the reference value of the convergence of the differences between modalities and the reference value of the information density entropy of each data source, and input these features into the trained model. The model will perform rapid inference and output a quantization coefficient (i.e., the information value density coefficient) representing the current "information value density" of the data source. This coefficient reflects the level of effective information content of the current data per unit time. If its value is lower than the preset threshold, it can be determined that the data source is in a high-frequency and low-value redundant state, and the system will dynamically adjust its fusion weight accordingly. Since the model is constructed based on prior knowledge and real cases, it can avoid the subjectivity and limitations brought by manually setting fixed thresholds, and can more flexibly and accurately adapt to data changes in different scenarios, improving the recognition efficiency and robustness of the entire system for redundant data sources. This approach embodies the integration idea of "data-driven + model-driven" and is the key means to achieve intelligent and automated data quality management.
[0037] The machine learning model is not limited here, and any machine learning model that can perform comprehensive analysis on the reference value of the convergence of the differences between modalities and the reference value of the information density entropy to generate the information value density change coefficient is acceptable. To implement the technical solution of the present invention, the present invention provides a specific implementation method; The formula for generating the information value density change coefficient is as follows: In the formula, and are the preset proportionality coefficients of the reference value of the convergence of the differences between modalities and the reference value of the information density entropy respectively, and and are both greater than 0.
[0038] The preset proportionality coefficients and refer to the weight factors given to the reference value of the convergence of the differences between modalities and the reference value of the information density entropy when calculating the information value density change coefficient , which are used to adjust their relative contributions to the final result. Since in different application scenarios, the influence of the reference value of the convergence of the differences between modalities and the reference value of the information density entropy on the information value may be different, it is necessary to set appropriate weights so that the calculation of the information value density change coefficient can more accurately reflect the information value density of the data source. and Determined through experience, data analysis, or machine learning model training to ensure the coefficient of variation of information value density maintains reasonable calculation results in different environments. Both are greater than 0, meaning the inter-modal difference convergence reference value and the information density entropy reference value both play a positive role in calculating the coefficient of variation of information value density and will not be completely ignored.
[0039] From the coefficient of variation of information value density, the smaller the inter-modal difference convergence reference value generated by comprehensively analyzing the difference degree between modalities at the information layer under the detection window, and the smaller the information density entropy reference value generated by analyzing the change rate of information entropy per unit time under the detection window, the smaller the coefficient of variation of information value density generated when the information value density of the current data source is intelligently evaluated by a pre-trained machine learning model, indicating a greater probability that the data of the current data source is in a high-frequency low-value redundant state. Conversely, it indicates a smaller probability that the data of the current data source is in a high-frequency low-value redundant state.
[0040] Compare and analyze the coefficient of variation of information value density generated when the information value density of the current data source is intelligently evaluated by a pre-trained machine learning model with the pre-set reference threshold of the coefficient of variation of information value density to determine whether the current data source is in a high-frequency low-value redundant state. The judgment steps are as follows: If the coefficient of variation of information value density is less than the pre-set reference threshold of the coefficient of variation of information value density, it is determined that the data of the current data source is in a high-frequency low-value redundant state; if the coefficient of variation of information value density is greater than or equal to the pre-set reference threshold of the coefficient of variation of information value density, it is determined that the data of the current data source is not in a high-frequency low-value redundant state.
[0041] When the machine learning model determines that the current data source enters a high-frequency low-value redundant state, based on the evaluation result of the machine learning model, reduce the fusion contribution degree of the current data source in the process of multi-modal fusion (weight assignment, attention mechanism, weighted average, etc.), weaken the interference of the current data source on the overall judgment result, and continuously monitor the data change of the current data source. When the "value of the data source rebounds", adaptively restore the fusion contribution degree; The core function of the above steps is to achieve intelligent dynamic regulation of high-frequency and low-value redundant data sources, ensuring data quality and the accuracy of fusion results during the multi-modal fusion process. When the machine learning model determines that a data source is in a high-frequency and low-value redundant state, it means that the data uploaded by this data source is frequent but has a low information value density during the current period, that is, its content lacks variability, has a low event trigger rate, or is highly repetitive. At this time, if a high fusion contribution degree is still maintained, it will cause the system to give unreasonable "trust" to it during the fusion stage, thus diluting the contributions of other high-value data sources, masking key risk signals, and even causing incorrect judgments. Therefore, based on the model evaluation results, the system dynamically reduces the proportion of this data source in fusion strategies such as weight allocation, attention mechanism, or feature weighting. Substantially, it weakens its influence on the final decision-making and enhances the robustness and recognition accuracy of the system.
[0042] Meanwhile, this step also undertakes the responsibility of maintaining the adaptive ability. The system does not simply "block" or "discard" redundant modalities, but continuously monitors the change in the information value density in its data stream and evaluates whether there are signs of "value recovery" (such as increased information entropy, increased event frequency, increased feature difference, etc.). Once the data source returns to an active and high-value state, the system will gradually restore its contribution degree in the fusion, realizing a closed-loop control mechanism of "dynamic suppression - continuous monitoring - intelligent recovery". This mechanism ensures that the system can not only respond quickly to redundant data but also retains the fusion space for it to regain value in the future, avoiding "one-size-fits-all" data elimination, thus achieving a multi-modal fusion ability with high flexibility, high precision, and high fault tolerance.
[0043] When the machine learning model determines that the current data source enters a high-frequency and low-value redundant state, based on the machine learning model evaluation results, the specific steps to reduce the fusion contribution degree of the current data source in the multi-modal fusion (weight allocation, attention mechanism, weighted average, etc.) process and continuously monitor the data change of the current data source and adaptively restore the fusion contribution degree after the data source "value recovers" are as follows: When the machine learning model determines that the current data source enters a high-frequency and low-value redundant state, according to the information value density change coefficient output by the model Compare it with the set reference threshold to trigger the weight adjustment mechanism and dynamically attenuate the fusion contribution degree of the current data source in the multi-modal fusion process. The calculation expression is as follows: , In the formula, is the effective fusion contribution degree after dynamic adjustment, is the original fusion contribution degree, is the attenuation rate factor, which controls the sensitivity of the downward adjustment of the fusion contribution degree. The larger the value, the faster the fusion contribution degree decreases. is the reference threshold of the information value density change coefficient, is the attenuation curve control factor, which is used to control the degree of non-linearity, represents accelerated downward adjustment, represents slow adjustment; This step uses the exponential attenuation mechanism to convert the gap with the threshold into a weight attenuation factor, enabling the system to quickly reduce the interference ability of redundant modes on the overall judgment when detecting redundant tendencies. This method is more flexible than linear downward adjustment, can avoid over-adjustment caused by small fluctuations, and at the same time significantly weakens severely redundant modes, effectively improving the perception ability of the fusion model for real key information.
[0044] After continuously monitoring this data source, once the machine learning model re-determines that the information value density change coefficient is greater than or equal to the threshold , it indicates that the content of the data source changes actively and the information quality recovers. At this time, based on the adaptive recovery strategy, the fusion contribution degree of the current data source should be gradually increased, and the formula is as follows: , In the formula, is the fusion weight after adaptive recovery, is the minimum weight after being weakened, is the recovery rate adjustment factor, and the larger the value, the faster the recovery, is the recovery sensitivity control parameter, which determines the non-linear influence of the information value density change coefficient on the recovery amplitude, represents that when the information value density change coefficient accelerates the recovery.
[0045] This step introduces a non-linear recovery mechanism to ensure that the fusion contribution degree of the data source is gradually restored only after the data source truly shows "substantial information recovery". Compared with directly restoring the original weight, this method is more robust, avoids system oscillations caused by short-term data fluctuations, and at the same time ensures the re-integration of valuable data sources, improving the comprehensiveness and dynamic adaptation ability of the fusion judgment.
[0046] The present invention realizes the real-time perception and dynamic regulation of the information value of data sources, significantly improving the quality of the fused data and the accuracy of the decision-making results. Specifically, by introducing an information value density evaluation and high-frequency low-value redundancy identification mechanism, this method avoids the problem of redundant modalities "dominating" in data fusion, ensuring the dominant position of truly critical and discriminative modality information in the generation of risk portraits. At the same time, the dynamic adjustment and adaptive recovery mechanism of the fusion contribution degree enable the system to respond in a timely manner according to changes in the data source state, maintaining the stability and robustness of the fusion process. The finally generated risk portrait of unsafe behaviors will be more accurate, dynamic, and targeted, capable of more effectively identifying high-risk behaviors, highlighting key risks, and providing more insightful auxiliary decision-making basis for safety management personnel, thereby enhancing the response ability of the overall safety monitoring system and the accident early warning level.
[0047] The above formulas are all dimensionless and take their numerical values for calculation. The formulas are obtained by collecting a large amount of data for software simulation to obtain a formula closest to the actual situation. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.
[0048] Only some exemplary embodiments of the present invention have been described by way of illustration. Undoubtedly, for those of ordinary skill in the art, without departing from the spirit and scope of the present invention, the described embodiments can be modified in various different ways. Therefore, the above drawings and descriptions are illustrative in nature and should not be construed as limiting the scope of the claims of the present invention.
[0049] It should be noted that in this article, if there are relational terms such as first and second, they are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of another identical element in the process, method, article or device comprising the element.
[0050] It should be understood that in various embodiments of the present application, the magnitudes of the serial numbers of the above processes do not mean the order of execution is prior or subsequent. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0051] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.
[0052] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.
[0053] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0054] In addition, the functional units in each embodiment of this application can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.
[0055] As described above, the above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in this application, and all should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
[0056] Only some exemplary embodiments of the present invention have been described above by way of illustration. Undoubtedly, for those of ordinary skill in the art, without departing from the spirit and scope of the present invention, the described embodiments can be modified in various different ways. Therefore, the above drawings and description are illustrative in nature and should not be construed as limiting the protection scope of the claims of the present invention.
Claims
1. A method for generating an unsafe behavior risk profile based on multimodal data fusion and dynamic visualization, characterized in that, Including the following steps: Through network communication, continuously and real-time collect the raw data from different modal data sources to ensure the freshness and integrity of the information relied on by the overall fusion process; After preprocessing the obtained raw data, organize the normalized data source data into a data set; Extract the key indicators reflecting the high-frequency low-value redundancy of the data source from the data set, conduct a comprehensive analysis of the extracted key indicators, and characterize the information value density of the data obtained by the data source; Input the key indicators after comprehensive analysis into a pre-trained machine learning model, and use the machine learning model to intelligently evaluate the information value density of the current data source to determine whether the current data source is in a high-frequency low-value redundancy state; When the machine learning model determines that the current data source enters the high-frequency low-value redundancy state, based on the evaluation result of the machine learning model, reduce the fusion contribution degree of the current data source in the multi-modal fusion process, weaken the interference of the current data source to the overall judgment result, and at the same time continuously monitor the data change of the current data source, and adaptively restore the fusion contribution degree when the "value of the data source rebounds".
2. The method for generating an unsafe behavior risk portrait based on multi-modal data fusion and dynamic visualization according to claim 1, characterized in that Continuously and real-time collect the raw data from different modal data sources through network communication, including the following specific steps: First, perform unified access configuration on various data sources to ensure the identification of their communication protocols, data formats, and sampling frequencies; Second, establish a stable data transmission channel, and use wired networks and wireless communication methods to achieve real-time reception of data from each modality; Subsequently, set reasonable sampling strategies and bandwidth allocation mechanisms according to the characteristics of different data sources to avoid data congestion or delay; for resource-constrained terminal devices, deploy edge computing modules to perform preliminary data filtering and compression in advance to reduce the burden on the network and the central server; Finally, through timestamp alignment and data caching mechanisms, orderly and synchronously send the real-time collected data stream into the fusion process to ensure that subsequent analysis and processing links are carried out based on multi-modal information with freshness, continuity, and integrity.
3. The method for generating an unsafe behavior risk portrait based on multimodal data fusion and dynamic visualization according to claim 1, wherein Extract the key indicators reflecting the high-frequency low-value redundancy of the data source from the data set, including the difference degree between modalities at the information layer and the change rate of information entropy per unit time. After comprehensively analyzing the change of the difference degree between modalities at the information layer and the change rate of information entropy per unit time under the detection window, generate a reference value for the convergence of differences between modalities and a reference value for the information density entropy respectively, and jointly characterize the information value density of the data obtained by the current data source through the reference value for the convergence of differences between modalities and the reference value for the information density entropy.
4. The method for generating an unsafe behavior risk portrait based on multimodal data fusion and dynamic visualization according to claim 3, wherein The specific steps for comprehensively analyzing the difference degree between modalities at the information layer under the detection window to generate a reference value for the convergence of differences between modalities are as follows: To characterize the similarity degree of different modalities at the information trend level, first, extract the local extreme points of each modality within the current time period, construct a trend structure feature set, and generate a trend similarity index. The calculation formula is as follows: , where is the trend similarity index, is the set of local extreme points of modality , is the set of local extreme points of modality , is the number of coincidences of modality and at the local extreme points, and are the total numbers of local extreme points of modality and modality respectively, is the minimum value of the numbers of local extreme points of modality and modality ; To further analyze the differences between modalities at the event recognition and response levels, an index for measuring the independence of modal responses, namely the event response overlap index, is constructed. Suppose within the detection window, modality and modality respectively identify their own event response sets, and some events are jointly recognized by both modalities. The formula for generating the event response overlap index is as follows: , where is the event response overlap index, is the number of times that modality and modality both respond to the same type of event at similar time points, is the number of parts in the event sets respectively identified by modality and modality that are not recognized by the other party; Based on the trend similarity index and the event response overlap index , a modal difference convergence reference value is generated, and the calculation expression is as follows: , where is the modal difference convergence reference value, is a very small positive number added to prevent the denominator from being zero.
5. The method for generating an unsafe behavior risk portrait based on multi-modal data fusion and dynamic visualization according to claim 3, characterized in that The specific steps for analyzing the change rate of information entropy per unit time under the detection window to generate a reference value for the information density entropy are as follows: Slice the original data within consecutive time periods into fixed data volumes, calculate the Shannon information entropy for each segment of data, and then perform a ratio calculation on the change in information entropy between every two adjacent segments of data to capture the intensity and direction of entropy value fluctuations. The calculation expression is as follows: , where is the information entropy change rate, that is, the information entropy change rate between the th segment of data and the th segment of data, is the Shannon information entropy of the th segment of data, is the Shannon information entropy of the th segment of data, is the normalization factor to avoid the scale inconsistency problem caused by different entropy value magnitudes, is a very small positive number added to prevent the denominator from being zero; All information entropy change rates are non-linearly transformed and compressed to generate an information density entropy reference value, which is used to comprehensively characterize the information activity of the data within the current monitoring window. The calculation expression is as follows: , where is the information density entropy reference value, is the natural base, is the sensitivity coefficient, which controls the sensitivity of the system to entropy changes.
6. The method for generating an unsafe behavior risk portrait based on multi-modal data fusion and dynamic visualization according to claim 3, wherein Input the modal difference convergence reference value and information density entropy reference value after comprehensive analysis into a pre-trained machine learning model. Generate an information value density coefficient through the machine learning model, and determine whether the current data source is in a high-frequency low-value redundancy state based on the information value density coefficient.
7. The method for generating an unsafe behavior risk portrait based on multi-modal data fusion and dynamic visualization according to claim 6, wherein Compare and analyze the information value density change coefficient generated when the pre-trained machine learning model intelligently evaluates the information value density of the current data source with the pre-set reference threshold of the information value density change coefficient to determine whether the current data source is in a high-frequency low-value redundancy state. The judgment steps are as follows: If the information value density change coefficient is less than the pre-set reference threshold of the information value density change coefficient, it is determined that the data of the current data source is in a high-frequency low-value redundancy state; if the information value density change coefficient is greater than or equal to the pre-set reference threshold of the information value density change coefficient, it is determined that the data of the current data source is not in a high-frequency low-value redundancy state.
8. The method for generating an unsafe behavior risk portrait based on multimodal data fusion and dynamic visualization according to claim 7, wherein When the machine learning model determines that the current data source enters a high-frequency low-value redundancy state, based on the evaluation result of the machine learning model, reduce the fusion contribution degree of the current data source in the multi-modal fusion process, and continuously monitor the data change of the current data source. The specific steps for adaptive recovery of the fusion contribution degree when the "value of the data source rebounds" are as follows: When the machine learning model determines that the current data source enters the high-frequency and low-value redundancy state, according to the information value density change coefficient output by the model Compare with the set reference threshold, trigger the weight adjustment mechanism, and dynamically attenuate the fusion contribution degree of the current data source in the multi-modal fusion process. The calculation expression is as follows: , where is the effective fusion contribution degree after dynamic adjustment, is the original fusion contribution degree, is the attenuation rate factor, is the reference threshold of the information value density change coefficient, is the attenuation curve control factor; After continuously monitoring the data source, once the machine learning model re-determines the coefficient of change in its information value density is greater than or equal to the threshold , it indicates that the content of the data source changes actively and the information quality is restored. At this time, based on the adaptive recovery strategy, the integration contribution degree of the current data source should be gradually increased, and the formula is as follows: , wherein, is the fused weight after adaptive restoration, is the minimum weight after being weakened, is the restoration rate adjustment factor, is the restoration sensitivity control parameter.
Citation Information
Patent Citations
Safety early warning system based on multi-modal data fusion
CN116881850A
Multi-modal data fusion control optimization system and method based on information bottleneck
CN117390432A
A security risk analysis system and method based on multimodal data processing
CN119784145A
Flight pushback state monitoring method based on multi-modal data fusion
US20220402626A1
Cited By
Sky-ground comprehensive operation inspection method and system based on multi-source cooperation
CN120974402A