A method for automatically monitoring construction operation violations of infrastructure construction by multi-modal data fusion

By combining multimodal data fusion and dynamic weight allocation technology with environmental adaptive perception and multi-source data conflict resolution, the problem of inconsistent violation judgment in infrastructure construction safety monitoring systems under complex environments has been solved, achieving efficient and reliable monitoring and judgment of violations.

CN120974243BActive Publication Date: 2026-01-23BEIJING HUALIAN POWER ENG SUPERVISION CO +2
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511501296.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-21
Publication Date
2026-01-23
Estimated Expiration
2045-10-21

AI Technical Summary

Technical Problem

Existing infrastructure construction safety monitoring systems suffer from reduced detection performance in low light or high noise environments, low efficiency in utilizing multi-source data, and a lack of effective data arbitration and fusion mechanisms, leading to inconsistent conclusions in violation determinations.

Method used

By employing multimodal data fusion and dynamic weight allocation techniques, combined with environmental adaptive perception and multi-source data conflict resolution, the system acquires multimodal source data, performs standardized processing to generate preprocessed feature data, and dynamically allocates weights based on environmental state parameters to judge and resolve multi-source data conflicts. Finally, it determines violations based on a behavioral rule pattern library.

Benefits of technology

It improves the accuracy and robustness of monitoring violations in complex environments, reduces the probability of false alarms and missed alarms, realizes intelligent supervision of the entire construction process, and generates consistent and reliable violation judgment results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120974243B_ABST
    Figure CN120974243B_ABST
Patent Text Reader

Abstract

The application discloses a kind of multi-modal data fusion's infrastructure construction operation violation automatic monitoring method, belong to graphics data processing and pattern recognition technical field, it includes obtaining the multi-modal source data of target scene, data recognition and standardization processing are carried out to multi-modal source data, generate pre-processing characteristic data;Environment condition is obtained in situ, and environment state parameter is generated;Based on environment state parameter, pre-processing characteristic data is dynamically weighted and distributed, and weighted characteristic data is generated;Multi-source data conflict is judged and is eliminated to weighted characteristic data, and conflict elimination event data is generated;Based on the behavior rule mode library of pre-established, conflict elimination event data is matched with behavior mode, and violation determination result is generated.The application adopts multi-modal space-time feature fusion and dynamic weight distribution mechanism, combines environment adaptive perception and multi-source data conflict elimination technology, can realize the automatic monitoring and pattern recognition of violation behavior under construction environment.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of pattern recognition and image data processing, and particularly relates to a method for automatically monitoring construction work violations based on multi-modal data fusion. BACKGROUND

[0002] In the field of construction safety monitoring, multi-modal data fusion technology has become a key means to improve the level of automated supervision. This technology involves the collection, processing and analysis of source data from multiple modalities such as video, audio and sensors, aiming to achieve intelligent identification and early warning of construction violations through image data processing and pattern recognition methods.

[0003] In the prior art, common monitoring solutions mainly rely on single type of sensor data or simple multi-source data juxtaposition. For example, some systems use fixed position cameras for video monitoring, identifying safety helmet wearing or personnel activity areas through image processing algorithms. Other solutions attempt to combine audio sensors and vibration sensors to detect abnormal sound and equipment operating status respectively, and then aggregate the results after independent judgment. These methods usually use static thresholds or pre-set rules to analyze various types of data.

[0004] However, the existing technical solutions have obvious defects. Systems based on a single data source are susceptible to environmental conditions, such as significant performance degradation in low light or high noise environments. In addition, multi-source data is often simply superimposed, lacking a dynamic weight adjustment mechanism based on environmental state, resulting in low data utilization efficiency. At the same time, when different sources of data conflict in judgment, existing systems lack effective arbitration and fusion mechanisms, making it difficult to generate consistent and reliable violation determination conclusions. SUMMARY

[0005] To solve the above problems, the present application provides a method for automatically monitoring construction work violations based on multi-modal data fusion, which adopts multi-modal spatio-temporal feature fusion and dynamic weight distribution mechanism, combines environmental self-adaptive perception and multi-source data conflict resolution technology, and can realize automatic monitoring and pattern recognition of violations in construction environment.

[0006] The above objectives can be achieved through the following solutions:

[0007] The application discloses a kind of multi-modal data fusion's infrastructure construction operation violation automatic monitoring method, including obtaining the multi-modal source data of target scene, to multi-modal source data is carried out data identification and standardization processing, generates pre-processing feature data;Environment condition is obtained in site, generates environment state parameter;Based on environment state parameter, pre-processing feature data is carried out dynamic weight distribution, generates weighted feature data;To weighted feature data is carried out multi-source data conflict judgment and resolution, generates conflict resolution event data;Based on the behavior rule mode library of preestablished, conflict resolution event data is carried out behavior mode matching, generates violation determination result.

[0008] Optionally, the multi-modal source data is carried out data identification and standardization processing, and the pre-processing feature data includes: the multi-modal source data is carried out key frame extraction and target area positioning, and video feature data is generated;The multi-modal source data is carried out noise reduction processing and acoustic feature extraction, and audio feature data is generated;The multi-modal source data is carried out frequency domain transformation and feature value extraction, and sensor feature data is generated;The video feature data, audio feature data and the sensor feature data are combined, and pre-processing feature data is generated.

[0009] Optionally, the pre-processing feature data is carried out dynamic weight distribution based on the environment state parameter, and the weighted feature data includes: the light intensity and noise level of the environment state parameter are obtained;According to the light intensity, the weight of the video feature data is adjusted, and according to the noise level, the weight of the audio feature data is adjusted, and dynamic weight coefficient is generated;The pre-processing feature data is weighted using the dynamic weight coefficient, and the weighted feature data is generated.

[0010] Optionally, the multi-source data conflict judgment and resolution of the weighted feature data includes: based on the weighted feature data, preliminary judgment conclusion is generated;The logical conflict between the preliminary judgment conclusion is detected, and conflict identifier is formed;When the conflict identifier is received, the running state data of construction equipment is obtained;The preliminary judgment conclusion is arbitrated and verified using the running state data, and conflict resolution event data is generated.

[0011] Optionally, the behavior mode matching of the conflict resolution event data based on the behavior rule mode library includes: the conflict resolution event data is time-series combined, and real-time behavior sequence is constructed;Compliance mode is extracted from the behavior rule mode library, and the behavior rule mode library includes typical violation action mode template based on image sequence identification;The real-time behavior sequence is matched with the compliance mode, and when the real-time behavior sequence deviates from the compliance mode, violation determination result is generated.

[0012] Optionally, generating a preliminary judgment conclusion based on the weighted feature data includes: performing multimodal spatiotemporal feature map convolution processing on the weighted feature data to obtain a structured event representation tensor; and performing adaptive semantic inference and conflict probability assessment on the structured event representation tensor to generate a preliminary judgment conclusion.

[0013] Optionally, generating environmental state parameters based on the environmental conditions includes: obtaining quantitative environmental indicators using the environmental conditions; and performing spatial state mapping on the quantitative environmental indicators to generate environmental state parameters.

[0014] Optionally, obtaining the quantified environmental index using the environmental conditions includes: performing multi-source signal acquisition and analog-to-digital conversion on the environmental conditions to obtain a digital environmental signal; and performing filtering, noise reduction, and eigenvalue calculation on the digital environmental signal to obtain the quantified environmental index.

[0015] Optionally, the step of combining the conflict resolution event data in a time series to construct a real-time behavior sequence includes: injecting the conflict resolution event data into spatiotemporal anchors to obtain a time-stamped monitoring event dataset; and performing construction behavior concept mapping and dimensional compression on the time-stamped monitoring event dataset to construct a real-time behavior sequence.

[0016] Based on the same inventive concept, this invention also provides an automatic monitoring system for violations in infrastructure construction operations using multimodal data fusion. The system includes: a data acquisition and processing module for acquiring multimodal source data from a target scene, performing data identification and standardization on the multimodal source data, and generating preprocessed feature data; an environmental perception module for acquiring on-site environmental conditions and generating environmental state parameters based on these conditions; a weight allocation module for dynamically allocating weights to the preprocessed feature data based on the environmental state parameters, generating weighted feature data; a conflict resolution module for judging and resolving multi-source data conflicts in the weighted feature data, generating conflict resolution event data; and a behavior pattern analysis module for matching behavior patterns in the conflict resolution event data based on a preset behavior rule pattern library, generating violation judgment results.

[0017] Compared with the prior art, the present invention has the following advantages:

[0018] This invention improves the accuracy and robustness of violation monitoring in complex environments through multimodal data fusion and dynamic weight allocation technology. The system can adaptively adjust the contribution of different modalities such as video and audio data, effectively overcoming the shortcomings of a single sensor in terms of reduced sensing capability under harsh conditions such as insufficient lighting or noise interference, thereby ensuring the continuous generation of reliable feature data in real and ever-changing infrastructure scenarios.

[0019] This invention helps reduce the probability of false alarms and missed alarms by establishing a multi-source data conflict judgment and arbitration verification mechanism. When there are logical conflicts in the preliminary judgment conclusions generated by different sensors, the system can introduce third-party data such as equipment operating status for cross-verification and intelligent arbitration, and finally output consistent event data after resolution, thereby improving the consistency of monitoring conclusions.

[0020] This invention utilizes a compliance pattern library based on image sequences to determine behavioral patterns, achieving intelligent supervision of the entire construction process. The system efficiently matches real-time behavioral sequences with templates of typical violations, enabling it to identify not only static violations but also dynamic violations in the work process.

[0021] This invention standardizes and modularizes the data processing workflow, forming a closed-loop operation from feature extraction and weight allocation to conflict resolution, thereby improving the system's engineering practicality and computational efficiency. Through temporal combination and dimensionality compression techniques, the system can transform multimodal data into concise real-time behavioral sequences, ultimately outputting intuitive and clear violation judgment results, providing an efficient technical tool for on-site safety management.

[0022] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures pointed out in the description, claims and drawings. Attached Figure Description

[0023] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0024] Figure 1 This is a flowchart illustrating an automatic monitoring method for violations in infrastructure construction operations based on multimodal data fusion, according to an embodiment of the present invention.

[0025] Figure 2 This is a video feature influence diagram of light intensity according to an embodiment of the present invention.

[0026] Figure 3 This is a spatiotemporal feature map of an embodiment of the present invention.

[0027] Figure 4 This is a schematic diagram of the structure of an automatic monitoring system for violations in infrastructure construction operations based on multimodal data fusion, according to an embodiment of the present invention. Detailed Implementation

[0028] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0029] Reference Figure 1 One embodiment of the present invention proposes an automatic monitoring method for violations in infrastructure construction operations based on multimodal data fusion. It adopts a multimodal spatiotemporal feature fusion and dynamic weight allocation mechanism, combined with environmental adaptive perception and multi-source data conflict resolution technology, which can realize automatic monitoring and pattern recognition of violations in the construction environment.

[0030] The method described in this embodiment specifically includes:

[0031] Acquire multimodal source data of the target scene, perform data identification and standardization processing on the multimodal source data, and generate preprocessed feature data;

[0032] Obtain the environmental conditions at the site, and generate environmental state parameters based on the environmental conditions;

[0033] Based on the environmental state parameters, the preprocessed feature data is dynamically weighted to generate weighted feature data;

[0034] The weighted feature data is subjected to multi-source data conflict judgment and resolution to generate conflict resolution event data;

[0035] Based on a preset behavioral rule pattern library, behavioral pattern matching is performed on the conflict resolution event data to generate violation judgment results.

[0036] Specifically, the method first transforms complex, heterogeneous visual and auditory information from the field into unified feature data through multimodal data acquisition and standardized preprocessing. The core innovation lies in the introduction of an environmental perception mechanism. This mechanism generates environmental state parameters by acquiring real-time environmental conditions and uses these parameters to dynamically evaluate and weight the reliability of each modal feature data. Subsequently, the method designs a conflict judgment and resolution logic to perform internal consistency checks on the weighted and fused data, proactively identifying and resolving contradictions between different information sources, thereby generating highly reliable event data that has undergone cross-validation. Finally, these discrete and reliable event data are placed in a behavioral rule pattern library for temporal context comparison. By determining whether real-time behavioral sequences deviate from standard operating procedures, the method achieves accurate judgment from low-level data perception to high-level behavioral logic violations.

[0037] Optionally, the step of performing data identification and standardization processing on the multimodal source data to generate preprocessed feature data includes:

[0038] Keyframe extraction and target region localization are performed on the multimodal source data to generate video feature data;

[0039] The multimodal source data is subjected to noise reduction and acoustic feature extraction to generate audio feature data;

[0040] The multimodal source data is subjected to frequency domain transformation and feature value extraction to generate sensor feature data;

[0041] The video feature data, audio feature data, and sensor feature data are combined to generate preprocessed feature data.

[0042] Specifically, the process begins with parallel processing of data streams from different modalities. For video data streams, keyframe extraction and target region localization are performed. Keyframe extraction aims to filter static images from continuous video streams that represent significant changes in scene content, reducing data redundancy and capturing core actions. Subsequently, target region localization techniques, such as object detection algorithms in deep learning, are used to accurately identify and box key targets such as construction workers and machinery, thus focusing the analysis on effective areas and ultimately generating video feature data containing information such as target location and posture. For audio data streams, noise reduction and acoustic feature extraction are performed. Noise reduction aims to filter out background noise common at construction sites, such as wind noise and equipment noise, improving the signal-to-noise ratio of speech or key event sounds. Next, acoustic features that characterize the sound content are extracted, such as Mel-frequency cepstral coefficients, which effectively describe the spectral characteristics of speech and sound effects, thereby generating audio feature data. For sensor data streams, frequency domain transformation and feature value extraction are performed. Sensor data is typically time-series signals, such as acceleration and angular velocity. Transforming these signals from the time domain to the frequency domain using techniques like Fast Fourier Transform (FFT) reveals their inherent periodicity and vibration modes. Then, key characteristic values ​​such as peak frequency and energy distribution are calculated from the frequency domain signal to generate sensor feature data. Finally, the independently extracted video, audio, and sensor feature data are combined to form a unified, high-dimensional preprocessed feature data vector. This combination process can be represented by the following equation:

[0043] ,

[0044] in This represents the final generated preprocessed feature data vector. This represents the video feature data obtained and encoded from the video. This represents acoustic feature data obtained from audio. This represents the feature value data obtained from sensor data. The formula indicates that by concatenating feature vectors from different modalities, a standardized data structure capable of comprehensively describing the instantaneous state of the construction site is created.

[0045] For example, in a scenario monitoring tower crane operations, the system acquires monitoring video from the tower crane cab, microphone recordings from the cab, and vibration sensor data mounted on the boom as multimodal source data. First, the system processes the video. When it detects the boom starting to move, it extracts the image at that moment as a keyframe and uses a target detection model to locate the boom and the suspended load below, generating video feature data containing their position coordinates and direction of movement. Simultaneously, the system processes the audio collected by the microphone, filters out engine noise through noise reduction, identifies the harsh metallic friction sound, extracts its acoustic features, and generates audio feature data. The system performs frequency domain transformation on the data collected by the vibration sensor, detects abnormal high-frequency vibrations, extracts the frequency and amplitude as feature values, and generates sensor feature data. Ultimately, these three sets of feature data are combined into a unified preprocessed feature data, which comprehensively characterizes the potential violation or danger of abnormal sounds and vibrations accompanying the boom's movement.

[0046] Optionally, the step of dynamically assigning weights to the preprocessed feature data based on the environmental state parameters to generate weighted feature data includes:

[0047] Obtain the light intensity and noise level of the environmental state parameters;

[0048] The weights of the video feature data are adjusted according to the light intensity, and the weights of the audio feature data are adjusted according to the noise level to generate dynamic weight coefficients.

[0049] The preprocessed feature data is weighted using the dynamic weighting coefficients to generate weighted feature data.

[0050] Specifically, the environmental parameters, particularly quantified light intensity and noise level, are first acquired from on-site environmental sensors. Based on a pre-defined mapping relationship, these two environmental indicators are converted into a set of dynamic weighting coefficients. For video feature data, the weight is positively correlated with light intensity; that is, the better the lighting conditions and the higher the video data quality, the greater the weight. Conversely, in dim lighting or insufficient illumination, the weight decreases. The influence of light intensity on video features is as follows: Figure 2As shown. For audio feature data, its weight is negatively correlated with the noise level; that is, the lower the background noise, the clearer and more reliable the audio data, and the higher its weight; conversely, in noisy environments, its weight will decrease. Sensor data is usually less affected by lighting and noise, so a fixed baseline weight can be set. Finally, using these generated dynamic weight coefficients, the corresponding parts of the preprocessed feature data vector output from the previous stage are weighted to generate weighted feature data. This weighting process can be expressed by the following formula:

[0051] ,

[0052] in This represents the final weighted feature data vector. , and These represent the video, audio, and sensor feature data portions of the preprocessed feature data, respectively. The video weights are calculated based on the real-time illumination intensity L. The audio weights are calculated based on the real-time noise level N, and these two together constitute the dynamic weight coefficients. This refers to the reference weights set for the sensor data, which are typically set to 1. This process scales the modal features using multiplication operations, ensuring logical consistency in the processing.

[0053] For example, at a construction site in the late afternoon, the system is monitoring a hoisting operation that requires both verbal commands and hand signals. Environmental parameters obtained by the system indicate low light intensity and high noise levels due to the operation of a nearby generator. At this time, the system calculates based on a preset function relationship; due to the low light intensity, the reliability of the video data decreases, therefore assigning it a lower video weight. For example, 0.4. Meanwhile, due to the high noise level, verbal instructions may be interfered with, so the system assigns a lower audio weight to the audio data. For example, 0.3. However, the data from the tension sensor mounted on the hook remains unaffected, and its weight... The value is kept at 1. After the system receives a set of preprocessed feature data containing gestures, commands, and pulling force readings, it uses this set of dynamic weighting coefficients {0.4, 0.3, 1.0} to weight the data and generate the final weighted feature data.

[0054] Optionally, the step of performing multi-source data conflict judgment and resolution on the weighted feature data to generate conflict resolution event data includes:

[0055] Based on the weighted feature data, a preliminary judgment conclusion is generated;

[0056] Detect logical conflicts between the preliminary judgment conclusions and generate conflict indicators;

[0057] When the conflict identifier is received, the operating status data of the construction equipment is obtained;

[0058] The preliminary judgment conclusion is arbitrated and verified using the operational status data to generate conflict resolution event data.

[0059] Specifically, firstly, based on the input weighted feature data, multiple analysis models running in parallel generate preliminary judgments. For example, a model based on video feature data might conclude that the person is wearing a safety helmet, while another model based on sensor feature data might conclude that the safety helmet is not fastened. These conclusions together constitute a multi-faceted preliminary judgment of the same event or state. Then, a logical conflict detection mechanism is initiated. This mechanism, based on a pre-set rule base or knowledge graph, examines whether there are logical contradictions between these preliminary judgments from different sources. For example, the rule base might define that it is reasonable for the conclusions of "person entering a high-risk work area" and "not wearing a safety helmet" to occur simultaneously, but the simultaneous occurrence of "safety helmet worn" and "safety helmet not fastened" constitutes a logical conflict. Once such a conflict is detected, a conflict identifier is generated. This conflict identifier triggers the next step: actively acquiring the operating status data of the construction equipment as the basis for arbitration. This data typically comes from the equipment's controller area network or programmable logic controller, providing objective and direct equipment information such as engine speed and hydraulic status. Finally, the acquired equipment operating status data is used to arbitrate and verify the conflicting preliminary judgments. If equipment operation data shows that the relevant equipment is not in operation, then the conclusions of previous equipment violation generated by other sensors will be deemed invalid. This arbitration process resolves the contradictions between data sources, ultimately outputting a verified, internally consistent conflict resolution event data, which will serve as reliable input for subsequent behavioral analysis.

[0060] For example, the system is monitoring the working area of ​​an excavator. The video analysis module, based on its weighted feature data, initially concludes that the excavator is performing digging operations. However, the analysis module corresponding to the vibration sensor installed on the excavator's boom initially concludes that the equipment's vibration frequency is within the idle range. The system's built-in logic detects a clear logical conflict between these two conclusions, as the vibration frequency during digging operations is necessarily different from that at idle. Therefore, the system generates a conflict flag. Upon receiving this conflict flag, the system immediately requests real-time operating status data from the excavator's onboard information system. The data returned by the onboard system shows that the excavator's engine speed and hydraulic pump pressure are both at their lowest idle values. The system uses this authoritative operating status data for arbitration verification, determining that the vibration sensor's conclusion is accurate, while the video analysis conclusion may have been misjudged due to lighting or background interference. Ultimately, the system generates conflict resolution event data, whose content is corrected and confirmed that the excavator is in an idling state, thus completing an effective conflict judgment and resolution.

[0061] Optionally, the step of performing behavior pattern matching on the conflict resolution event data based on a preset behavior rule pattern library to generate a violation determination result includes:

[0062] The conflict resolution event data are combined in a time sequence to construct a real-time behavior sequence;

[0063] Compliance patterns are extracted from the behavior rule pattern library, which contains templates of typical violation behavior patterns based on image sequence recognition;

[0064] The real-time behavior sequence is matched with the compliance pattern, and a violation determination result is generated when the real-time behavior sequence deviates from the compliance pattern.

[0065] Specifically, the discrete conflict resolution event data from the previous stage are first arranged in order according to their timestamps, and then combined sequentially. This process connects independent event points into a real-time behavioral sequence that reflects the continuous actions of construction personnel or equipment. This sequence is a dynamic and structured description of the actual situation on site. Simultaneously, a pre-defined behavioral rule pattern library is accessed. This library stores various standardized operating procedures, i.e., compliance patterns, which define in detail the correct sequence of steps to be followed when performing a specific task. These compliance patterns are templates built based on safety regulations and operating manuals, using techniques such as image sequence recognition to analyze a large number of standard operating images. Then, the real-time behavioral sequence is matched with the corresponding task's compliance pattern extracted from the behavioral rule pattern library. This matching process can be implemented using sequence alignment algorithms such as dynamic time warping to calculate the similarity or difference between the two sequences. When the matching result shows that the real-time behavioral sequence deviates significantly from the compliance pattern, such as missing key safety steps or performing an incorrect sequence of actions, it is judged as a violation, and a final violation judgment result is generated. The degree of deviation can be represented by the following logic:

[0066] ,

[0067] In this expression, This represents a function used to calculate the difference between two sequences, such as based on edit distance or dynamic time warping algorithms. It is a real-time behavioral sequence constructed by combining conflict resolution event data in a time sequence. It is a compliance pattern extracted from the behavioral rule pattern library that is relevant to the current construction task. It is a pre-set threshold, representing the maximum acceptable deviation. When the calculated difference... Greater than this threshold At that time, a violation determination result is generated.

[0068] For example, suppose a construction site regulation requires workers to wear safety harnesses before performing high-altitude work. A behavioral rule pattern library pre-defines corresponding compliance patterns, with the event sequence being "entering the work area," "wearing a safety harness," and "climbing scaffolding." The monitoring system, through preliminary processing, obtains a series of conflict resolution event data related to a particular worker. The system first combines this data temporally to construct the worker's real-time behavioral sequence, identifying the behavior as "entering the work area" and "climbing scaffolding." Subsequently, the system extracts the compliance pattern for high-altitude work from the behavioral rule pattern library. During behavioral pattern matching, the system compares the real-time "entering the work area" and "climbing scaffolding" sequence with the compliant "entering the work area," "wearing a safety harness," and "climbing scaffolding" sequence. The comparison reveals that the real-time behavioral sequence is missing the crucial event of "wearing a safety harness." Since this deviation exceeds a preset safety threshold, the system determines the behavior is in violation and immediately generates a violation determination result, issuing an alert to management.

[0069] Optionally, generating a preliminary judgment conclusion based on the weighted feature data includes:

[0070] The weighted feature data is subjected to multimodal spatiotemporal feature map convolution processing to obtain a structured event representation tensor;

[0071] Adaptive semantic inference and conflict probability assessment are performed on the structured event representation tensor to generate a preliminary judgment conclusion.

[0072] Specifically, the input weighted feature data is first processed by multimodal spatiotemporal feature map convolution. In this step, the weighted feature data, which integrates video, audio, and sensor information, is reconstructed into a spatiotemporal feature map with spatial and temporal dimensions. The spatiotemporal features are as follows: Figure 3 As shown. Subsequently, a 3D convolutional neural network model capable of processing spatiotemporal information is used to perform depthwise convolution on the feature map. The output of this process is a structured event representation tensor, which is a high-dimensional array containing condensed numerical representations of the dynamic and static patterns of events learned from the original data. This transformation process can be represented by the following equation:

[0073] ,

[0074] in The structured event representation tensor representing the output. It is the input weighted feature data. This represents a deep learning network model that performs multimodal spatiotemporal feature map convolution processing. Next, adaptive semantic inference and conflict probability assessment are performed on this structured event representation tensor. This tensor is then input into subsequent fully connected layers and the classifier of the neural network. Adaptive semantic inference refers to the model mapping this abstract numerical tensor to one or more semantic labels with clear business meanings based on the learned features. Simultaneously, conflict probability assessment provides an internal consistency evaluation score along with the output conclusion. Finally, a preliminary judgment conclusion is generated, and this inference process can be represented as:

[0075] ,

[0076] in This represents the initial judgment conclusion that is ultimately generated. These are specific semantic tags, such as "person not wearing a safety helmet." It is a probability vector associated with the label, containing information such as confidence level and conflict probability. The network portion that performs inference and evaluation functions. It is the structured event representation tensor output from the previous step.

[0077] For example, the system receives a set of weighted feature data, where the video portion shows a worker moving rapidly along the edge of a high-altitude work platform, the audio portion contains muffled shouts, and the sensor portion comes from a gyroscope on the worker's safety harness, showing drastic angle changes. This set of weighted feature data... The data is fed into the multimodal spatiotemporal feature map convolution processing module. (Model) Through 3D convolution operations, the rapid movement trajectory of a human figure, the acoustic pattern of shouts, and the sharp flips of the gyroscope in the video were captured simultaneously. These features were fused into a structured event representation tensor. This tensor numerically encapsulates the concepts of height, speed, and instability. Subsequently, this tensor... Sent to the inference layer The model's adaptive semantic inference function interprets it as semantic labels. This provides a preliminary assessment of the risk of falls from height. During conflict probability assessment, because multiple sources of evidence strongly point to the same event, the model generates a vector containing high confidence and low conflict probability. This ultimately produces a clear preliminary assessment of the risk of falls from height, accompanied by a high confidence level, for use in subsequent conflict resolution or alarm procedures.

[0078] Optionally, generating environmental state parameters based on the environmental conditions includes:

[0079] Using the aforementioned environmental conditions, quantitative environmental indicators are obtained;

[0080] Spatial state mapping is performed on the quantified environmental indicators to generate environmental state parameters.

[0081] Specifically, the first step involves obtaining quantified environmental indicators by utilizing the environmental conditions. This step utilizes various environmental sensors deployed at the construction site, such as illuminometers, sound level meters, and temperature and humidity sensors, to measure the environmental conditions in real time. The sensors convert the collected analog signals of light and sound into specific, continuous numerical values, such as a light intensity of 500 lux and a noise level of 85 dB. These precise values ​​are the quantified environmental indicators. Subsequently, these quantified environmental indicators are spatially mapped to generate the final environmental state parameters. Spatial state mapping is a process of discretizing and labeling continuously changing numerical values. Its purpose is to summarize an infinite number of possible quantified indicator values ​​into a finite number of states with clear semantics. According to preset thresholds or rules, the quantified indicators are mapped to corresponding state levels. This mapping process can be represented by the following function:

[0082] ,

[0083] in The output environmental state parameter is a collection of multiple environmental dimensions. It is an input quantitative environmental indicator, such as a specific lux or decibel value. This represents the function that performs spatial state mapping. This function encapsulates a series of classification rules or threshold judgment logic to complete the transformation from continuous values ​​to discrete states.

[0084] For example, suppose a construction site transitions from daytime to nighttime. At 3 PM, the illuminance meter measures a quantitative environmental index of 15,000 lux, and the sound level meter measures 75 dB. The system inputs these indicators into a spatial state mapping function. The function maps light intensity greater than 1000 lux to a sufficiently lit state and noise levels between 70 and 90 decibels to a moderately noisy state, based on preset rules. Therefore, the system generates the following environmental state parameters at this point: {Light intensity: Sufficient light, Noise level: Moderate noise}. At 7 PM, the illuminance meter reading drops to 50 lux, while the sound level meter reading rises to 95 decibels due to concentrated nighttime construction work. The system performs the mapping again; this time, 50 lux is mapped to insufficient light, and 95 decibels is mapped to high noise. The system then updates the environmental state parameters to {Light intensity: Insufficient light, Noise level: High noise}.

[0085] Optionally, obtaining quantitative environmental indicators using the environmental conditions includes:

[0086] The environmental conditions are subjected to multi-source signal acquisition and analog-to-digital conversion to obtain digital environmental signals;

[0087] The digital environmental signal is filtered and denoised, and its eigenvalues ​​are calculated to obtain quantified environmental indicators.

[0088] Specifically, the first step involves multi-source signal acquisition and analog-to-digital conversion of the environmental conditions at the site. This step utilizes sensor arrays deployed on-site, such as photodiodes for sensing light intensity and microphones for sensing sound intensity, to capture continuously changing physical quantities. These sensors convert the received physical signals, such as light and sound, into continuous analog electrical signals. Subsequently, these analog signals are sampled and quantized at high speed using an analog-to-digital converter, transforming them into a time series composed of discrete values, i.e., a digital environmental signal. Next, the acquired digital environmental signal undergoes filtering, noise reduction, and eigenvalue calculation. Since the original digital signal may contain noise and glitches caused by circuit interference or sudden environmental changes, digital filters, such as moving average filters, are used to smooth the signal sequence and eliminate these interferences. After obtaining a clean signal, eigenvalue calculation is required to obtain a value that stably represents the environmental conditions over a certain period. This typically refers to calculating a statistical characteristic, such as the mean, variance, or root mean square, of the signal sequence within a preset time window. This process can be summarized by the following formula:

[0089] ,

[0090] in This represents the final quantitative environmental indicator. It is the first in the time series of the filtered and denoised digital environmental signal. The value of each sampling point. This is the total number of sampling points used in the calculation, i.e., the size of the time window. This formula uses the calculation of the mean as an example to demonstrate how to extract a stable feature value from a signal sequence.

[0091] For example, to obtain a quantitative environmental indicator of light intensity at a construction site, the system first acquires environmental signals using a light sensor. The sensor outputs an analog voltage signal proportional to the light intensity. An analog-to-digital converter samples this voltage signal at a fixed frequency, generating a series of discrete digital environmental signals. This raw digital signal may fluctuate drastically due to rapid cloud movement or momentary flashes of light. To eliminate these effects, the system applies a moving average filtering algorithm to this digital signal, resulting in a smoother signal curve. Subsequently, the system sets a calculation window of 1 second and performs an eigenvalue calculation to average the values ​​of all filtered sampling points within the window. The calculated average value, for example, 500, is used as the quantitative environmental indicator of light intensity at the current moment. This value stably reflects the average light conditions over the past second, rather than an extreme value at a certain instant, and is therefore more reliable.

[0092] Optionally, the step of combining the conflict resolution event data in a time sequence to construct a real-time behavior sequence includes:

[0093] The conflict resolution event data is injected into the spatiotemporal anchor point to obtain a time-stamped monitoring event dataset;

[0094] The construction behavior concept mapping and dimensional compression are performed on the time-stamped monitoring event dataset to construct a real-time behavior sequence.

[0095] Specifically, the input conflict resolution event data is first injected with spatiotemporal anchors to obtain a time-stamped monitoring event dataset. In this step, precise time and spatial information is appended to each verified, discrete conflict resolution event data point. The time information is the timestamp of the event, while the spatial information is the specific location of the event, such as coordinates obtained through GPS or Building Information Modeling. By injecting spatiotemporal anchors, the originally independent events are given a clear spatiotemporal context, forming a dataset containing information such as [event content, time, and location], i.e., a time-stamped monitoring event dataset. Subsequently, construction behavior concept mapping and dimensionality compression are performed on this time-stamped monitoring event dataset. Construction behavior concept mapping is the core semantic enhancement step. Based on a professional knowledge base in the construction field, it maps one or more consecutive, logically related low-level monitoring events to a higher-level behavioral concept with clear construction significance. Dimensional compression, in this mapping process, summarizes and refines the information from multiple low-level events, forming a more concise data unit describing high-level behavior, thereby constructing the final real-time behavior sequence. This abstraction process can be conceptually represented by the following function:

[0096] ,

[0097] in, Represents the final constructed real-time sequence of behaviors, which is an ordered list of a series of high-level construction behavior concepts. It is the input time-stamped monitoring event dataset. This represents a complex mapping and compression function that encapsulates the logical rules for mapping, combining, and refining underlying spatiotemporal event points into upper-level behavioral concepts.

[0098] For example, suppose the system is monitoring the hoisting process of precast wall panels and receives a series of conflict resolution event data. First, the system injects spatiotemporal anchors into these events, resulting in a time-stamped monitoring event dataset, which may include: Event A {Content: "Crane hook", Time: 10:01:05, Location: Wall panel stacking area}; Event B {Content: "Wall panel lifted", Time: 10:01:30, Location: Wall panel stacking area}; Event C {Content: "Wall panel moved to installation point", Time: 10:02:15, Location: Floor X, Axis Y}; Event D {Content: "Wall panel installation completed", Time: 10:03:00, Location: Floor X, Axis Y}. Next, the system performs construction behavior concept mapping and dimensional compression on this dataset. Based on the definitions in the knowledge base, the system identifies that events A, B, C, and D are spatiotemporally continuous and logically related, collectively constituting a complete wall panel hoisting operation. Therefore, the system maps these four low-level events to a high-level behavioral concept: Execute wall panel hoisting operation. During the dimensionality compression process, the system uses the time of event A as the start time of this behavior, the time of event D as the end time, and summarizes the location information as "from the stacking area to the installation point". Finally, after these four low-level events are compressed and refined, they form an element in the real-time behavioral sequence, namely [Behavioral concept: "Execute wall panel hoisting operation", start and end time: 10:01:05-10:03:00].

[0099] Based on the same inventive concept, such as Figure 4 As shown, the present invention also provides an automatic monitoring system for violations in infrastructure construction operations based on multimodal data fusion, the system comprising:

[0100] The data acquisition and processing module is used to acquire multimodal source data of the target scene, perform data identification and standardization processing on the multimodal source data, and generate preprocessed feature data;

[0101] An environmental perception module is used to acquire on-site environmental conditions and generate environmental state parameters based on those conditions.

[0102] The weight allocation module is used to dynamically allocate weights to the preprocessed feature data based on the environmental state parameters, and generate weighted feature data.

[0103] The conflict resolution module is used to perform multi-source data conflict judgment and resolution on the weighted feature data, and generate conflict resolution event data;

[0104] The behavior pattern analysis module is used to perform behavior pattern matching on the conflict resolution event data based on a preset behavior rule pattern library, and generate violation judgment results.

[0105] It should be noted that the electrical connections between the various units described above do not necessarily represent direct or indirect connections. Any indirect connection method can be applied to the embodiments of the present invention as long as it achieves the purpose of the present invention. The above descriptions are merely exemplary embodiments of the present invention and should not be construed as limiting the scope of the present invention.

[0106] All equivalent changes and modifications made in accordance with the teachings of this invention are still within the scope of this invention. Those skilled in the art will readily conceive of other embodiments of this invention upon considering the specification and the disclosure of practical truth. This application is intended to cover any variations, uses, or adaptations of this invention that follow the general principles of this invention and include common knowledge or conventional techniques in the art not described herein.

Claims

1. A method for automatic monitoring of violations in infrastructure construction operations based on multimodal data fusion, characterized in that, The method includes: The process involves acquiring multimodal source data of a target scene, performing data identification and standardization on the multimodal source data to generate preprocessed feature data, including: extracting keyframes and locating target regions from the multimodal source data to generate video feature data; performing noise reduction and acoustic feature extraction on the multimodal source data to generate audio feature data; performing frequency domain transformation and feature value extraction on the multimodal source data to generate sensor feature data; and combining the video feature data, audio feature data, and sensor feature data to generate preprocessed feature data. Obtain the environmental conditions at the site, and generate environmental state parameters based on the environmental conditions; Based on the environmental state parameters, the preprocessed feature data is dynamically weighted to generate weighted feature data; The weighted feature data is subjected to multi-source data conflict judgment and resolution to generate conflict resolution event data, including: generating preliminary judgment conclusions based on the weighted feature data; detecting logical conflicts between the preliminary judgment conclusions to form conflict identifiers; when the conflict identifiers are received, acquiring the operating status data of the construction equipment; and using the operating status data to arbitrate and verify the preliminary judgment conclusions to generate conflict resolution event data. Based on a preset behavioral rule pattern library, behavioral pattern matching is performed on the conflict resolution event data to generate a violation judgment result. This includes: combining the conflict resolution event data in a time sequence to construct a real-time behavioral sequence; extracting compliance patterns from the behavioral rule pattern library, which contains typical violation action pattern templates based on image sequence recognition; matching the real-time behavioral sequence with the compliance patterns, and generating a violation judgment result when the real-time behavioral sequence deviates from the compliance patterns. This includes: injecting the conflict resolution event data into spatiotemporal anchors to obtain a time-stamped monitoring event dataset; and performing construction behavior concept mapping and dimensional compression on the time-stamped monitoring event dataset to construct a real-time behavioral sequence.

2. The method for automatic monitoring of violations in infrastructure construction operations based on multimodal data fusion according to claim 1, characterized in that, The step of dynamically assigning weights to the preprocessed feature data based on the environmental state parameters to generate weighted feature data includes: Obtain the light intensity and noise level of the environmental state parameters; The weights of the video feature data are adjusted according to the light intensity, and the weights of the audio feature data are adjusted according to the noise level to generate dynamic weight coefficients. The preprocessed feature data is weighted using the dynamic weighting coefficients to generate weighted feature data.

3. The method for automatic monitoring of violations in infrastructure construction operations based on multimodal data fusion according to claim 1, characterized in that, The preliminary judgment conclusion generated based on the weighted feature data includes: The weighted feature data is subjected to multimodal spatiotemporal feature map convolution processing to obtain a structured event representation tensor; Adaptive semantic inference and conflict probability assessment are performed on the structured event representation tensor to generate a preliminary judgment conclusion.

4. The method for automatic monitoring of violations in infrastructure construction operations based on multimodal data fusion according to claim 1, characterized in that, The generation of environmental state parameters based on the environmental conditions includes: Using the aforementioned environmental conditions, quantitative environmental indicators are obtained; Spatial state mapping is performed on the quantified environmental indicators to generate environmental state parameters.

5. The method for automatic monitoring of violations in infrastructure construction operations based on multimodal data fusion according to claim 4, characterized in that, The process of obtaining quantitative environmental indicators using the aforementioned environmental conditions includes: The environmental conditions are subjected to multi-source signal acquisition and analog-to-digital conversion to obtain digital environmental signals; The digital environmental signal is filtered and denoised, and its eigenvalues ​​are calculated to obtain quantified environmental indicators.

6. A multimodal data fusion-based automatic monitoring system for violations in infrastructure construction operations, applied to the multimodal data fusion-based automatic monitoring method for violations in infrastructure construction operations as described in any one of claims 1-5, characterized in that, The system includes: The data acquisition and processing module is used to acquire multimodal source data of the target scene, perform data identification and standardization processing on the multimodal source data, and generate preprocessed feature data. This includes: extracting keyframes and locating target regions from the multimodal source data to generate video feature data; performing noise reduction and acoustic feature extraction on the multimodal source data to generate audio feature data; performing frequency domain transformation and feature value extraction on the multimodal source data to generate sensor feature data; and combining the video feature data, audio feature data, and sensor feature data to generate preprocessed feature data. An environmental perception module is used to acquire on-site environmental conditions and generate environmental state parameters based on those conditions. The weight allocation module is used to dynamically allocate weights to the preprocessed feature data based on the environmental state parameters, and generate weighted feature data. The conflict resolution module is used to perform multi-source data conflict judgment and resolution on the weighted feature data and generate conflict resolution event data. This includes: generating preliminary judgment conclusions based on the weighted feature data; detecting logical conflicts between the preliminary judgment conclusions and forming conflict identifiers; acquiring the operating status data of the construction equipment when the conflict identifiers are received; and using the operating status data to arbitrate and verify the preliminary judgment conclusions to generate conflict resolution event data. The behavior pattern analysis module is used to perform behavior pattern matching on the conflict resolution event data based on a preset behavior rule pattern library to generate violation judgment results. This includes: combining the conflict resolution event data in a time sequence to construct a real-time behavior sequence; extracting compliance patterns from the behavior rule pattern library, which contains templates of typical violation action patterns based on image sequence recognition; matching the real-time behavior sequence with the compliance patterns, and generating a violation judgment result when the real-time behavior sequence deviates from the compliance pattern. This includes: injecting the conflict resolution event data into spatiotemporal anchors to obtain a time-stamped monitoring event dataset; and performing construction behavior concept mapping and dimensional compression on the time-stamped monitoring event dataset to construct a real-time behavior sequence.

Citation Information

Patent Citations

  • Event detection method and device based on multiple modes, electronic equipment and storage medium

    CN119989258A

  • Emergency state alarm system

    CN120656287A

  • Video analysis-based multi-scene operator violation behavior identification method and system

    CN120726699A