Automatic monitoring method for infrastructure construction operation violation based on multi-modal data fusion
By employing multimodal data fusion and dynamic weight allocation techniques, the problem of data interference in complex environments for infrastructure construction safety monitoring systems has been solved, enabling efficient and reliable identification and monitoring of violations, and improving the system's engineering practicality and computational efficiency.
Patent Information
- Application Number
- CN202511501296.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-21
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2045-10-21
AI Technical Summary
Existing infrastructure construction safety monitoring systems are susceptible to environmental interference due to a single data source, have low efficiency in utilizing multi-source data, and lack effective data arbitration and fusion mechanisms, leading to inconsistent conclusions in determining violations.
By employing a multimodal data fusion and dynamic weight allocation mechanism, combined with environmental adaptive perception and multi-source data conflict resolution technology, automatic monitoring and pattern recognition in the construction environment are achieved through multimodal spatiotemporal feature fusion, dynamic weight allocation, and behavioral pattern matching.
It improves the accuracy and robustness of violation monitoring, reduces the probability of false alarms and missed alarms, realizes intelligent full-chain supervision of the construction process, and outputs reliable violation judgment results.
Smart Images

Figure CN120974243A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of pattern recognition and image data processing, and particularly relates to a method for automatically monitoring construction operation violations in infrastructure construction by multi-modal data fusion. BACKGROUND
[0002] In the field of infrastructure construction safety monitoring, multi-modal data fusion technology has become a key means to improve the level of automated supervision. This technology involves the collection, processing and analysis of source data from multiple modalities such as video, audio and sensors, aiming to achieve intelligent identification and early warning of construction violations through image data processing and pattern recognition methods.
[0003] In the prior art, common monitoring solutions mainly rely on single type of sensor data or simple multi-source data juxtaposition. For example, some systems use fixed position cameras for video monitoring, identifying safety helmet wearing or personnel activity areas through image processing algorithms. Other solutions attempt to combine audio sensors and vibration sensors to detect abnormal sound and equipment operating status respectively, and then aggregate the results after independent judgment. These methods usually use static thresholds or pre-set rules to analyze various types of data.
[0004] However, the existing technical solutions have obvious defects. Systems based on a single data source are susceptible to environmental conditions, such as significant performance degradation in low light or high noise environments. In addition, multi-source data is often simply superimposed, lacking a dynamic weight adjustment mechanism based on environmental state, resulting in low data utilization efficiency. At the same time, when different sources of data conflict in judgment, existing systems lack effective arbitration and fusion mechanisms, making it difficult to generate consistent and reliable violation determination conclusions. SUMMARY
[0005] To solve the above problems, the present application provides a method for automatically monitoring construction operation violations in infrastructure construction by multi-modal data fusion, which adopts multi-modal spatio-temporal feature fusion and dynamic weight distribution mechanism, combines environmental self-adaptive perception and multi-source data conflict resolution technology, and can realize automatic monitoring and pattern recognition of violations in construction environment.
[0006] The above objectives can be achieved through the following solutions: The application discloses a kind of multi-modal data fusion's infrastructure construction operation violation automatic monitoring method, including obtaining the multi-modal source data of target scene, to multi-modal source data is carried out data identification and standardization processing, generates pre-processing feature data;Environment condition is obtained in site, generates environment state parameter;Based on environment state parameter, pre-processing feature data is carried out dynamic weight distribution, generates weighted feature data;To weighted feature data is carried out multi-source data conflict judgment and resolution, generates conflict resolution event data;Based on the behavior rule mode library of preestablished, conflict resolution event data is carried out behavior mode matching, generates violation determination result.
[0007] Optionally, the multi-modal source data is carried out data identification and standardization processing, and the pre-processing feature data includes: the multi-modal source data is carried out key frame extraction and target area positioning, and video feature data is generated;The multi-modal source data is carried out noise reduction processing and acoustic feature extraction, and audio feature data is generated;The multi-modal source data is carried out frequency domain transformation and feature value extraction, and sensor feature data is generated;The video feature data, audio feature data and the sensor feature data are combined, and pre-processing feature data is generated.
[0008] Optionally, the pre-processing feature data is carried out dynamic weight distribution based on the environment state parameter, and the weighted feature data includes: the light intensity and noise level of the environment state parameter are obtained;According to the light intensity, the weight of the video feature data is adjusted, and according to the noise level, the weight of the audio feature data is adjusted, and dynamic weight coefficient is generated;The pre-processing feature data is weighted using the dynamic weight coefficient, and the weighted feature data is generated.
[0009] Optionally, the multi-source data conflict judgment and resolution of the weighted feature data is carried out, and the conflict resolution event data includes: based on the weighted feature data, preliminary judgment conclusion is generated;The logical conflict between the preliminary judgment conclusion is detected, and conflict identifier is formed;When the conflict identifier is received, the running state data of construction equipment is obtained;The preliminary judgment conclusion is arbitrated and verified using the running state data, and the conflict resolution event data is generated.
[0010] Optionally, the behavior mode matching of the conflict resolution event data is carried out based on the behavior rule mode library, and the violation determination result includes: the conflict resolution event data is time-series combined, and real-time behavior sequence is constructed;Compliance mode is extracted from the behavior rule mode library, and the behavior rule mode library includes the typical violation action mode template based on image sequence identification;The real-time behavior sequence is matched with the compliance mode, and when the real-time behavior sequence deviates from the compliance mode, the violation determination result is generated.
[0011] Optionally, the generating a preliminary judgment conclusion based on the weighted feature data comprises: performing multi-modal spatio-temporal feature map convolution processing on the weighted feature data to obtain a structured event representation tensor; and performing adaptive semantic inference and conflict probability evaluation on the structured event representation tensor to generate the preliminary judgment conclusion.
[0012] Optionally, the generating an environment state parameter based on the environment condition comprises: obtaining a quantitative environment index using the environment condition; and performing spatial state mapping on the quantitative environment index to generate the environment state parameter.
[0013] Optionally, the obtaining a quantitative environment index using the environment condition comprises: performing multi-source signal acquisition and analog-to-digital conversion on the environment condition to obtain a digital environment signal; and performing filtering and denoising and eigenvalue calculation on the digital environment signal to obtain the quantitative environment index.
[0014] Optionally, the combining the conflict resolution event data in time sequence to construct a real-time behavior sequence comprises: injecting the conflict resolution event data into a space-time anchor to obtain a time-standardized monitoring event data set; and performing construction behavior concept mapping and dimension compression on the time-standardized monitoring event data set to construct the real-time behavior sequence.
[0015] Based on the same inventive concept, the application also provides a multi-modal data fusion-based construction operation violation automatic monitoring system, which comprises: a data acquisition and processing module, configured to acquire multi-modal source data of a target scene, perform data recognition and standardization processing on the multi-modal source data, and generate preprocessed feature data; an environment perception module, configured to acquire an environment condition on site and generate an environment state parameter based on the environment condition; a weight distribution module, configured to perform dynamic weight distribution on the preprocessed feature data based on the environment state parameter to generate weighted feature data; a conflict resolution module, configured to perform multi-source data conflict judgment and resolution on the weighted feature data to generate conflict resolution event data; and a behavior pattern analysis module, configured to perform behavior pattern matching on the conflict resolution event data based on a pre-set behavior rule pattern library to generate a violation determination result.
[0016] Compared with the prior art, the application has the following advantages: The application improves the accuracy and robustness of violation behavior monitoring in a complex environment through multi-modal data fusion and dynamic weight distribution technology. The system can adaptively adjust the contribution of different modal data such as video and audio, effectively overcoming the defect of decreased sensing ability of a single sensor under poor conditions such as insufficient light or noise interference, thereby ensuring the continuous generation of reliable feature data in real and variable construction scenes.
[0017] This invention helps reduce the probability of false alarms and missed alarms by establishing a multi-source data conflict judgment and arbitration verification mechanism. When there are logical conflicts in the preliminary judgment conclusions generated by different sensors, the system can introduce third-party data such as equipment operating status for cross-verification and intelligent arbitration, and finally output consistent event data after resolution, thereby improving the consistency of monitoring conclusions.
[0018] This invention utilizes a compliance pattern library based on image sequences to determine behavioral patterns, achieving intelligent supervision of the entire construction process. The system efficiently matches real-time behavioral sequences with templates of typical violations, enabling it to identify not only static violations but also dynamic violations in the work process.
[0019] This invention standardizes and modularizes the data processing workflow, forming a closed-loop operation from feature extraction and weight allocation to conflict resolution, thereby improving the system's engineering practicality and computational efficiency. Through temporal combination and dimensionality compression techniques, the system can transform multimodal data into concise real-time behavioral sequences, ultimately outputting intuitive and clear violation judgment results, providing an efficient technical tool for on-site safety management.
[0020] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures pointed out in the description, claims and drawings. Attached Figure Description
[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 This is a flowchart illustrating an automatic monitoring method for violations in infrastructure construction operations based on multimodal data fusion, according to an embodiment of the present invention.
[0023] Figure 2 This is a video feature influence diagram of light intensity according to an embodiment of the present invention.
[0024] Figure 3 This is a spatiotemporal feature map of an embodiment of the present invention.
[0025] Figure 4 This is a schematic diagram of the structure of an automatic monitoring system for violations in infrastructure construction operations based on multimodal data fusion, according to an embodiment of the present invention. Detailed Implementation
[0026] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0027] Reference Figure 1 One embodiment of the present invention proposes an automatic monitoring method for violations in infrastructure construction operations based on multimodal data fusion. It adopts a multimodal spatiotemporal feature fusion and dynamic weight allocation mechanism, combined with environmental adaptive perception and multi-source data conflict resolution technology, which can realize automatic monitoring and pattern recognition of violations in the construction environment.
[0028] The method described in this embodiment specifically includes: Acquire multimodal source data of the target scene, perform data identification and standardization processing on the multimodal source data, and generate preprocessed feature data; Obtain the environmental conditions at the site, and generate environmental state parameters based on the environmental conditions; Based on the environmental state parameters, the preprocessed feature data is dynamically weighted to generate weighted feature data; The weighted feature data is subjected to multi-source data conflict judgment and resolution to generate conflict resolution event data; Based on a preset behavioral rule pattern library, behavioral pattern matching is performed on the conflict resolution event data to generate violation judgment results.
[0029] Specifically, the method first transforms complex, heterogeneous visual and auditory information from the field into unified feature data through multimodal data acquisition and standardized preprocessing. The core innovation lies in the introduction of an environmental perception mechanism. This mechanism generates environmental state parameters by acquiring real-time environmental conditions and uses these parameters to dynamically evaluate and weight the reliability of each modal feature data. Subsequently, the method designs a conflict judgment and resolution logic to perform internal consistency checks on the weighted and fused data, proactively identifying and resolving contradictions between different information sources, thereby generating highly reliable event data that has undergone cross-validation. Finally, these discrete and reliable event data are placed in a behavioral rule pattern library for temporal context comparison. By determining whether real-time behavioral sequences deviate from standard operating procedures, the method achieves accurate judgment from low-level data perception to high-level behavioral logic violations.
[0030] Optionally, the step of performing data identification and standardization processing on the multimodal source data to generate preprocessed feature data includes: key frame extraction and target region localization are performed on the multi-modal source data to generate video feature data; noise reduction processing and acoustic feature extraction are performed on the multi-modal source data to generate audio feature data; frequency domain transformation and feature value extraction are performed on the multi-modal source data to generate sensor feature data; The video feature data, audio feature data and sensor feature data are combined to generate pre-processing feature data.
[0031] Specifically, first, the data streams from different modalities are processed in parallel. For video data streams, key frame extraction and target region localization are performed. Key frame extraction aims to filter out static images that can represent significant changes in scene content from continuous video streams, in order to reduce data redundancy and capture core actions. Subsequently, through target region localization techniques such as object detection algorithms in deep learning, key targets such as construction personnel and mechanical equipment are accurately identified and framed, so as to focus analysis on the effective area, and finally generate video feature data containing target position, posture and other information. For audio data streams, noise reduction processing and acoustic feature extraction are performed. Noise reduction processing aims to filter out background noise such as wind noise and equipment roar common in construction sites, to improve the signal-to-noise ratio of voice or key event sound. Then, acoustic features that can represent sound content are extracted, such as Mel-frequency cepstral coefficients, which can effectively describe the spectral characteristics of voice and sound effects, thereby generating audio feature data. For sensor data streams, frequency domain transformation and feature value extraction are implemented. Sensor data is usually a time series signal such as acceleration and angular velocity, which is converted from time domain to frequency domain through fast Fourier transform, etc., which can reveal its internal periodicity and vibration mode. Then, key feature values such as peak frequency and energy distribution are calculated from the frequency domain signal to generate sensor feature data. Finally, the video feature data, audio feature data and sensor feature data extracted independently are combined to form a unified, high-dimensional pre-processing feature data vector. This combination process can be represented by the following formula: , wherein represents the final generated pre-processing feature data vector, represents the video feature data obtained from the video and encoded, represents the acoustic feature data obtained from the audio, represents the feature value data obtained from the sensor data. This formula represents the splicing of feature vectors from different modalities, creating a standardized data structure that can comprehensively describe the instantaneous state of the construction site.
[0032] Exemplarily, in a scenario of monitoring the operation of a tower crane, the system obtains the monitoring video of the tower crane cab, the microphone recording in the cab, and the vibration sensor data installed on the boom as multi-modal source data. First, the system processes the video, and when it detects that the boom starts to move, it extracts the image at that moment as a key frame, and uses a target detection model to locate the boom and the heavy object suspended below to generate video feature data containing the position coordinates and movement direction . At the same time, the system processes the audio collected by the microphone, filters out the engine noise through noise reduction, identifies the harsh metal friction sound, and extracts its acoustic features to generate audio feature data . For the data collected by the vibration sensor, the system performs frequency domain transformation, finds abnormal high-frequency vibration, extracts the frequency and amplitude as feature values, and generates sensor feature data . Finally, the three sets of feature data are combined into a unified pre-processing feature data, which comprehensively represents the potential violation or dangerous event of abnormal sound and vibration accompanying the movement of the boom.
[0033] Optionally, the dynamic weight allocation of the pre-processing feature data based on the environmental state parameters to generate weighted feature data comprises: Obtain the light intensity and noise level of the environmental state parameters; Adjust the weight of the video feature data according to the light intensity, and adjust the weight of the audio feature data according to the noise level to generate dynamic weight coefficients; Weight the pre-processing feature data using the dynamic weight coefficients to generate weighted feature data.
[0034] Specifically, first, the environmental state parameters provided by the on-site environmental sensors are obtained, especially the quantized light intensity and noise level. According to the pre-set mapping relationship, these two environmental indicators are converted into a set of dynamic weight coefficients. For video feature data, its weight is positively correlated with light intensity, i.e. the better the lighting conditions, the higher the quality of video data, and its weight also increases accordingly; on the contrary, in dim light or insufficient illumination, its weight will decrease. The influence of light intensity on video feature is shown in Figure 2 . For audio feature data, its weight is negatively correlated with noise level, i.e. the smaller the background noise, the clearer and more reliable the audio data, and its weight is higher; on the contrary, in a noisy environment, its weight will decrease. Sensor data is usually less affected by light and noise, so a fixed reference weight can be set. Finally, these generated dynamic weight coefficients are used to weight the corresponding parts in the pre-processing feature data vector output in the previous stage to generate weighted feature data. This weighting process can be represented by the following formula: , wherein represents the final generated weighted feature data vector, , and represent the video, audio and sensor feature data parts in the pre-processed feature data respectively. is the video weight calculated according to the real-time light intensity L, is the audio weight calculated according to the real-time noise level N, and the two together constitute the dynamic weight coefficient. is the reference weight set for the sensor data, which can usually be set to 1. This process scales each modality feature through multiplication, ensuring the logical consistency of the processing.
[0035] Exemplarily, at a construction site in the evening, the system is monitoring a hoisting operation that requires both verbal instructions and hand signals. The environmental state parameters obtained by the system show that the current light intensity is low, and the operation of a nearby generator causes the noise level on site to be high. At this time, the system calculates according to the preset function relationship, and since the light intensity is low, the reliability of the video data decreases, so it allocates a lower video weight , for example, 0.4. At the same time, since the noise level is high, the verbal instructions may be disturbed, and the system allocates a lower audio weight to the audio data, for example, 0.3. The tension sensor data installed on the hook is not affected, and its weight remains 1. When the system receives a set of pre-processed feature data containing hand signals, passwords and tension readings, it uses this set of dynamic weight coefficients {0.4, 0.3, 1.0} to weight it, generating the final weighted feature data.
[0036] Optionally, the multi-source data conflict judgment and elimination of the weighted feature data to generate conflict resolution event data comprises: generate preliminary judgment conclusions based on the weighted feature data; detect logical conflicts between the preliminary judgment conclusions to form conflict identifiers; when receiving the conflict identifiers, obtain the operating state data of the construction equipment; use the operating state data to arbitrate and verify the preliminary judgment conclusions to generate conflict resolution event data.
[0037] Specifically, first, based on the input weighted feature data, preliminary judgment conclusions are generated by multiple analysis models running in parallel. For example, a model based on video feature data may output the conclusion that the personnel have worn safety helmets, while another model based on sensor feature data may output the conclusion that the safety helmets are not fastened, which together constitute multi-angle preliminary judgments on the same event or state. Then, a logical conflict detection mechanism is started, which checks whether there is a logical contradiction between these preliminary judgment conclusions from different sources according to a pre-set rule base or knowledge graph. For example, the rule base may define that the simultaneous occurrence of the two conclusions that personnel enter a high-risk operation area and do not wear safety helmets is reasonable, but the simultaneous occurrence of the two conclusions that safety helmets are worn and safety helmets are not fastened constitutes a logical conflict. Once such a conflict is detected, a conflict identifier is generated. The conflict identifier triggers the next action, that is, actively obtaining the operation state data of the construction equipment as the arbitration basis. Such data usually comes from the controller area network or programmable logic controller of the equipment, which can provide objective and direct equipment information such as engine speed, hydraulic state, etc. Finally, the device operation state data obtained is used to arbitrate and verify the preliminary judgment conclusions with conflicts. If the device operation data shows that the relevant device is in a non-working state, the previous conclusion of device violation operation generated by other sensors will be judged as invalid. Through this arbitration process, the contradiction between the data sources is resolved, and a verified and internally logically consistent conflict resolution event data is finally output, which will be used as a reliable input for subsequent behavior analysis.
[0038] Illustratively, the system is monitoring the working range of a excavator. Based on its weighted feature data, the video analysis module draws the preliminary judgment conclusion that the excavator is performing a digging operation. However, the analysis module corresponding to the vibration sensor installed on the excavator's boom draws the preliminary judgment conclusion that the device vibration frequency is in the idle speed range. The system's built-in logic detects that there is a clear logical conflict between the two conclusions, because the vibration frequency when performing a digging operation must be different from the idle state, so the system generates a conflict identifier. Upon receiving the conflict identifier, the system immediately requests the real-time operation state data from the excavator's on-board information system. The data returned by the on-board system shows that the engine speed and hydraulic pump pressure of the excavator are both at the lowest idle speed. The system uses this authoritative operation state data to arbitrate and verify, determining that the conclusion of the vibration sensor is accurate, while the conclusion of the video analysis may be a false judgment due to light and shadow or background interference. Finally, the system generates a conflict resolution event data, the content of which is corrected and confirmed as the excavator being in an idle state, thus completing an effective conflict judgment and resolution.
[0039] Optionally, the behavior pattern matching of the conflict resolution event data based on the pre-set behavior rule mode base to generate the violation judgment result comprises: combining the conflict resolution event data in time sequence to construct a real-time behavior sequence; extracting a compliance pattern from a behavior rule pattern library, the behavior rule pattern library containing typical violation action pattern templates identified based on image sequences; matching the real-time behavior sequence with the compliance pattern, and generating a violation determination result when the real-time behavior sequence deviates from the compliance pattern.
[0040] Specifically, first, the discrete conflict resolution event data from the previous link is arranged in order according to its time stamp, and time sequence combination is performed. This process connects independent event points into a real-time behavior sequence that can reflect the continuous action of construction personnel or equipment. This sequence is a dynamic and structured description of the actual situation on site. At the same time, a pre-set behavior rule pattern library is accessed. The library stores a variety of standardized operation processes, i.e. compliance patterns, which define in detail the correct step sequence to be followed when performing a specific task. These compliance patterns are templates established based on safety specifications and operation manuals, using image sequence identification and other techniques to analyze a large number of standard operation images. Then, the real-time behavior sequence constructed in real time is matched with the compliance pattern of the corresponding task extracted from the behavior rule pattern library. This matching process can be achieved through sequence comparison algorithms such as dynamic time warping, to calculate the similarity or difference between the two sequences. When the matching result shows that the real-time behavior sequence deviates significantly from the compliance pattern, for example, a key safety step is missing or an incorrect action sequence is performed, it is determined to be a violation, and a final violation determination result is generated. The degree of deviation can be represented by the following logic: , In this expression, represents a function for calculating the difference between two sequences, such as based on edit distance or dynamic time warping algorithm. is the real-time behavior sequence constructed by time sequence combination of conflict resolution event data. is the compliance pattern related to the current construction task extracted from the behavior rule pattern library. is a pre-set threshold value representing the maximum acceptable deviation. When the calculated difference is greater than this threshold , a violation determination result is generated.
[0041] Exemplarily, it is assumed that the construction site stipulates that workers must wear safety belts before performing high-altitude work. The corresponding compliance mode in the behavior rule mode library is preset, and the event sequence thereof is "entering the work area", "wearing a safety belt", and "climbing the scaffold". The monitoring system obtains a series of conflict resolution event data about a worker through early processing. The system first performs time sequence combination on the data to construct the real-time behavior sequence of the worker, and finds that the behavior is "entering the work area" and "climbing the scaffold". Then, the system extracts the compliance mode of high-altitude work from the behavior rule mode library. When performing behavior mode matching, the system compares the real-time "entering the work area" and "climbing the scaffold" sequence with the compliance "entering the work area", "wearing a safety belt", and "climbing the scaffold" sequence. It is found through comparison that the "wearing a safety belt" key event is missing in the real-time behavior sequence. Since this deviation exceeds the preset safety threshold, the system determines that the behavior is in violation, and immediately generates a violation judgment result and sends an alarm to the management personnel.
[0042] Optionally, the generating a preliminary judgment conclusion based on the weighted feature data comprises: performing multi-modal spatio-temporal feature map convolution processing on the weighted feature data to obtain a structured event representation tensor; performing adaptive semantic inference and conflict probability evaluation on the structured event representation tensor to generate a preliminary judgment conclusion.
[0043] Specifically, first, the input weighted feature data is subjected to multi-modal spatio-temporal feature map convolution processing. In this step, the weighted feature data fused with video, audio, and sensor information is reconstructed into a spatio-temporal feature map with spatial and temporal dimensions, as shown in FIG. 1. Figure 3 Subsequently, a three-dimensional convolutional neural network or other model capable of processing spatio-temporal information is used to perform deep convolution operation on the feature map. The output of this process is a structured event representation tensor, which is a high-dimensional array containing condensed numerical representations of event dynamics and static patterns learned from the original data. This conversion process can be represented by the following formula: , wherein represents the output structured event representation tensor, is the input weighted feature data, is a deep learning network model performing multi-modal spatio-temporal feature map convolution processing. Then adaptive semantic inference and conflict probability evaluation are performed on this structured event representation tensor. The tensor is input into the subsequent fully connected layers and classifier of the neural network. Adaptive semantic inference refers to the model mapping this abstract numerical tensor to one or more semantic labels with explicit business meaning according to the learned features. At the same time, conflict probability evaluation refers to the model giving an evaluation score of internal consistency while outputting the conclusion. A preliminary judgment conclusion is finally generated, and the inference process can be represented as: , wherein represents the final generated preliminary judgment conclusion, is a specific semantic label, such as personnel not wearing a safety helmet, is a probability vector associated with the label, containing information such as confidence and conflict probability. represents the network part performing inference and evaluation functions, is the structured event representation tensor output from the previous step.
[0044] For example, the system receives a set of weighted feature data, in which the video part shows a worker moving quickly at the edge of a high-altitude platform, the audio part contains blurred shouting, and the sensor part comes from a gyroscope on the worker's safety belt, showing a sharp angle change. The set of weighted feature data is sent to the multi-modal spatio-temporal feature map convolution processing module. The model captures the rapid movement trajectory of the human form in the video, the acoustic pattern of the shouting, and the sharp turning of the gyroscope through three-dimensional convolution operations. These features are fused into a structured event representation tensor which numerically highly generalizes the concept of high, fast, and unstable. Subsequently, the tensor is sent to the inference layer , and the adaptive semantic inference function of the model interprets it as a semantic label as the preliminary judgment conclusion of the personnel high-fall risk. During conflict probability evaluation, since the multi-source evidence is highly directed to the same event, the model gives a vector containing high confidence and low conflict probability. Thus, a clear preliminary judgment conclusion is finally generated, which is the personnel high-fall risk with high confidence, for subsequent conflict resolution or alarm process.
[0045] Optionally, the generating an environment state parameter based on the environmental condition comprises: obtaining a quantitative environment index using the environmental condition; performing spatial state mapping on the quantitative environment index to generate an environment state parameter.
[0046] Specifically, first, the quantitative environment indicators are obtained by using the environmental conditions. This step is realized by deploying various environmental sensors such as illuminometers, sound level meters, temperature and humidity sensors, etc. at the construction site to measure the environmental conditions in real time. The sensors convert the collected analog signals such as light and sound into specific and continuous numerical values, for example, the light intensity is 500 lux, and the noise level is 85 decibels. These accurate numerical values are the quantitative environment indicators. Subsequently, the quantitative environment indicators are subjected to spatial state mapping to generate the final environmental state parameters. Spatial state mapping is a process of discretizing and labeling continuous numerical values, and its purpose is to induce a limited number of states with clear semantics from an infinite number of quantitative indicator values. According to the preset threshold or rule, the quantitative indicators are mapped to the corresponding state level. This mapping process can be represented by the following function: , wherein represents the output environmental state parameters, which is a set containing multiple environmental dimension states. is the input quantitative environment indicator, for example, a specific lux or decibel value. represents the function of performing spatial state mapping, which encapsulates a series of classification rules or threshold judgment logic inside, for the conversion from continuous numerical values to discrete states.
[0047] Exemplarily, assume that a construction site enters the night from the daytime. At 3 pm, the illuminometer measures the quantitative environment indicator as 15000 lux, and the sound level meter measures the indicator as 75 decibels. The system inputs these indicators into the spatial state mapping function . According to the preset rules, the illuminance intensity greater than 1000 lux is mapped to the state of sufficient light, and the noise level of 70 to 90 decibels is mapped to the state of moderate noise. Therefore, the environmental state parameters generated by the system at this time are {light state: sufficient light, noise state: moderate noise}. At 7 pm, the illuminometer measures the indicator as 50 lux, and due to the centralized operation of construction machinery at night, the sound level meter measures the indicator as 95 decibels. The system performs mapping again, and this time 50 lux is mapped to the state of insufficient light, and 95 decibels is mapped to the state of high noise. Therefore, the system updates the environmental state parameters as {light state: insufficient light, noise state: high noise}.
[0048] Optionally, the obtaining of the quantitative environment indicators by using the environmental conditions comprises: performing multi-source signal acquisition and analog-to-digital conversion on the environmental conditions to obtain digital environmental signals; performing filtering and denoising and characteristic value calculation on the digital environmental signals to obtain the quantitative environment indicators.
[0049] Specifically, first, the environmental conditions on site are collected by multi-source signals and converted into digital signals. This step uses sensor arrays deployed on site, such as photosensitive diodes for sensing light intensity and microphones for sensing sound intensity, to capture continuously changing physical quantities. These sensors convert the received physical signals such as light and sound into continuous analog electrical signals. Subsequently, these analog signals are sampled and quantized at high speed by an analog-to-digital converter, converting them into a time series composed of discrete numerical values, i.e., digital environmental signals. Next, the acquired digital environmental signals are filtered and denoised, and feature values are calculated. Since the original digital signals may contain noise and glitches caused by circuit interference or environmental mutations, a digital filter such as a moving average filter is used to smooth the signal sequence to eliminate these disturbances. After obtaining the pure signal, in order to obtain a value that can stably represent the environmental conditions in a certain period of time, feature value calculation is needed. This usually refers to calculating a statistical feature quantity such as mean, variance or root mean square within a predetermined time window. This process can be summarized as follows: , where represents the final quantitative environmental indicator, is the value of the th sampling point in the time series of the filtered and denoised digital environmental signal. is the total number of sampling points used for calculation, i.e., the size of the time window. This formula takes the calculation of the mean value as an example to show how to extract a stable feature value from a signal sequence.
[0050] Exemplarily, in order to obtain the light intensity of the construction site as a quantitative environmental indicator, the system first collects signals of environmental conditions through a light sensor. The sensor outputs an analog voltage signal proportional to the light intensity. The analog-to-digital converter samples the voltage signal at a fixed frequency to generate a string of discrete digital environmental signals. This string of original numbers may fluctuate dramatically due to the rapid movement of clouds or the instantaneous flicker of lights. In order to eliminate these influences, the system applies a moving average filter algorithm to this string of digital signals to obtain a smoother signal curve. Subsequently, the system sets a calculation window of 1 second and performs mean value calculation on the values of all filtered sampling points within the window. The calculated mean value, for example, 500, is taken as the quantitative environmental indicator of the light intensity at the current time. This value stably reflects the average light intensity in the past 1 second, rather than an extreme value at a certain moment, and is therefore more reliable.
[0051] Optionally, the time-series combination of the conflict resolution event data to construct real-time behavior sequences comprises: injecting the conflict resolution event data into spatio-temporal anchors to obtain a time-scaled monitoring event dataset; performing construction behavior concept mapping and dimension compression on the time-scaled monitoring event dataset to construct a real-time behavior sequence.
[0052] Specifically, first, the input conflict resolution event data is injected into spatio-temporal anchors to obtain a time-scaled monitoring event dataset. In this step, for each verified and discrete conflict resolution event data, accurate time and space information is attached. The time information is the timestamp of the event occurrence, and the space information is the specific location of the event occurrence, such as coordinates obtained through the global positioning system or building information model. By injecting spatio-temporal anchors, the originally independent events are given a clear spatio-temporal context, forming a data set containing [event content, time, location] information, i.e., the time-scaled monitoring event dataset. Then, the construction behavior concept mapping and dimension compression are performed on this time-scaled monitoring event dataset. Construction behavior concept mapping is the core semantic enhancement step, which maps one or more continuously occurring and logically related bottom-level monitoring events into a higher-level behavior concept with clear construction meaning according to the construction domain knowledge base. Dimension compression is to summarize and refine the information of multiple bottom-level events during this mapping process to form a more refined data unit describing the high-level behavior, thereby constructing the final real-time behavior sequence. This abstraction process can be conceptually represented by the following function: , wherein, represents the final constructed real-time behavior sequence, which is an ordered list composed of a series of high-level construction behavior concepts. is the input time-scaled monitoring event dataset. represents a complex mapping and compression function that encapsulates the logic rules of mapping, combining, and refining bottom-level spatio-temporal event points into upper-level behavior concepts.
[0053] Exemplarily, assume that the system is monitoring the hoisting process of a prefabricated wallboard and receives a series of conflict resolution event data. First, the system injects spatio-temporal anchors for these events to obtain a time-scaled monitoring event data set, which can include: event A {content: "crane hook", time: 10:01:05, location: wallboard stacking area}; event B {content: "wallboard is hoisted", time: 10:01:30, location: wallboard stacking area}; event C {content: "wallboard moves to installation point", time: 10:02:15, location: floor X, axis Y}; event D {content: "wallboard installation is completed", time: 10:03:00, location: floor X, axis Y}. Next, the system performs construction behavior concept mapping and dimension compression on the data set. According to the definition in the knowledge base, the system identifies that events A, B, C, and D are continuous and logically related in space and time, and collectively constitute a complete wallboard hoisting operation. Therefore, the system maps the four underlying events to a high-level behavior concept: performing wallboard hoisting operation. In the dimension compression process, the system takes the time of event A as the start time of the behavior, takes the time of event D as the end time, and generalizes the location information as "from the stacking area to the installation point". Finally, the four underlying events are compressed and refined to form an element in the real-time behavior sequence, i.e. [behavior concept: "performing wallboard hoisting operation", start and end time: 10:01:05-10:03:00].
[0054] Based on the same inventive concept, as shown in Figure 4 The present application also provides a multi-modal data fusion construction operation violation automatic monitoring system, which comprises: a data acquisition and processing module for acquiring multi-modal source data of a target scene, performing data recognition and standardization processing on the multi-modal source data, and generating pre-processed feature data; an environment perception module for acquiring environmental conditions on site and generating environmental state parameters based on the environmental conditions; a weight allocation module for performing dynamic weight allocation on the pre-processed feature data based on the environmental state parameters to generate weighted feature data; a conflict resolution module for performing multi-source data conflict judgment and resolution on the weighted feature data to generate conflict resolution event data; a behavior pattern analysis module for performing behavior pattern matching on the conflict resolution event data based on a pre-set behavior rule pattern library to generate a violation determination result.
[0055] It should be noted that the electrical connection between the various units described above does not necessarily indicate a direct connection, and the indirect connection mode can also be applied to the embodiments of the present application as long as the purpose of the present application is achieved. The above is only an exemplary embodiment of the present application, and cannot limit the scope of the present application.
[0056] That is, any equivalent changes and modifications made in accordance with the teachings of the present application are still within the scope of the present application. Other embodiments of the present application will be readily apparent to those skilled in the art upon considering the description and practice of the principles disclosed herein. The present application is intended to cover any variations, uses, or adaptive changes of the present application that follow the general principles of the present application and include common knowledge or conventional techniques in the art that are not described in the present application.
Claims
1. A method for automatic monitoring of violations in infrastructure construction operations based on multimodal data fusion, characterized in that, The method includes: Acquire multimodal source data of the target scene, perform data identification and standardization processing on the multimodal source data, and generate preprocessed feature data; Obtain the environmental conditions at the site, and generate environmental state parameters based on the environmental conditions; Based on the environmental state parameters, the preprocessed feature data is dynamically weighted to generate weighted feature data; The weighted feature data is subjected to multi-source data conflict judgment and resolution to generate conflict resolution event data; Based on a preset behavioral rule pattern library, behavioral pattern matching is performed on the conflict resolution event data to generate violation judgment results.
2. The method for automatic monitoring of violations in infrastructure construction operations based on multimodal data fusion according to claim 1, characterized in that, The step of performing data identification and standardization processing on the multimodal source data to generate preprocessed feature data includes: Keyframe extraction and target region localization are performed on the multimodal source data to generate video feature data; The multimodal source data is subjected to noise reduction and acoustic feature extraction to generate audio feature data; The multimodal source data is subjected to frequency domain transformation and feature value extraction to generate sensor feature data; The video feature data, audio feature data, and sensor feature data are combined to generate preprocessed feature data.
3. The method for automatic monitoring of violations in infrastructure construction operations based on multimodal data fusion according to claim 2, characterized in that, The step of dynamically assigning weights to the preprocessed feature data based on the environmental state parameters to generate weighted feature data includes: Obtain the light intensity and noise level of the environmental state parameters; The weights of the video feature data are adjusted according to the light intensity, and the weights of the audio feature data are adjusted according to the noise level to generate dynamic weight coefficients. The preprocessed feature data is weighted using the dynamic weighting coefficients to generate weighted feature data.
4. The method for automatic monitoring of violations in infrastructure construction operations based on multimodal data fusion according to claim 1, characterized in that, The step of performing multi-source data conflict judgment and resolution on the weighted feature data to generate conflict resolution event data includes: Based on the weighted feature data, a preliminary judgment conclusion is generated; Detect logical conflicts between the preliminary judgment conclusions and generate conflict indicators; When the conflict identifier is received, the operating status data of the construction equipment is obtained; The preliminary judgment conclusion is verified by arbitration using the operational status data, and conflict resolution event data is generated.
5. The method for automatic monitoring of violations in infrastructure construction operations based on multimodal data fusion according to claim 1, characterized in that, The method of matching behavioral patterns in the conflict resolution event data based on a preset behavioral rule pattern library to generate violation determination results includes: The conflict resolution event data are combined in a time sequence to construct a real-time behavior sequence; Compliance patterns are extracted from the behavior rule pattern library, which contains templates of typical violation behavior patterns based on image sequence recognition; The real-time behavior sequence is matched with the compliance pattern, and a violation determination result is generated when the real-time behavior sequence deviates from the compliance pattern.
6. The method for automatic monitoring of violations in infrastructure construction operations based on multimodal data fusion according to claim 4, characterized in that, The preliminary judgment conclusion generated based on the weighted feature data includes: The weighted feature data is subjected to multimodal spatiotemporal feature map convolution processing to obtain a structured event representation tensor; Adaptive semantic inference and conflict probability assessment are performed on the structured event representation tensor to generate a preliminary judgment conclusion.
7. The method for automatic monitoring of violations in infrastructure construction operations based on multimodal data fusion according to claim 1, characterized in that, The generation of environmental state parameters based on the environmental conditions includes: Using the aforementioned environmental conditions, quantitative environmental indicators are obtained; Spatial state mapping is performed on the quantified environmental indicators to generate environmental state parameters.
8. The method for automatic monitoring of violations in infrastructure construction operations based on multimodal data fusion according to claim 7, characterized in that, The process of obtaining quantitative environmental indicators using the aforementioned environmental conditions includes: The environmental conditions are subjected to multi-source signal acquisition and analog-to-digital conversion to obtain digital environmental signals; The digital environmental signal is filtered and denoised, and its eigenvalues are calculated to obtain quantified environmental indicators.
9. The method for automatic monitoring of violations in infrastructure construction operations based on multimodal data fusion according to claim 5, characterized in that, The step of combining the conflict resolution event data in a time sequence to construct a real-time behavior sequence includes: The conflict resolution event data is injected into the spatiotemporal anchor point to obtain a time-stamped monitoring event dataset; The construction behavior concept mapping and dimensional compression are performed on the time-stamped monitoring event dataset to construct a real-time behavior sequence.
10. A multimodal data fusion-based automatic monitoring system for violations in infrastructure construction operations, applied to the multimodal data fusion-based automatic monitoring method for violations in infrastructure construction operations as described in any one of claims 1-9, characterized in that, The system includes: The data acquisition and processing module is used to acquire multimodal source data of the target scene, perform data identification and standardization processing on the multimodal source data, and generate preprocessed feature data; An environmental perception module is used to acquire on-site environmental conditions and generate environmental state parameters based on those conditions. The weight allocation module is used to dynamically allocate weights to the preprocessed feature data based on the environmental state parameters, and generate weighted feature data. The conflict resolution module is used to perform multi-source data conflict judgment and resolution on the weighted feature data, and generate conflict resolution event data; The behavior pattern analysis module is used to perform behavior pattern matching on the conflict resolution event data based on a preset behavior rule pattern library, and generate violation judgment results.
Citation Information
Patent Citations
Multi-sensor data fusion method based on reinforcement learning and D-S evidence theory
CN113283516A
Power construction safety management method, device, equipment and medium
CN119379014A
Event detection method and device based on multiple modes, electronic equipment and storage medium
CN119989258A
Three-dimensional automatic modeling method and system based on AI large model technology
CN120088409A
Dynamic calibration method and system of vehicle-mounted emotion recognition system
CN120448917A
Cited By
Deep learning-fused exploration scene monitoring illegal behavior automatic identification method and system
CN121659076A
Survey scene monitoring violation behavior automatic identification method and system fusing deep learning
CN121659076B
Visual construction management method and system
CN122176245A
A method and system for visualizing construction management
CN122176245B
Intelligent construction safety log filling and reporting system
CN122264315A