Machine vision-based deep foundation pit construction safety early warning method
By using spatiotemporal alignment of multi-source heterogeneous data and causal inference graph models, combined with video micro-vibration characteristics and physical constraints, the problems of data independence and environmental interference in the safety monitoring of deep foundation pit construction were solved. This enabled comprehensive and refined monitoring and intelligent hierarchical early warning of deep foundation pit support structures, improving the efficiency and accuracy of early warning.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SAIKEN ENGINEERING TECHNOLOGY (CHANGZHOU) CO LTD
- Filing Date
- 2026-04-08
- Publication Date
- 2026-07-03
AI Technical Summary
Existing methods for monitoring the safety of deep foundation pit construction suffer from problems such as independent data sources, susceptibility to dynamic environmental interference, simple early warning logic, and inability to reveal the causal relationship of risks, resulting in insufficient monitoring continuity and early warning accuracy.
By combining spatiotemporal alignment of multi-source heterogeneous data, extraction of video micro-vibration features, and three-dimensional reconstruction under physical constraints, and solving through a causal inference graph model, the root causes and transmission paths of risks are identified, enabling in-depth analysis and intelligent hierarchical early warning.
It enables comprehensive and refined monitoring of the dynamic response of deep foundation pit support structures, improves early warning efficiency and accuracy, can detect structural anomalies at an early stage, and provides in-depth risk causal analysis support.
Smart Images

Figure CN122333167A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer vision and safety monitoring technology, and in particular to a machine vision-based method for safety early warning in deep foundation pit construction. Background Technology
[0002] Deep foundation pit engineering is a crucial link in urban underground space development. Its construction process involves complex environments and numerous uncertainties, and the safety of the support structure directly affects the safety of the entire project and its surrounding environment. Machine vision technology, as a core component of image data processing and pattern recognition, is increasingly being applied in civil engineering to monitor and analyze construction sites non-contactly, thereby improving safety management.
[0003] Currently, safety monitoring in deep foundation pit construction primarily relies on embedding or installing various physical sensors, such as displacement gauges, inclinometers, strain gauges, and vibration sensors, in key structural components to obtain point-like data on structural deformation and internal forces. Simultaneously, some machine vision-based applications are emerging, such as using cameras to monitor whether construction workers are wearing safety helmets, using image recognition technology to monitor the unauthorized operation of large machinery, and employing digital image correlation to measure large deformations on structural surfaces. These methods have improved the efficiency of safety management to some extent.
[0004] However, existing technologies have several shortcomings. First, various monitoring data sources are independent of each other; video surveillance data, sensor data, and personnel positioning data are usually processed in different systems, lacking effective fusion under a unified spatiotemporal benchmark, making it difficult to form a comprehensive and integrated understanding of construction safety risks. Second, machine vision-based monitoring methods are highly susceptible to interference from the dynamic environment of the construction site; for example, moving machinery and personnel frequently obstruct the field of view, leading to the loss of key data and compromising the continuity and reliability of monitoring. Finally, most existing early warning systems are based on threshold judgments of single monitoring indicators, with simple early warning logic that fails to reveal the complex causal relationships behind risk events. When an alarm is triggered, managers find it difficult to quickly locate the root cause and transmission path of the risk. Summary of the Invention
[0005] To address the aforementioned issues, this invention provides a machine vision-based method for early warning of safety during deep foundation pit construction. It employs a combination of multi-source heterogeneous data spatiotemporal alignment, video micro-vibration feature extraction, and 3D reconstruction under physical constraints. The data is then input into a causal reasoning graph model for calculation. This method accurately identifies the root causes and transmission paths of risks, enabling in-depth analysis and intelligent hierarchical early warning of safety risks during deep foundation pit construction.
[0006] The above objectives can be achieved through the following approach:
[0007] A machine vision-based safety early warning method for deep foundation pit construction includes: acquiring heterogeneous data, including multi-view video streams, sensor data, and positioning data of the construction area, and performing spatiotemporal alignment based on a unified time axis and spatial coordinate system to generate multi-source synchronous monitoring data; performing temporal domain amplification processing on the multi-view video streams to make pixel-level micro-vibration signals explicit, and generating an initial spectrum map characterizing the frequency and energy distribution of each pixel through time-frequency transformation, mapping it to a reference three-dimensional structural model to assign spatial attributes, and extracting the vibration spectrum features of the support structure surface; identifying personnel and equipment targets in the multi-view video streams and positioning data and tracking dynamic behavior features, matching the dynamic behavior features with a preset high-risk behavior rule library to identify high-risk behavior events and their three-dimensional locations, and determining the local area affected by the event based on the topological model of the mechanical transmission relationship of the support structure. The region is designated as a high-risk area of concern. Vibration spectrum features within this region are extracted for focused analysis, and the energy change gradient and dominant frequency offset within a time window are calculated to generate regional risk vibration features. These vibration spectrum features are then incorporated into a three-dimensional geometric reconstruction algorithm as physical continuity constraints to fill in data gaps caused by occlusion and generate three-dimensional structural features. The regional risk vibration features, three-dimensional structural features, and sensor data are aligned and correlated under a unified spatiotemporal reference to generate spatiotemporally aligned fusion features. These spatiotemporally aligned fusion features are input into a pre-defined causal inference graph model. This model includes behavioral nodes representing external disturbances, structural state nodes representing structural responses, and physical signal nodes representing physical measurements. By solving the causal paths between nodes, the causal risk chain with the highest confidence is identified, and graded early warning information is generated and output accordingly.
[0008] Optionally, acquiring heterogeneous data, including multi-view video streams, sensor data, and positioning data of the construction area, and performing spatiotemporal alignment based on a unified time axis and spatial coordinate system, includes: synchronously acquiring high-definition video streams of the construction area using several non-collinearly deployed industrial cameras to form multi-view video streams; periodically polling vibration sensors and displacement monitoring devices deployed at key stress nodes of the support structure to acquire sensor data containing dynamic changes in structural physical quantities; receiving wireless signals emitted by positioning tags installed on construction personnel and construction machinery, calculating three-dimensional coordinate information to form positioning data; using a central server as a time reference source, applying a unified timestamp to each frame of the multi-view video stream, each set of sensor data readings, and each coordinate point of the positioning data, and converting the heterogeneous data to a unified three-dimensional spatial coordinate system of the foundation pit based on calibrated camera parameters to generate multi-source synchronous monitoring data.
[0009] Optionally, the multi-view video stream is subjected to temporal amplification to make the pixel-level micro-vibration signal explicit, and an initial spectrum map characterizing the frequency and energy distribution of each pixel is generated by time-frequency transformation. This spectrum map is then mapped to a reference three-dimensional structural model to give it spatial attributes. This includes: using a temporal filter to amplify the amplitude of pixel brightness fluctuations within a specific frequency range in the multi-view video stream to make the pixel-level micro-vibration signal on the support structure surface explicit; performing a short-time Fourier transform on the pixel-level micro-vibration signal of each pixel to generate an initial spectrum map characterizing the frequency components and energy distribution of the pixel within different time windows; mapping the two-dimensional pixels in the initial spectrum map to three-dimensional spatial coordinates on the reference three-dimensional structural model using ray projection based on calibrated camera parameters; and weighting and fusing multiple spectral information corresponding to the same point on the reference three-dimensional structural model that is simultaneously covered by multiple views to generate a vibration spectrum feature with three-dimensional spatial attributes.
[0010] Optionally, matching the dynamic behavioral features with a preset high-risk behavior rule base to identify high-risk behavioral events and their three-dimensional locations, and determining the local area affected by the event as a high-risk area based on the topological model of the mechanical transmission relationship of the support structure, includes: comparing the target trajectory and velocity information in the dynamic behavioral features with the behavior types in the rule base to identify high-risk behavioral events and their three-dimensional locations; retrieving the support structure topological model representing the physical connection and mechanical transmission relationship between each support component, centered on the three-dimensional location; determining an initial range of concern within a preset influence radius, and dynamically expanding the boundary of the initial range of concern according to the duration and intensity of the high-risk behavioral event to form a high-risk area.
[0011] Optionally, focusing the analysis on the vibration spectrum characteristics within the high-risk area and calculating the energy change gradient and dominant frequency offset within a time window includes: extracting corresponding local spectrum data from the vibration spectrum characteristics based on the three-dimensional spatial coordinate range of the high-risk area; performing bandpass filtering on the local spectrum data to filter out noise frequency components generated by construction background activities and retain preset frequency components associated with the structural response; calculating the first derivative of the filtered local spectrum data with time within a sliding time window to obtain the energy change gradient characterizing the instantaneous growth rate of vibration energy; identifying the frequency corresponding to the energy peak within the sliding time window and calculating the offset value relative to the steady-state reference frequency to obtain the dominant frequency offset characterizing the change in structural response characteristics; and combining the energy change gradient and the dominant frequency offset to generate regional risk vibration characteristics characterizing the quantitative index of regional physical risk.
[0012] Optionally, the vibration spectrum features are introduced as physical continuity constraints into the 3D geometric reconstruction algorithm to complete the data missing regions caused by occlusion and generate 3D structural features. Aligning and associating the regional risk vibration features, 3D structural features, and sensor data under a unified spatiotemporal reference includes: identifying data missing regions in the reference 3D structural model that cannot be directly observed due to dynamic object occlusion; extracting the spectral continuity information of the vibration spectrum features at the boundary of the data missing regions, and quantifying the spectral continuity information into spatial continuity constraints, which are added as physical constraint terms to the optimization objective function of the 3D geometric reconstruction algorithm; iteratively solving the optimization objective function, using the continuity of physical vibration information to guide the geometric completion of the data missing regions, generating completed 3D structural features; and under a unified 3D spatial coordinate system and time axis, using the structural topology nodes within the high-risk region as units, associating each node with the regional risk vibration features, the completed 3D structural features, and the sensor data to generate a spatiotemporally aligned fusion feature matrix.
[0013] Optionally, the spatiotemporally aligned fusion features are input into a preset causal inference graph model. This model includes behavioral nodes representing external disturbances, structural state nodes representing structural responses, and physical signal nodes representing physical measurements. Identifying the causal risk chain with the highest confidence by solving the causal paths between nodes includes: updating the state occurrence probability of each node in the causal inference graph model using the spatiotemporally aligned fusion feature matrix; traversing all candidate risk transmission paths starting from the activated behavioral nodes and ending at the physical signal nodes; calculating the confidence of each candidate risk transmission path, where the confidence is obtained by multiplying the state value of the starting behavioral node of the path with the preset weights of each causal relationship edge on the path; selecting the path with the highest confidence as the causal risk chain, and parsing out the risk root node and transmission path sequence from it.
[0014] Optionally, the method further includes: collecting historically output graded early warning information and corresponding on-site handling feedback to construct a case library containing positive and negative sample cases; extracting and analyzing the causal risk chains corresponding to false alarm cases and missed alarm cases in the case library; and iteratively optimizing the early warning model by adjusting the weights of each causal relationship edge in the causal inference graph model, or by adjusting the guiding weights of the vibration spectrum features in the three-dimensional geometric reconstruction algorithm for filling in missing data regions, based on the analysis results.
[0015] Optionally, generating and outputting graded early warning information includes: pre-setting a set of confidence thresholds associated with risk levels; comparing the confidence of the identified causal risk chain with the set of confidence thresholds, and determining the final risk level through a step-by-step mapping; extracting the corresponding risk root cause information and transmission path information from the causal risk chain, and matching the corresponding handling suggestions from a preset emergency plan database; integrating the final risk level, risk root cause information, transmission path information, and handling suggestions to generate structured graded early warning information and push it to a preset terminal.
[0016] Based on the same inventive concept, this invention also provides a machine vision-based deep foundation pit construction safety early warning system. The system includes: a data acquisition and alignment module, used to acquire heterogeneous data including multi-view video streams, sensor data, and positioning data from the construction area, and perform spatiotemporal alignment based on a unified time axis and spatial coordinate system to generate multi-source synchronous monitoring data; a spectrum feature extraction module, used to perform temporal domain amplification processing on the multi-view video streams to make pixel-level micro-vibration signals explicit, and generate an initial spectrum map characterizing the frequency and energy distribution of each pixel point through time-frequency transformation, mapping it to a reference three-dimensional structural model to assign spatial attributes, and extracting the vibration spectrum features of the support structure surface; and a risk area delineation module, used to identify personnel and equipment targets in the multi-view video streams and positioning data and track dynamic behavior characteristics, matching the dynamic behavior characteristics with a preset high-risk behavior rule library to identify high-risk behavior events and their three-dimensional locations, and determining the affected areas based on the topological model of the mechanical transmission relationship of the support structure. The affected local area is designated as a high-risk area. A spectrum focusing analysis module is used to extract vibration spectrum features within the high-risk area for focused analysis, and calculate the energy change gradient and dominant frequency offset within a time window to generate regional risk vibration features. A fusion and geometric reconstruction module is used to introduce the vibration spectrum features as physical continuity constraints into a three-dimensional geometric reconstruction algorithm, filling in data gaps caused by occlusion and generating three-dimensional structural features. The regional risk vibration features, three-dimensional structural features, and sensor data are aligned and correlated under a unified spatiotemporal reference to generate spatiotemporally aligned fused features. A causal path calculation and early warning generation module is used to input the spatiotemporally aligned fused features into a preset causal inference graph model. This model includes behavioral nodes representing external disturbances, structural state nodes representing structural responses, and physical signal nodes representing physical measurements. By calculating the causal paths between nodes, the causal risk chain with the highest confidence is identified, and graded early warning information is generated and output.
[0017] Compared with the prior art, the present invention has the following advantages:
[0018] This invention enables comprehensive and refined monitoring of the dynamic response of deep foundation pit support structures. By employing time-domain amplification technology to extract pixel-level micro-vibration signals from multi-view video streams, this method can acquire a high-density dynamic response field covering the entire structural surface in a non-contact manner. This overcomes the limitations of traditional discrete point-based sensors in terms of spatial coverage and information dimension, thus enabling earlier and more comprehensive detection of structural anomalies.
[0019] This invention proposes a targeted analysis mechanism that correlates external disturbances with internal structural responses, improving the efficiency and accuracy of early warning. By identifying high-risk behaviors of personnel and equipment at the construction site and combining this with a mechanical transfer topology model of the support structure to delineate high-risk areas, this method can focus computational resources on the most likely localized areas of danger, avoiding the blind processing of massive amounts of global data and realizing a shift from passive monitoring to proactive risk tracing.
[0020] This invention enhances the robustness and data integrity of visual monitoring by introducing physical information constraints. By incorporating the continuity of vibration spectrum characteristics as a physical constraint into the 3D geometric reconstruction algorithm, it can effectively fill in dynamic occlusion areas caused by the movement of construction equipment or personnel. This solves the problem of data loss in complex dynamic scenarios using pure visual methods, ensuring the spatiotemporal continuity and reliability of monitoring data, and providing a high-quality data foundation for subsequent fusion analysis.
[0021] This invention constructs a deep risk diagnosis and early warning model based on causal reasoning. Unlike traditional simple alarms based on threshold triggers, this method, by solving a causal reasoning graph model, can identify and present the complete causal risk chain from the risk source behavior to the final physical signal manifestation. It can not only determine whether the risk exists and its level, but also clearly explain the cause and transmission path of the risk, providing managers with profound and interpretable decision support, and realizing the knowledge-based and intelligent nature of early warning information.
[0022] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures pointed out in the description, claims, and drawings. Attached Figure Description
[0023] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0024] Figure 1 This is a flowchart illustrating a machine vision-based safety early warning method for deep foundation pit construction according to an embodiment of the present invention.
[0025] Figure 2 This is a comparison curve of the brightness fluctuation signal of the surface pixels of the support structure before and after time-domain amplification processing according to an embodiment of the present invention.
[0026] Figure 3 This is a vibration energy comparison curve on the cross-section of the embodiment of the present invention, which uses vibration spectrum continuity constraints to guide the three-dimensional geometric completion of the data missing area under occlusion.
[0027] Figure 4 This is a schematic diagram of a machine vision-based safety early warning system for deep foundation pit construction according to an embodiment of the present invention. Detailed Implementation
[0028] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0029] Reference Figure 1 One embodiment of the present invention proposes a machine vision-based method for early warning of safety in deep foundation pit construction. It adopts a combination of spatiotemporal alignment of multi-source heterogeneous data, extraction of video micro-vibration features and three-dimensional reconstruction under physical constraints, and inputs them into a causal reasoning graph model for calculation. This method can accurately identify the root causes and transmission paths of risks, and realize in-depth analysis and intelligent hierarchical early warning of safety risks in deep foundation pit construction.
[0030] The method described in this embodiment specifically includes:
[0031] It acquires heterogeneous data, including multi-view video streams, sensor data, and positioning data from the construction area, and performs spatiotemporal alignment based on a unified time axis and spatial coordinate system to generate multi-source synchronous monitoring data.
[0032] The multi-view video stream is subjected to temporal domain amplification to make the pixel-level micro-vibration signal explicit, and an initial spectrum diagram characterizing the frequency and energy distribution of each pixel is generated by time-frequency transformation. This spectrum diagram is then mapped onto a reference three-dimensional structural model to give it spatial attributes, and the vibration spectrum characteristics of the support structure surface are extracted.
[0033] By identifying personnel and equipment targets in multi-view video streams and positioning data and tracking dynamic behavior characteristics, the dynamic behavior characteristics are matched with a preset high-risk behavior rule base to identify high-risk behavior events and their three-dimensional locations. Based on the topological model of the mechanical transmission relationship of the support structure, the local area affected by the event is determined as a high-risk area of concern.
[0034] Vibration spectrum characteristics within the high-risk area are extracted for focused analysis, and the energy change gradient and dominant frequency shift within the time window are calculated to generate regional risk vibration characteristics.
[0035] The vibration spectrum features are introduced into the three-dimensional geometric reconstruction algorithm as physical continuity constraints to fill in the data missing areas caused by occlusion and generate three-dimensional structural features. The regional risk vibration features, three-dimensional structural features and sensor data are aligned and correlated under a unified spatiotemporal reference to generate spatiotemporally aligned fusion features.
[0036] The spatiotemporally aligned fusion features are input into a preset causal reasoning graph model. This model includes behavioral nodes representing external disturbances, structural state nodes representing structural responses, and physical signal nodes representing physical measurements. By solving the causal paths between each node, the causal risk chain with the highest confidence is identified, and hierarchical early warning information is generated and output accordingly.
[0037] Specifically, this invention proposes a machine vision-based safety early warning method for deep foundation pit construction. The core principle is to construct a closed-loop information processing framework from multi-source heterogeneous data perception to causal reasoning and decision-making. First, through spatiotemporal alignment technology, data from different modalities, such as multi-view video, sensor physical quantities, and personnel and equipment positioning, are unified under a common temporal and spatial reference, forming a synchronized monitoring dataset. Next, innovatively utilizing temporal amplification and time-frequency transformation techniques, minute vibrations on the surface of the support structure are extracted and quantified non-contactly from the video stream as high-density, spatialized dynamic response features. Simultaneously, through behavior recognition and a mechanical topology model, external disturbances generated by construction activities are accurately mapped to high-risk areas on the structure. This method further uses the extracted vibration features as physical continuity constraints to guide a three-dimensional geometric reconstruction algorithm to complete visually occluded areas, ensuring the integrity of structural features. Finally, the multimodal fusion features, including behavior, vibration, geometry, and sensor readings, are input into a pre-defined causal inference graph model. By solving the causal transmission path from external disturbance behavior nodes to internal structural state nodes, and then to observable physical signal nodes, the risk chain with the highest confidence is identified, thereby achieving the diagnosis and root cause tracing of safety risks.
[0038] Optionally, acquiring heterogeneous data, including multi-view video streams, sensor data, and positioning data of the construction area, and performing spatiotemporal alignment based on a unified time axis and spatial coordinate system, includes:
[0039] High-definition video streams of the construction area are simultaneously acquired by several non-co-linearly deployed industrial cameras to form a multi-view video stream;
[0040] The vibration sensors and displacement monitoring devices deployed at key stress nodes of the support structure are periodically polled to obtain sensor data containing dynamic changes in the physical quantities of the structure.
[0041] It receives wireless signals transmitted by positioning tags installed on construction workers and construction machinery, calculates three-dimensional coordinate information, and forms positioning data;
[0042] Using the central server as the time reference source, a unified timestamp is applied to each frame of the multi-view video stream, each set of readings of the sensor data, and each coordinate point of the positioning data. Based on the calibrated camera parameters, the heterogeneous data is converted to a unified three-dimensional spatial coordinate system of the foundation pit to generate multi-source synchronous monitoring data.
[0043] Specifically, for time synchronization, a synchronization network is constructed with an industrial switch supporting the IEEE 1588 Precision Time Protocol (PTP) as the master clock. All industrial cameras, sensor data acquisition (DAQ) modules, and ultra-wideband (UWB) positioning base stations are connected to this network to achieve microsecond-level clock synchronization. To further eliminate differences in camera exposure times, a synchronization signal generator with its clock phase-locked with the PTP master clock is added. The generated hardware trigger pulses are simultaneously sent to all cameras to ensure that each frame of the multi-view video stream is strictly synchronized at the moment of capture. Sensor data is synchronously sampled at a fixed frequency by DAQ modules supporting PTP, while the positioning coordinates of personnel and machinery are calculated by the UWB base station with synchronized timestamps based on the Time Difference of Arrival (TDOA) algorithm. For spatial alignment, a unified three-dimensional world coordinate system for the foundation pit is first established. For example, the design corner point A of the foundation pit is used as the origin O(0,0,0), with east, north, and vertically upward as the X, Y, and Z axes, respectively. The position parameters of all sensing devices in this coordinate system were precisely calibrated using a total station: the intrinsic parameter matrix K, lens distortion coefficient, and extrinsic parameter matrix of each camera were obtained using the Zhang Zhengyou calibration method. Furthermore, the parameter matrix defines the rotation from the world coordinate system to the camera coordinate system. With translation That is, satisfying ,in As a world coordinate system, For camera coordinates; directly measure and record the world coordinates of each vibration sensor and displacement monitoring device installation point. , The term "sensor" refers to vibration sensors and displacement monitoring devices deployed at key stress nodes of the support structure; and it accurately measures the coordinates of each UWB positioning base station. Used for location calculation. The term "base station" refers to the positioning base station in a UWB positioning system. During data processing, the central server appends a PTP timestamp, recorded at the time of its generation, to each received data unit. For video data, this applies to each pixel. Using camera intrinsic parameters Distortion is removed by adjusting the distortion coefficients to obtain normalized coordinates. , and These represent the horizontal and vertical indices of the image pixel coordinate system, respectively. and This represents the normalized image coordinates of a pixel after distortion correction. It retrieves the depth value corresponding to that pixel. Depth value The following methods can be used to obtain the pixel's coordinates: (1) Based on a calibrated synchronous multi-view image, calculate the disparity using a multi-view stereo vision algorithm (such as the PatchMatch algorithm) and convert it; (2) From a high-precision point cloud reference model of the foundation pit obtained in advance by laser scanning, find the distance to the surface point closest to the pixel's line of sight. Calculate the coordinates of the pixel in the world coordinate system using the formula. :
[0044] ;
[0045] The spatial location of sensor data is determined by its preset installation coordinates. The positioning coordinates are directly defined, and during the calculation, they are converted from the calibrated base station coordinates to values in the world coordinate system. Finally, the precise timestamp of each video frame is used. Reference time Retrieve all timestamps satisfy The sensor readings and positioning coordinates, where Using a preset time tolerance threshold, such as 1 millisecond, these data are combined with the current video frame to generate a structured multi-source synchronous monitoring data packet. The data packet contains a base timestamp. The process involves generating a set of multi-view image frames, an array of sensor readings, and an array of dynamic target coordinates. Through these steps, a high-precision data input with a strictly consistent spatiotemporal reference is generated, which can be directly used for subsequent feature extraction and fusion analysis.
[0046] For example, a world coordinate system is established with the design corner point A of the foundation pit as the origin. The coordinates of the location where vibration sensor No. 1 is installed on the support pile are measured by a total station. The system's central server, acting as the PTP master clock, sends synchronization signals to all sensing devices, capturing a high-definition image frame at the reference time of 10:00:00.500. The camera's intrinsic parameter matrix has been calibrated. Given the rotation matrix in the extrinsic parameter matrix The reverse With translation vector Confirmed. For a specific pixel on the surface of the support pile in the image. After distortion correction, the normalized coordinates are: The depth value of the point is calculated using a multi-view stereo vision algorithm. The value is 15.0 meters. Substituting this into the coordinate transformation formula, the calculation is as follows: The spatial coordinates of the pixel in the world coordinate system are calculated as follows: At this time, the central server retrieves the reading of sensor 1 (2.5 mm / s²) within the timestamp of 10:00:00.500±1 milliseconds, and the excavator coordinates calculated by the UWB positioning system. The aforementioned visual spatial coordinates, sensor physical quantities, and positioning data are encapsulated to generate a spatiotemporally consistent multi-source synchronous monitoring data packet.
[0047] Optionally, the multi-view video stream is subjected to temporal domain magnification to make pixel-level micro-vibration signals explicit, and an initial spectrum representing the frequency and energy distribution of each pixel is generated through time-frequency transformation. This spectrum is then mapped onto a reference three-dimensional structural model to assign spatial properties, including:
[0048] A temporal filter is used to enhance the amplitude of pixel brightness fluctuations within a specific frequency range in a multi-view video stream, thereby making the pixel-level micro-vibration signals on the surface of the support structure explicit.
[0049] A short-time Fourier transform is performed on the pixel-level micro-vibration signal of each pixel to generate an initial spectrum that characterizes the frequency components and energy distribution of the pixel in different time windows.
[0050] Based on the calibrated camera parameters, the two-dimensional pixels in the initial spectrogram are mapped to the three-dimensional spatial coordinates on the reference three-dimensional structural model by ray casting.
[0051] For the same point on the reference three-dimensional structural model that is simultaneously covered by multiple perspectives, the corresponding multiple spectral information is weighted and fused to generate vibration spectral features with three-dimensional spatial attributes.
[0052] Specifically, the phase differential amplification algorithm is first used to enhance the amplitude of pixel brightness or phase fluctuations within a specific frequency range in a multi-view video stream. This algorithm extracts the phase information of pixels by performing complex manipulable pyramid decomposition on the video sequence and then performs temporal bandpass filtering, thereby making the minute displacement signals of the support structure surface, which are invisible to the naked eye, explicit. For the enhanced video stream, the brightness fluctuation signal of each pixel in the temporal domain is then analyzed. Perform a short-time Fourier transform to generate an initial spectrum representing the frequency components and energy distribution of the pixel within different time windows. The calculation formula is as follows:
[0053] ;
[0054] in, The coordinates after time-domain magnification are The pixels in The brightness value at any given moment is derived from an image sequence acquired by an industrial camera and then magnified. For The window function centered is used to extract the time slices required for calculation, and its parameters are preset according to the monitoring frequency requirements. The vibration frequency; That is, the pixel point in The power spectral density at time t reflects the vibration energy distribution on the structure's surface. Subsequently, based on the calibrated camera intrinsic parameter matrix... and extrinsic parameter matrix Two-dimensional pixel coordinates are projected using ray casting. Convert the ray into a 3D spatial ray and calculate the coordinates of the intersection point between the ray and the surface of the reference 3D structural model. This enables reprojection from the image plane to physical space. The world coordinate system is used to refer to physical quantities within a uniformly established three-dimensional world coordinate system for the foundation pit. For the same point on the baseline three-dimensional structural model that is simultaneously covered by multiple viewpoints, a weighted fusion strategy is employed to generate vibration spectrum features with three-dimensional spatial attributes. Its fusion formula is expressed as:
[0055] ;
[0056] in, To observe this point simultaneously The number of cameras; For the first The initial spectral data calculated by each camera at the projected pixel of this point; The weighting factors are set based on the observation angle and distance, and satisfy the following conditions: Weight The value is derived from the cosine of the angle between the camera's optical axis and the normal vector at that point, as well as the reciprocal of the shooting distance. In this way, pure visual pixel fluctuations are transformed into a vibration field on the surface of the support structure that possesses physical spatial location information, enabling precise capture of early micro-vibration responses caused by internal stress variations. For example... Figure 2 As shown, the comparison of temporal vibration signals before and after enhancement is illustrated. The solid line represents the original pixel brightness signal, and the dashed line represents the vibration signal that has been made explicit after temporal amplification processing according to this invention.
[0057] For example, a phase differential amplification algorithm is used to process a high-definition video stream of the support pile surface. The passband frequency of the time-domain bandpass filter is set to 10-30Hz, amplifying the minute displacements (0.1 pixel level) invisible to the naked eye within this frequency band by 50 times, making the micro-vibration characteristics of the support pile surface visible in the image. For a specific pixel after enhancement... In its time-domain brightness fluctuation signal Apply a window with a duration of 2 seconds. By performing the short-time Fourier transform formula Perform the solution, assuming in At time 15Hz, the power spectral density value of the pixel was calculated. This generates the initial spectrogram of the pixel. Subsequently, based on the calibrated camera parameters, the pixel is located using ray casting. Spatial points on the corresponding reference three-dimensional structural model This point was observed by both cameras 1 and 2. The spectral energy calculated by camera 1 is known. Its observation weighting factor The distance is relatively short; the spectral energy calculated by camera 2 Observation weighting factor Substituting into the weighted fusion formula calculate: Ultimately, this results in three-dimensional coordinate points. A vibration spectrum characteristic value of 0.81 is assigned.
[0058] Optionally, the dynamic behavioral characteristics are matched with a preset high-risk behavior rule base to identify high-risk behavioral events and their three-dimensional locations. Based on the topological model of the mechanical transmission relationship of the support structure, the local areas affected by the events are determined as high-risk areas of concern, including:
[0059] The target trajectory and speed information in the dynamic behavior features are compared with the behavior types in the rule base to identify high-risk behavior events and their three-dimensional locations.
[0060] Centered on the three-dimensional location of occurrence, a topological model of the support structure representing the physical connection and mechanical transmission relationship between each support component is retrieved;
[0061] An initial scope of concern is determined within a preset radius of influence, and the boundary of the initial scope of concern is dynamically expanded according to the duration and intensity of the high-risk behavioral events to form a high-risk area of concern.
[0062] Specifically, firstly, based on multi-source synchronous monitoring data, target detection algorithms (such as the YOLO algorithm) are used to identify personnel and equipment targets in the video stream and positioning data. Then, tracking algorithms such as Kalman filtering are used to obtain their three-dimensional trajectories and speeds, forming dynamic behavioral characteristics. Subsequently, these dynamic behavioral characteristics are matched in real-time against a pre-defined high-risk behavior rule base. This rule base defines logic in the form of production rules (IF-THEN), such as "IF Target Category = Heavy Excavator AND Distance to Support Pile < 3 meters THEN Event Type = 'Heavy Equipment Close-Range Operation'". Once a match is successful, the high-risk behavioral event and its three-dimensional location are locked. Next, the topological model of the mechanical transmission relationship of the support structure is retrieved. Based on the foundation pit support design drawings, the connection points of key components such as support piles, support beams, and lintels are abstracted into a graphical structure. vertex set The physical connections between components are abstracted into edge sets. This model can be manually constructed or automatically generated from the foundation pit BIM model. It stores the 3D coordinates of vertices in the world coordinate system and the connection attributes of edges. Centered on a point, the radius of interest is dynamically calculated using a range expansion function. The parameters of this function are explicitly defined, and the formula for calculating the expansion radius is:
[0063] ;
[0064] in, The initial radius of influence is determined based on the foundation pit design parameters and the geotechnical investigation report. For example, it can be based on industry experience. Perform initialization. This refers to the depth of the foundation pit excavation. This is an empirical coefficient related to the internal friction angle of the soil layer; for example, it can be taken as 0.7 for soft clay. The duration growth factor, whose calculation corrects for circular dependencies, is defined as follows: . For high-risk behavioral events, the location after triggering The cumulative duration This is a preset baseline time threshold, such as 300 seconds. The growth factor is a preset value, for example, 0.2. This is the disturbance intensity coefficient, whose value is retrieved from a predefined mapping table based on the target category of the triggering event. For example: heavy-duty hydraulic vibratory hammer. Tracked excavators: Dump trucks: Construction workers: After calculating the dynamic radius Then, an algorithm for generating a region based on the force transfer path is executed, rather than simply delineating a spherical region. Starting with the topological model of the support structure In the middle, along the connecting edge Breadth-first search (BFS) is performed, and the physical distance limit for the search is... All in Vertices reachable within range These nodes were marked as affected. Ultimately, high-risk areas were defined as these affected nodes. The convex hull geometry formed in three-dimensional space. This region accurately reflects the spatial range that external disturbances may be transmitted through the mechanical network of the supporting structure. Through the above steps, a closed loop is achieved from identifying global risk behaviors to intelligently delineating local risk regions on the structural mechanical model, ensuring that subsequent vibration spectrum analysis can be targeted and focused, thereby improving the real-time performance and signal-to-noise ratio of the early warning.
[0065] For example, the YOLO algorithm is used to identify the disturbance intensity coefficient of a tracked excavator in a video stream. The table shows a value of 1.5. The excavator continuously operated near the foundation pit support piles, triggering the "heavy equipment operating at close range" rule, thus locking the 3D location of the incident. for Set the initial radius of influence. It is 5.0 meters, with a reference time threshold. For 300 seconds, the growth factor is... The value is 0.2. This refers to the continuous operating time of the excavator. When the duration reaches 600 seconds, calculate the duration growth factor. Substitute into the formula for the radius of extension to calculate: Rice. With Starting with the topological model of the support structure A breadth-first search is performed along the connecting edge to find all reachable component nodes within a 10.5-meter physical distance. The affected nodes found are fitted into convex hull geometry in three-dimensional space to dynamically generate high-risk areas of concern.
[0066] Optionally, the vibration spectrum characteristics within the high-risk area are extracted for focused analysis, and the energy change gradient and dominant frequency shift within the time window are calculated, including:
[0067] Based on the three-dimensional spatial coordinate range of the high-risk area, the corresponding local spectrum data is extracted from the vibration spectrum features;
[0068] The local spectrum data is subjected to bandpass filtering to remove noise frequency components generated by construction background activities, and retain preset frequency components associated with structural response.
[0069] The first derivative of the filtered local spectrum data with time is calculated within a sliding time window to obtain the energy change gradient characterizing the instantaneous growth rate of vibration energy.
[0070] Identify the frequency corresponding to the energy peak within the sliding time window, and calculate the offset value relative to the steady-state reference frequency to obtain the main frequency offset characterizing the change in structural response characteristics;
[0071] By combining the energy change gradient with the dominant frequency offset, regional risk vibration characteristics are generated to represent the quantitative index of regional physical risk.
[0072] Specifically, key quantitative indicators that can directly characterize the precursors of structural instability are extracted from the massive amount of raw vibration signals in the identified local risk areas. This process first involves retrieving and extracting spectral data points whose spatial locations fall within the defined geometric boundaries from a globally distributed vibration spectrum feature dataset, based on the determined three-dimensional spatial coordinate range of the high-risk area, thus forming local spectral data with time-series attributes. Next, a preset digital bandpass filter is applied to this local spectral data, with its passband range set to 5-80 Hz based on the inherent frequency distribution characteristics of the deep foundation pit support structure. After filtering, within a length of... Feature extraction operations are performed within the sliding time window: firstly, the energy change gradient is calculated. It calculates the total energy of all frequency points within the current time window. The first derivative with respect to time is obtained, i.e. ; The first parameter, derived from the energy integral of the filtered local spectrum in the frequency domain, reflects the instantaneous rate of increase in vibration energy. The second parameter is the calculation of the dominant frequency offset. It identifies the actual dominant frequency corresponding to the peak vibration energy within the current time window. Calculate the reference frequency relative to the structure in a steady state. The absolute value of the offset is:
[0073] ;
[0074] in, The reference frequency is derived from the maximum power spectral density point within the current time window. This parameter is derived from the statistical mean of the system under conditions of no construction disturbance during the initial stage of operation; it characterizes the change in structural stiffness or constraint characteristics. Finally, the calculated energy change gradient is... With the main frequency offset By combining vectors, a regional risk vibration feature vector representing the physical risk state of the area is generated. This feature vector serves as direct input evidence for subsequent causal inference graph models, enabling the transformation from visual vibration fields to structural dynamic risk parameters.
[0075] For example, using three-dimensional coordinate points within the designated high-risk area Centered on this, spectral data points within this local range are extracted from the global vibration spectrum characteristics, and then applied... A Butterworth bandpass filter removes construction background noise. (The last part, "in length," appears to be an unrelated fragment and is left untranslated.) Within a 2-second sliding time window, if the total energy of all frequency points in this local region... from The number of times increased from 10J to At 18J, according to the formula Calculate the energy change gradient This reflects the instantaneous surge in vibrational energy. Simultaneously, identifying the maximum point of the power spectral density within this window yields the actual dominant frequency. for The reference frequency of the structural steady state recorded during the initial operation period for Substitute into the formula The main frequency offset was calculated. This indicates a change in structural stiffness. Ultimately, the combined data generates a risk vibration characteristic vector for this region. This provides a quantitative physical risk indicator for subsequent causal reasoning.
[0076] Optionally, the vibration spectrum features are introduced as a physical continuity constraint into the three-dimensional geometric reconstruction algorithm to fill in the data missing regions caused by occlusion and generate three-dimensional structural features. Aligning and correlating the regional risk vibration features, three-dimensional structural features, and sensor data under a unified spatiotemporal reference includes:
[0077] Identify data gap areas in the baseline 3D structural model that cannot be directly observed due to occlusion by dynamic objects;
[0078] The vibration spectrum features are extracted to obtain the spectrum continuity information at the boundary of the data missing region, and the spectrum continuity information is quantified into spatial continuity constraints, which are then added as physical constraints to the optimization objective function of the three-dimensional geometric reconstruction algorithm.
[0079] By iteratively solving the optimization objective function, the continuity of physical vibration information is used to guide the geometric completion of the missing data region, generating the completed three-dimensional structural features.
[0080] Under a unified three-dimensional spatial coordinate system and time axis, taking the structural topology nodes within the high-risk area as units, each node is associated with the regional risk vibration characteristics, the completed three-dimensional structural characteristics, and the sensor data to generate a spatiotemporally aligned fusion feature matrix.
[0081] Specifically, by utilizing the continuity of physical vibration fields propagating on the surface of the support structure, the algorithm is guided to infer and fill in the geometric shape of visual blind spots. This process first uses multi-view stereo vision or neural radiation field algorithms to generate a preliminary baseline 3D structural model, and then identifies in real time the areas to be filled in due to visual feature loss caused by dynamic objects obstructing the view at the construction site. Its edge limit is defined as the boundary. To achieve geometry completion guided by physical information, physical continuity information is quantified into a physical vibration constraint loss term. And incorporate it into the overall optimization objective function of the 3D geometric reconstruction algorithm. middle:
[0082] ;
[0083] in, As the fundamental reconstruction loss term driving 3D geometric reconstruction, it is used to ensure the consistency between the generated 3D model and multi-view visual observation data. In algorithms based on differentiable rendering, it is manifested as photometric loss composed of pixel residuals, and in algorithms based on feature matching, it is manifested as reprojection error of 3D points mapped to the image plane. This is the weighting coefficient, with a value ranging from 0.1 to 0.5. As the core constraint term, it is quantified as the value at the boundary in each reconstruction iteration. Upsample a set of reference points with known vibration spectrum characteristics The observed spectrum vector extracted from the video stream is denoted as Meanwhile, in the missing regions currently predicted by the model... Internal surface sampling corresponding points It also uses a pre-defined lightweight multilayer perceptron (MLP) network to regress and predict its virtual spectral vector. The network is obtained through supervised training using the vibration spectrum characteristics of unobstructed areas in historical monitoring data and their corresponding three-dimensional coordinates. The physical vibration constraint loss term is defined as:
[0084] ;
[0085] in, The squared Euclidean distance between the internal predicted spectrum and the boundary observed spectrum is used to measure the degree of deviation of physical properties. Indicated by reference point Center of the sphere, radius is All internal sampling points within the spatial neighborhood ; For distance-based spatial smoothing weights, The preset spatial attenuation constant is derived from the vibration conduction attenuation coefficient of the support structure material. This loss function penalizes unreasonable vibration faults generated on the reconstructed surface, forcing the geometry of the missing region generated by the algorithm to be physically consistent with the surrounding known structures. After completing the geometric completion, using the structural topology nodes within the high-risk area as basic units, each node is associated with the regional risk vibration characteristics through spatial location mapping. The completed 3D structural features, including changes in surface normals at the node, local curvature or deformation displacement vectors relative to the baseline model, and sensor physical readings, ultimately generate a spatiotemporally aligned fusion feature matrix. This matrix fully encapsulates risk indicators for a local region, from external characterization to internal stress, eliminating early warning failures caused by data breaks. For example... Figure 3 As shown, the original missing state (a) is illustrated: missing region. Vibration data set to 0; Traditional completion result (b): The result using the traditional geometric completion method, showing that the completed data is... A significant numerical step occurred at the regional boundary point compared to the surrounding known reference points, indicating a physical vibration property fault. The completion result (c) of this invention employs the physical continuity constraint-based approach proposed in this invention. The guided completion results show that the completed data is... A smooth transition is achieved within the region, and it is physically consistent with the surrounding known reference points. The clear numerical comparison of each line type on the cross-sectional view directly demonstrates the technical advantage of this invention in maintaining physical continuity.
[0086] For example, during the 3D reconstruction iteration process, areas with missing surface data of the foundation pit support piles caused by obstruction from the excavator boom were identified. Surrounding boundaries There is a reference point above. Known reference point The observed spectrum vector at the location The energy value is 0.81, while the sampling points within the area to be completed... Virtual spectrum vectors predicted by MLP networks The energy value is 0.78, and the square of the Euclidean distance between them is... It is 0.0009. Let the spatial attenuation constant be... The distance between the reference point and the internal points is 5.0. The calculation space smoothing weight is 0.2 meters. Substitute the values into the formula to calculate the physical vibration constraint loss at that point. Set weighting coefficients. The loss due to basic reconstruction is 0.3. The value is 0.02. Substitute this value into the overall optimization objective function to calculate... By solving iteratively... Minimization, leveraging spectral continuity, successfully guided the algorithm to complete the geometric normal vectors and curvature features of the missing region. The completed 3D structural features were then associated with the regional risk vibration features of the node. and measured sensor strain values Perform multi-dimensional vector concatenation to generate the fusion feature matrix of the node. .
[0087] Optionally, the spatiotemporally aligned fusion features are input into a preset causal inference graph model. This model includes behavioral nodes representing external disturbances, structural state nodes representing structural responses, and physical signal nodes representing physical measurements. By solving the causal paths between the nodes, the causal risk chains with the highest confidence are identified, including:
[0088] The state occurrence probability of each node in the causal inference graph model is updated using the spatiotemporally aligned fusion feature matrix;
[0089] Traverse all candidate risk propagation paths that start from the activated behavior node and terminate at the physical signal node;
[0090] Calculate the confidence level of each candidate risk transmission path. The confidence level is obtained by multiplying the state value of the starting behavior node of the path with the preset weight of each causal relationship edge on the path.
[0091] The path with the highest confidence level is selected as the causal risk chain, and the root cause node and transmission path sequence of the risk are extracted from it.
[0092] Specifically, discrete physical monitoring indicators are transformed into a logically rigorous chain of evidence for risk evolution. This process first utilizes the aforementioned generated fusion feature matrix. Update the probability of each node's state in the causal inference graph model. The causal inference graph model is a pre-defined Bayesian network, and its specific graph structure consists of behavioral nodes representing external perturbations. Such as the "Heavy equipment close-range operation" node "Overloaded soil piling around the foundation pit" node Structural state nodes characterizing structural response Such as the "stress concentration in the support pile" node "Instability of supporting components" node and physical signal nodes characterizing physical measurements For example, the node "Frequency Offset Exceeds Standard". "Vibrational energy gradient surge" node The resulting directed acyclic graph is defined as follows: The connections between nodes (edges) represent causal transmission paths, such as edges... This indicates stress concentration caused by heavy equipment operation. When obtaining the real-time status of nodes, the fused feature matrix is used. The evidence in the network is used to update the posterior probability of each node through a connection tree algorithm. Among these, the physical signal nodes... The activation probability is determined by its corresponding feature value, such as the regional risk vibration feature. In or Determined through a preset likelihood function. The state quantization value of the initial behavior node. From the disturbance intensity coefficient Through mapping function Calculated and mapped to The interval represents the activation level of the behavior. Traverse from the already activated behavior nodes... Departure and termination at physical signal node All candidate risk transmission paths For each candidate path, calculate its path confidence. :
[0093] ;
[0094] in, This represents the total number of causal edges contained in the path; For the first on the path The predefined weights of the causal relationship edges represent the conditional probability of the parent node's influence on the child node. The initial values were determined using the Analytic Hierarchy Process (AHP) combined with scores from domain experts on historical accident cases, such as edge... weight The confidence level was set to 0.85. By sorting all candidate paths in descending order of confidence, the path with the highest confidence was selected as the causal risk chain, and the root cause node and complete transmission path sequence were extracted from it. This process achieves complete causal attribution from the risk source behavior to the physical signal manifestation, providing managers with scientific and interpretable decision support.
[0095] For example, suppose high-risk behavior nodes are identified. For "close-range operation of heavy equipment", its disturbance intensity coefficient The value is 1.5, and its state quantization value is calculated through a mapping function. This behavior node One path in the causal reasoning graph model was activated. The path passes through the structural state node. "Stress concentration in the support pile body" terminates at the physical signal node. "Frequency offset exceeds limit". The path is known to contain two causal edges with preset weights of [weights to be filled in]. ,correspond and ,correspond Substitute into the path confidence calculation formula. The confidence level of the path is obtained. If, after traversal, the path is found to have the highest confidence among all candidate risk transmission paths, it is identified as a causal risk chain. Analysis reveals the root cause to be improper operation of heavy equipment, with the transmission process manifesting as a shift in the dominant frequency caused by stress anomalies in the support piles, thus providing clear evidence for on-site management.
[0096] Optionally, the method further includes:
[0097] Collect the historical output of the graded early warning information and the corresponding on-site handling feedback, and construct a case library containing positive and negative sample cases;
[0098] Extract and analyze the causal risk chains corresponding to false positives and false negatives in the case library;
[0099] Based on the analysis results, the early warning model is iteratively optimized by adjusting the weights of each causal relationship edge in the causal reasoning graph model, or by adjusting the guiding weights of the vibration spectrum features in the three-dimensional geometric reconstruction algorithm for filling in missing data regions.
[0100] Specifically, through a closed-loop learning mechanism, the conditional probability parameters in the causal inference graph model and the physical constraint strength during the 3D reconstruction process are continuously corrected, enabling the system to adapt to complex background noise changes at different construction stages. This process first collects all risk levels, risk sources, and transmission paths output by the system throughout its historical operating cycles, and combines this with actual handling feedback provided by on-site safety management personnel to construct a comprehensive case library containing positive samples (real risks) and negative samples (false alarms). Subsequently, an automated feature analysis tool is used to deeply mine the case library. This tool works by comparing the causal risk chains output by the model in false alarm / missed alarm cases with the actual evolution paths confirmed by on-site feedback, statistically identifying specific causal relationships in the model where logical judgment errors frequently occur. This was then used as the focus of subsequent parameter optimization. Based on the analysis results, an iterative parameter optimization program was initiated: On the one hand, for the causal inference graph model, when the accumulated labeled cases reached a preset threshold, the maximum likelihood estimation (MLE) algorithm was used to re-estimate the conditional probability table (CPT) of each node in the network, thereby realizing the edge weights. The update formula is:
[0101] ;
[0102] in, The updated causal edge weights are mathematically equivalent to the conditional probability of a child node appearing in the state of the parent node. Parent node in the case library The total number of samples that are active at the same time as their child nodes; This represents the total number of samples where the parent node is individually active. This probabilistic statistical law can correct for biases in expert experience weights. On the other hand, it relates to the guiding weights in 3D geometric reconstruction algorithms. Dynamic calibration is performed. To avoid non-convergence caused by unidirectional adjustment, a performance search strategy is adopted: on the validation set, the weighted sum of reconstruction accuracy and warning accuracy is used as the objective function. The optimal weight is searched within a fixed step size, expressed by the formula:
[0103] ;
[0104] in, The adjusted guiding weight; The error rate for early warning identification under this weight; The average residual of the geometric reconstruction; The preset balance coefficients are used. Through the coordinated adjustment of the above parameters, false alarms caused by background environmental interference can be reduced, and the sensitivity to capturing real structural response signals can be enhanced. This achieves a dual evolution of the early warning model at the levels of physical modeling and logical reasoning, ensuring the accuracy and reliability of the early warning system during long-term operation.
[0105] For example, when the number of cases accumulated in the case library reaches a preset threshold, the parent node of "heavy equipment close-range operation" in the causal reasoning graph model is targeted. With the sub-node of "stress concentration in the support pile body" Causal edges between Optimization was performed. Statistics showed that the total number of samples in the case library where the parent node was active was [number missing]. There are 100 groups, and the total number of samples in which both parent and child nodes are simultaneously active. There are 88 groups. Substitute them into the maximum likelihood estimation update formula. This corrects the probability weight of that side from the initial expert-empirical value of 0.85 to 0.88. Simultaneously, regarding the guiding weight... Initiate performance search calibration and set the balance factor. It is 0.5. On the validation set, when When the value is 0.3, the measured early warning recognition error rate is... The mean residual of geometric reconstruction is 0.05. The value is 0.08. Substitute this value into the objective function to calculate the score: Already in Step-size search revealed that this score was the highest among all values, therefore... The value was set to 0.3. Through the coordinated adjustment of the above parameters, the early warning model achieved a dual evolution at both the probabilistic logic and physical modeling levels.
[0106] Optionally, the generation and output of tiered early warning information includes:
[0107] A set of confidence thresholds associated with risk levels is pre-defined;
[0108] The confidence level of the identified causal risk chain is compared with the set of confidence thresholds, and the final risk level is determined by ladder mapping.
[0109] Extract the corresponding risk root cause information and transmission path information from the causal risk chain, and match the corresponding handling suggestions from the preset emergency plan database;
[0110] The system integrates the final risk level, risk source information, transmission path information, and handling suggestions to generate structured, tiered early warning information and pushes it to preset terminals.
[0111] Specifically, generating and outputting tiered early warning information involves pre-setting a set of confidence thresholds associated with risk levels, comparing the confidence levels of identified causal risk chains with these thresholds, and determining the final risk level through a step-by-step mapping. In this process, the set of confidence thresholds... This data is derived from a statistical analysis of the confidence distribution of causal chains in a historical accident case database. This indicates the low confidence threshold. Indicates the medium confidence threshold. This indicates the high confidence threshold. , and These constitute the confidence threshold boundaries used to determine different risk levels, for example, setting... , and Final risk level The ladder mapping rule is defined as: when the confidence level of the causal risk chain... satisfy hour, It is a blue alert (attention level); when hour, It is a yellow alert (alert level); when hour, An orange or red alert (danger level) is issued. The value is derived from the highest confidence result obtained through causal path calculation. After determining the risk level, the corresponding risk root cause information, such as the initial behavioral node, is extracted from the causal risk chain. With transmission path information, such as intermediate structural state nodes The system then matches corresponding response suggestions from a pre-set emergency response plan database. The matching logic for these suggestions is based on a composite search of the risk root cause node type and risk level. For example, when the risk root cause is "close-range operation of heavy equipment" and the level is "yellow alert," the system automatically matches a suggestion to "limit equipment speed and increase the frequency of support shaft force monitoring." Finally, the system integrates the final risk level, risk root cause information, transmission path information, and response suggestions to generate structured, tiered early warning information, which is then sent to pre-set management terminals using a message queue or push protocol. This combination of hierarchical mapping and causal tracing transforms early warning information from a single numerical value into logically supported and actionable guidelines, effectively improving the scientific rigor of on-site safety decision-making.
[0112] For example, a set of confidence thresholds Preset At a certain monitoring point, the calculated highest confidence causal risk chain path is "heavy equipment close-range operation". Stress Concentration in Support Pile Main frequency offset exceeds the standard "The confidence level of this path" It is 0.476. Compared with the threshold set, because it satisfies The ladder mapping rule determines the final risk level. A blue alert was issued. Subsequently, the root cause information was extracted from the risk chain as "close-range operation of heavy equipment," and the transmission path information was "stress concentration in the support piles leading to frequency shift." The corresponding response recommendation was retrieved from the emergency response database: "Strengthen patrols in the area and check the stability of the support piles." Finally, the risk level, root cause, path, and recommendations were integrated to generate a structured early warning message, which was pushed to management personnel terminals in real time.
[0113] Based on the same inventive concept, such as Figure 4 As shown, the present invention also provides a machine vision-based deep foundation pit construction safety early warning system, the system comprising:
[0114] The data acquisition and alignment module is used to acquire heterogeneous data, including multi-view video streams, sensor data, and positioning data from the construction area, and perform spatiotemporal alignment based on a unified time axis and spatial coordinate system to generate multi-source synchronous monitoring data.
[0115] The spectrum feature extraction module is used to perform time-domain amplification processing on the multi-view video stream to make pixel-level micro-vibration signals explicit, and generate an initial spectrum map characterizing the frequency and energy distribution of each pixel point through time-frequency transformation, which is then mapped to a reference three-dimensional structural model to give it spatial attributes, and the vibration spectrum features of the support structure surface are extracted.
[0116] The risk zone delineation module is used to identify personnel and equipment targets in multi-view video streams and positioning data and track dynamic behavior characteristics. It matches the dynamic behavior characteristics with a preset high-risk behavior rule library to identify high-risk behavior events and their three-dimensional locations. Based on the topological model of the mechanical transmission relationship of the support structure, it determines the local area affected by the event as a high-risk area of concern.
[0117] The spectrum focusing analysis module is used to extract the vibration spectrum characteristics of the high-risk area for focusing analysis, and calculate the energy change gradient and dominant frequency offset within the time window to generate regional risk vibration characteristics.
[0118] The fusion and geometric reconstruction module is used to introduce the vibration spectrum features as physical continuity constraints into the three-dimensional geometric reconstruction algorithm, fill in the data missing areas caused by occlusion and generate three-dimensional structural features, and align and correlate the regional risk vibration features, three-dimensional structural features and sensor data under a unified spatiotemporal reference to generate spatiotemporally aligned fusion features.
[0119] The causal path calculation and early warning generation module is used to input spatiotemporally aligned fusion features into a preset causal reasoning graph model. This model includes behavioral nodes representing external disturbances, structural state nodes representing structural responses, and physical signal nodes representing physical measurements. By calculating the causal paths between each node, the module identifies the causal risk chain with the highest confidence level and generates and outputs graded early warning information accordingly.
[0120] It should be noted that all equivalent changes and modifications made in accordance with the teachings of this invention are still within the scope of this invention. Those skilled in the art will readily conceive of other embodiments of this invention upon considering the specification and the disclosure of practical truth. This application is intended to cover any variations, uses, or adaptations of this invention that follow the general principles of this invention and include common knowledge or conventional techniques in the art not described herein.
Claims
1. A machine vision-based deep foundation pit construction safety warning method, characterized in that, The method includes: It acquires heterogeneous data, including multi-view video streams, sensor data, and positioning data from the construction area, and performs spatiotemporal alignment based on a unified time axis and spatial coordinate system to generate multi-source synchronous monitoring data. The multi-view video stream is subjected to temporal domain amplification to make the pixel-level micro-vibration signal explicit, and an initial spectrum diagram characterizing the frequency and energy distribution of each pixel is generated by time-frequency transformation. This spectrum diagram is then mapped onto a reference three-dimensional structural model to give it spatial attributes, and the vibration spectrum characteristics of the support structure surface are extracted. By identifying personnel and equipment targets in multi-view video streams and positioning data and tracking dynamic behavior characteristics, the dynamic behavior characteristics are matched with a preset high-risk behavior rule base to identify high-risk behavior events and their three-dimensional locations. Based on the topological model of the mechanical transmission relationship of the support structure, the local area affected by the event is determined as a high-risk area of concern. Vibration spectrum characteristics within the high-risk area are extracted for focused analysis, and the energy change gradient and dominant frequency shift within the time window are calculated to generate regional risk vibration characteristics. The vibration spectrum features are introduced into the three-dimensional geometric reconstruction algorithm as physical continuity constraints to fill in the data missing areas caused by occlusion and generate three-dimensional structural features. The regional risk vibration features, three-dimensional structural features and sensor data are aligned and correlated under a unified spatiotemporal reference to generate spatiotemporally aligned fusion features. The spatiotemporally aligned fusion features are input into a preset causal reasoning graph model. This model includes behavioral nodes representing external disturbances, structural state nodes representing structural responses, and physical signal nodes representing physical measurements. By solving the causal paths between each node, the causal risk chain with the highest confidence is identified, and hierarchical early warning information is generated and output accordingly.
2. The method for early warning of deep foundation pit construction safety based on machine vision according to claim 1, characterized in that, The acquisition of heterogeneous data, including multi-view video streams, sensor data, and positioning data of the construction area, and spatiotemporal alignment based on a unified time axis and spatial coordinate system, includes: High-definition video streams of the construction area are simultaneously acquired by several non-co-linearly deployed industrial cameras to form a multi-view video stream; The vibration sensors and displacement monitoring devices deployed at key stress nodes of the support structure are periodically polled to obtain sensor data containing dynamic changes in the physical quantities of the structure. It receives wireless signals transmitted by positioning tags installed on construction workers and construction machinery, calculates three-dimensional coordinate information, and forms positioning data; Using the central server as the time reference source, a unified timestamp is applied to each frame of the multi-view video stream, each set of readings of the sensor data, and each coordinate point of the positioning data. Based on the calibrated camera parameters, the heterogeneous data is converted to a unified three-dimensional spatial coordinate system of the foundation pit to generate multi-source synchronous monitoring data.
3. The method for early warning of deep foundation pit construction safety based on machine vision according to claim 2, characterized in that, The multi-view video stream is subjected to temporal domain magnification to make pixel-level micro-vibration signals explicit, and an initial spectrum representing the frequency and energy distribution of each pixel is generated through time-frequency transformation. This spectrum is then mapped onto a reference three-dimensional structural model to assign spatial properties, including: A temporal filter is used to enhance the amplitude of pixel brightness fluctuations within a specific frequency range in a multi-view video stream, thereby making the pixel-level micro-vibration signals on the surface of the support structure explicit. A short-time Fourier transform is performed on the pixel-level micro-vibration signal of each pixel to generate an initial spectrum that characterizes the frequency components and energy distribution of the pixel in different time windows. Based on the calibrated camera parameters, the two-dimensional pixels in the initial spectrogram are mapped to the three-dimensional spatial coordinates on the reference three-dimensional structural model by ray casting. For the same point on the reference three-dimensional structural model that is simultaneously covered by multiple perspectives, the corresponding multiple spectral information is weighted and fused to generate vibration spectral features with three-dimensional spatial attributes.
4. The method for early warning of deep foundation pit construction safety based on machine vision according to claim 1, characterized in that, The dynamic behavioral characteristics are matched with a preset high-risk behavior rule base to identify high-risk behavioral events and their three-dimensional locations. Based on the topological model of the mechanical transmission relationship of the support structure, the local areas affected by the events are identified as high-risk areas of concern, including: The target trajectory and speed information in the dynamic behavior features are compared with the behavior types in the rule base to identify high-risk behavior events and their three-dimensional locations. Centered on the three-dimensional location of occurrence, a topological model of the support structure representing the physical connection and mechanical transmission relationship between each support component is retrieved; An initial scope of concern is determined within a preset radius of influence, and the boundary of the initial scope of concern is dynamically expanded according to the duration and intensity of the high-risk behavioral events to form a high-risk area of concern.
5. The method for early warning of deep foundation pit construction safety based on machine vision according to claim 1, characterized in that, The vibration spectrum characteristics within the high-risk area of concern are extracted for focused analysis, and the energy change gradient and dominant frequency shift within the time window are calculated, including: Based on the three-dimensional spatial coordinate range of the high-risk area, the corresponding local spectrum data is extracted from the vibration spectrum features; The local spectrum data is subjected to bandpass filtering to remove noise frequency components generated by construction background activities, and retain preset frequency components associated with structural response. The first derivative of the filtered local spectrum data with time is calculated within a sliding time window to obtain the energy change gradient characterizing the instantaneous growth rate of vibration energy. Identify the frequency corresponding to the energy peak within the sliding time window, and calculate the offset value relative to the steady-state reference frequency to obtain the main frequency offset characterizing the change in structural response characteristics; By combining the energy change gradient with the dominant frequency offset, regional risk vibration characteristics are generated to represent the quantitative index of regional physical risk.
6. The method for early warning of deep foundation pit construction safety based on machine vision according to claim 3, characterized in that, The vibration spectrum features are introduced as physical continuity constraints into the 3D geometric reconstruction algorithm to fill in data gaps caused by occlusion and generate 3D structural features. The regional risk vibration features, 3D structural features, and sensor data are then aligned and correlated under a unified spatiotemporal reference, including: Identify data gap areas in the baseline 3D structural model that cannot be directly observed due to occlusion by dynamic objects; The vibration spectrum features are extracted to obtain the spectrum continuity information at the boundary of the data missing region, and the spectrum continuity information is quantified into spatial continuity constraints, which are then added as physical constraints to the optimization objective function of the three-dimensional geometric reconstruction algorithm. By iteratively solving the optimization objective function, the continuity of physical vibration information is used to guide the geometric completion of the missing data region, generating the completed three-dimensional structural features. Under a unified three-dimensional spatial coordinate system and time axis, taking the structural topology nodes within the high-risk area as units, each node is associated with the regional risk vibration characteristics, the completed three-dimensional structural characteristics, and the sensor data to generate a spatiotemporally aligned fusion feature matrix.
7. A method for safety early warning of deep foundation pit construction based on machine vision according to claim 6, characterized in that, The spatiotemporally aligned fusion features are input into a preset causal inference graph model. This model includes behavioral nodes representing external disturbances, structural state nodes representing structural responses, and physical signal nodes representing physical measurements. By solving the causal paths between the nodes, the causal risk chains with the highest confidence are identified, including: The state occurrence probability of each node in the causal inference graph model is updated using the spatiotemporally aligned fusion feature matrix; Traverse all candidate risk propagation paths that start from the activated behavior node and terminate at the physical signal node; Calculate the confidence level of each candidate risk transmission path. The confidence level is obtained by multiplying the state value of the starting behavior node of the path with the preset weight of each causal relationship edge on the path. The path with the highest confidence level is selected as the causal risk chain, and the root cause node and transmission path sequence of the risk are extracted from it.
8. A method for safety early warning of deep foundation pit construction based on machine vision according to claim 7, characterized in that, The method further includes: Collect the historical output of the graded early warning information and the corresponding on-site handling feedback, and construct a case library containing positive and negative sample cases; Extract and analyze the causal risk chains corresponding to false positives and false negatives in the case library; Based on the analysis results, the early warning model is iteratively optimized by adjusting the weights of each causal relationship edge in the causal reasoning graph model, or by adjusting the guiding weights of the vibration spectrum features in the three-dimensional geometric reconstruction algorithm for filling in missing data regions.
9. The method for early warning of deep foundation pit construction safety based on machine vision according to claim 1, characterized in that, The generation and output of tiered early warning information includes: A set of confidence thresholds associated with risk levels is pre-defined; The confidence level of the identified causal risk chain is compared with the set of confidence thresholds, and the final risk level is determined by ladder mapping. Extract the corresponding risk root cause information and transmission path information from the causal risk chain, and match the corresponding handling suggestions from the preset emergency plan database; The system integrates the final risk level, risk source information, transmission path information, and handling suggestions to generate structured, tiered early warning information and pushes it to preset terminals.
10. A machine vision-based safety early warning system for deep foundation pit construction, characterized in that, The system is used for a machine vision-based safety early warning method for deep foundation pit construction as described in any one of claims 1-9, the system comprising: The data acquisition and alignment module is used to acquire heterogeneous data, including multi-view video streams, sensor data, and positioning data from the construction area, and perform spatiotemporal alignment based on a unified time axis and spatial coordinate system to generate multi-source synchronous monitoring data. The spectrum feature extraction module is used to perform time-domain amplification processing on the multi-view video stream to make pixel-level micro-vibration signals explicit, and generate an initial spectrum map characterizing the frequency and energy distribution of each pixel point through time-frequency transformation, which is then mapped to a reference three-dimensional structural model to give it spatial attributes, and the vibration spectrum features of the support structure surface are extracted. The risk area delineation module is used to identify personnel and equipment targets in multi-view video streams and positioning data and track dynamic behavior characteristics. It matches the dynamic behavior characteristics with a preset high-risk behavior rule library to identify high-risk behavior events and their three-dimensional locations. Based on the topological model of the mechanical transmission relationship of the support structure, it determines the local area affected by the event as a high-risk area of concern. The spectrum focusing analysis module is used to extract the vibration spectrum characteristics of the high-risk area for focusing analysis, and calculate the energy change gradient and dominant frequency offset within the time window to generate regional risk vibration characteristics. The fusion and geometric reconstruction module is used to introduce the vibration spectrum features as physical continuity constraints into the three-dimensional geometric reconstruction algorithm, fill in the data missing areas caused by occlusion and generate three-dimensional structural features, and align and correlate the regional risk vibration features, three-dimensional structural features and sensor data under a unified spatiotemporal reference to generate spatiotemporally aligned fusion features. The causal path calculation and early warning generation module is used to input the spatiotemporally aligned fusion features into a preset causal reasoning graph model. This model includes behavioral nodes representing external disturbances, structural state nodes representing structural responses, and physical signal nodes representing physical measurements. By calculating the causal paths between each node, the module identifies the causal risk chain with the highest confidence level and generates and outputs graded early warning information accordingly.