AI vision-based water transportation construction project safety hazard intelligent monitoring method and system
By collecting video streams of underwater pile foundation construction and safety specification texts, and combining feature enhancement and cross-modal fusion technologies, the problems of easy loss of small target features, static background interference, and insufficient recognition of continuous behavior in underwater pile foundation construction have been solved, achieving accurate identification of safety hazards and compliance judgment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHANGSHA JIAOTONG LOGISTICS CO LTD
- Filing Date
- 2026-02-09
- Publication Date
- 2026-05-29
AI Technical Summary
Traditional safety hazard monitoring technologies for waterway construction projects suffer from poor identification accuracy and inability to link with safety regulations in waterborne pile foundation construction scenarios. They also cannot adapt to the needs of small target features, static background interference, and continuous behavior, resulting in inaccurate hazard identification and a lack of compliance judgment.
High-definition industrial cameras or drones are used to collect video stream data. Combined with water transport safety regulations, feature enhancement, cross-modal fusion and temporal fusion technologies are used to extract small target features, filter static background interference, integrate continuous frame features, and make behavioral compliance judgments in conjunction with water transport safety regulations.
It achieves precise capture of small target features during the construction of underwater pile foundations, filters static background interference, fully represents continuous behavior, accurately identifies safety hazards and provides compliance judgment, and meets the monitoring needs of underwater pile foundation construction.
Smart Images

Figure CN122116132A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of safety hazard monitoring in waterway construction projects, and in particular to an intelligent monitoring method and system for safety hazards in waterway construction projects based on AI vision. Background Technology
[0002] The AI-based vision-based intelligent monitoring technology for safety hazards in waterway construction projects uses computer vision as its core to extract and analyze features from on-site video data, ultimately generating hazard identification results. The general process involves real-time frame parsing of the on-site construction video stream, extracting visual features such as personnel operations, equipment status, and working environment, and then identifying and judging safety hazard behaviors.
[0003] Among them, the construction of waterborne pile foundations is a core and critical link in waterway construction projects. Its safety monitoring is more difficult than that of ordinary waterway operations and is also different from land construction projects such as building construction and municipal engineering. The specific characteristics of the scenario are as follows: the monitoring targets are mainly small targets such as workers. These small targets are small in size and have a low feature ratio. They are easily obscured by the pile foundation, working platform, water surface, etc., and their identification is weak. After multi-step feature processing, their features are even more easily weakened; the working background includes a large area of static water surface, fixed pile foundation and working platform. The static background has a high proportion and can easily interfere with the dynamic movement information of construction workers and small operating equipment. At the same time, the water environment is easily affected by natural factors such as wind, waves, tides, fog and rain, which further increases the difficulty of feature extraction; the construction process is continuous and coherent. The behavioral features of workers' operation, equipment movement, pile foundation erection, etc. are distributed in multiple frames of video stream, rather than isolated single frame static actions. Moreover, the boundaries of the working area are blurred and highly mobile.
[0004] Traditional approaches mainly rely on a combination of manual inspections and fixed-point video surveillance. Manual inspections are limited by factors such as high labor costs, significant differences in subjective judgment, and the hazardous environment of water operations, which can easily lead to the omission of minor hidden dangers and misjudgment of violations. Fixed-point surveillance can only record and store the operation footage and cannot actively identify violations. Even if a suspected hidden danger is found, it is still necessary to manually review a large number of paper or offline stored fragmented standard texts for comparison and verification, which is not only time-consuming and labor-intensive, but also prone to judgment errors, seriously affecting the timeliness and accuracy of hidden danger rectification.
[0005] While existing deep learning-based AI visual monitoring technologies can automatically detect some safety hazards, they still have technical limitations and cannot adapt to the specific needs of underwater pile foundation construction. These technologies often employ single-frame image feature extraction, failing to specifically enhance the features of small targets such as workers, and lacking effective static background filtering measures. They cannot accurately distinguish between dynamic construction activities and static backgrounds, and are susceptible to interference from non-construction dynamics such as floating objects and distant, unrelated vessels, leading to the loss of small target features and distorted motion trajectory capture. Furthermore, they fail to consider the continuity of underwater pile foundation construction, severing the temporal correlation between video frames and only extracting static features from a single frame, thus failing to fully represent the continuous construction process. During construction, most existing technologies can only identify simple safety hazards such as not wearing life jackets or helmets as required, but cannot identify hazards such as personnel crossing boundaries or unauthorized removal of protective equipment. Furthermore, these technologies generally employ a single visual feature recognition mode, failing to incorporate waterway safety regulations and textual data, and do not attempt to establish a correlation between visual behavioral features and regulatory textual features. Lacking guidance from waterway safety regulations, they cannot assess the compliance of construction activities, only performing basic behavioral identification. The detection results lack sufficient interpretability, making it difficult for regulators to quickly obtain evidence of violations corresponding to hazards. Hazard rectification guidance lacks precise support, failing to meet the application requirements of accurate and compliant hazard monitoring in waterborne pile foundation construction. Summary of the Invention
[0006] In view of this, the present invention aims to provide an intelligent monitoring method and system for safety hazards in waterway construction projects based on AI vision, so as to solve the problems of poor accuracy in identifying safety hazards in waterway construction projects by traditional methods and the difficulty in linking with safety specification texts.
[0007] A method for intelligent monitoring of safety hazards in waterway construction projects based on AI vision includes: A1: Collect video stream data of underwater pile foundation construction and text data of waterway safety specifications, and preprocess them respectively to obtain preprocessed video stream data of underwater pile foundation construction and preprocessed text data of waterway safety specifications; A2: Based on the preprocessed video stream data of underwater pile foundation construction, extract the foundation features and small target feature gain weights of underwater pile foundation construction, and then calculate the motion trajectory mask matrix and temporal fusion features in sequence; A3: Based on the basic characteristics of underwater pile foundation construction, extract the structural mask matrix of the water transport scene and the construction characteristics of underwater pile foundation after physical constraints in sequence, and generate safety hazard behavior characteristics; A4: Guided by the behavioral characteristics of safety hazards, based on the preprocessed waterway safety specification text data, initial text features and weighted text features are extracted, and then cross-modal fusion is performed to obtain cross-modal fused features; A5: Enhance the characteristics of safety hazard behaviors and then combine them with cross-modal fusion features to calculate the compliance matching degree of the behavior; A6: Calculate the behavioral compliance result based on the behavioral compliance matching degree, calculate the text matching score based on the behavioral compliance result, and extract the content from the corresponding pre-processed waterway safety specification text data as the basis for determining non-compliance of waterway construction project safety hazards.
[0008] Furthermore, step A1 also includes: A11: Acquire video stream data of underwater pile foundation construction using high-definition industrial cameras or drone equipment; the type of the underwater pile foundation construction video stream data is continuous frame image sequence data; A12: The video stream data of the underwater pile foundation construction is preprocessed by Gaussian filtering for noise reduction, frame rate linear interpolation normalization and histogram equalization to obtain the preprocessed video stream data of the underwater pile foundation construction. A13: Collect waterway safety regulations text data and waterway construction terminology data; after deduplicating the waterway construction terminology data, construct a custom word segmentation dictionary using terminology hierarchical annotation and word frequency weighting; preprocess the waterway safety regulations text data using forced matching word segmentation based on the custom word segmentation dictionary, invalid character removal using regular expressions, and stop word filtering to obtain preprocessed waterway safety regulations text data.
[0009] Furthermore, step A2 also includes: A21: Based on the preprocessed video stream data of underwater pile foundation construction, extract the foundation features and small target feature gain weights of underwater pile foundation construction. A22: Based on the characteristics of underwater pile foundation construction, inter-frame motion information is captured by optical flow method, and then a self-attention module is introduced to filter static background interference, and the motion trajectory mask matrix is calculated. A23: Based on the motion trajectory mask matrix and the characteristics of the underwater pile foundation construction, the characteristics of the underwater pile foundation construction in consecutive frames are fused in a temporal sequence. The temporal fusion characteristics are calculated through a temporal feature gating fusion mechanism.
[0010] Furthermore, step A2 also includes: First, the single-frame image in the preprocessed video stream data of underwater pile foundation construction is processed by a convolutional layer with residual connection to obtain the underwater pile foundation construction features of the current frame image. Then, the underwater pile foundation construction features are subjected to max pooling processing, and then processed by a multilayer perceptron and a sigmoid function to obtain the small target feature gain weight of the s-th frame image. Based on the water-based pile foundation construction features of the current frame image and the water-based pile foundation construction features of the previous frame image, the inter-frame motion information is captured by the optical flow calculation function, and then the inter-frame motion information is processed by the self-attention mechanism. After processing by the multilayer perceptron and the Sigmoid function, the motion trajectory mask matrix of the current frame image is obtained. The inter-frame transposed fusion features of the current frame image and the previous frame image are transposed and then multiplied by a Hadamard product to obtain the inter-frame transposed fusion features of the current frame image. The inter-frame transposed fusion features are processed by a graph convolutional network and then multiplied by a Hadamard product with the motion trajectory mask matrix of the current frame image. After processing by a gated multilayer perceptron and a Sigmoid function, the temporal gating weights of the current frame are obtained. The inter-frame transposed fusion features of the current frame image are multiplied by a Hadamard product with the temporal gating weights. At the same time, the inter-frame transposed fusion features of the previous frame image are multiplied by a Hadamard product with the result of 1 minus the temporal gating weights. The two results are then added element-wise to obtain the temporal fusion features of the current frame image.
[0011] It should be further explained that the monitoring targets in the underwater pile foundation construction scenario are mainly the workers. These small targets are easily obscured by the pile foundation, work platform, water surface, etc. Their characteristics account for a low proportion and have weak recognition in the construction video stream. At the same time, the scene background includes a large area of static water surface, fixed pile foundation and work platform, etc. The static background accounts for a large proportion and is easy to interfere with the dynamic movement information of construction personnel and equipment. In addition, underwater pile foundation construction is a continuous operation, and its behavioral characteristics are distributed in multiple frames of the video stream, making it difficult for traditional monitoring methods to accurately capture the dynamic behavior of construction and fully represent the continuous operation process. In step A21, this invention first uses a convolutional layer with residual connections to process the preprocessed single-frame image of the video stream of underwater pile foundation construction, effectively extracting the basic features of the underwater pile foundation construction in the current frame image. Simultaneously, through a combination of max pooling, multilayer perceptron, and the sigmoid function, the gain weights for small target features are calculated. Max pooling preserves key features of small targets, multilayer perceptron enhances feature representation, and the sigmoid function normalizes the weights, thereby strengthening the features of small targets in the scene. This ensures that even if small targets are partially occluded, their features can still be effectively captured, avoiding subsequent behavior recognition failures due to the loss of small target features; thus solving the problems of low feature ratio and easy loss of small targets. Building upon this, step A22 utilizes optical flow to capture inter-frame motion information between the current frame and the previous frame's features related to the construction of the underwater pile foundation, capturing the dynamic movement trajectories of construction personnel. A self-attention module is then introduced to process the motion information, filtering out interference from static backgrounds such as the water surface and fixed pile foundations. Further optimization using the Sigmoid function and a multilayer perceptron yields a motion trajectory mask matrix. The self-attention mechanism focuses on dynamic motion areas and suppresses static background areas, ensuring the accuracy of motion trajectory capture. The motion trajectory mask matrix precisely filters out non-construction dynamic interference such as floating objects on the water surface and distant unrelated vessels, while also eliminating interference from static water surfaces and fixed work platforms. This ensures that only the dynamic movement trajectories of personnel and equipment within the construction area are captured, aligning with the actual characteristics of underwater pile foundation construction: dynamic operations and a large proportion of static background. Finally, step A23 introduces a temporal feature gating fusion mechanism. First, the features of the current frame's underwater pile foundation construction are transposed with those of the previous frame's underwater pile foundation construction and then subjected to Hadamard product operation to obtain the inter-frame transposed fused features. This fused feature is then processed by a graph convolutional network. Combined with the motion trajectory mask matrix, the temporal gating weights are calculated using a gated multilayer perceptron and the Sigmoid function. Then, weighted fusion is performed to obtain the temporal fused features. By strengthening the inter-frame feature association through the graph convolutional network and dynamically adjusting the proportion of features in adjacent frames through the temporal gating weights, the feature information of multiple frames is integrated, solving the problem of separating the temporal association between frames and fully representing the continuous construction behavior process. For example, it can completely capture the entire continuous process of workers walking from the work platform to the pile foundation and returning after completing the operation, realizing the analysis of continuous construction behavior, rather than the traditional method that can only identify static actions in a single frame. In the current technology for safety monitoring of underwater pile foundation construction, the method of extracting features from single-frame images is mostly used for video stream data. It does not specifically enhance the features of small targets, nor does it take effective static background filtering measures. Furthermore, it does not consider the continuity of construction behavior. It can only extract static features from single-frame images and cannot capture continuous construction behavior trajectories. This results in the loss of small target features, severe background interference, and the inability to achieve behavior analysis. It is difficult to adapt to the actual needs of underwater pile foundation construction. Most of them can only identify simple safety hazards such as not wearing life jackets or safety helmets according to regulations. Compared with existing technologies, the advantages of this invention are as follows: First, by designing small target feature gain weights, it specifically enhances the features of small targets in the construction scenario of underwater pile foundations, effectively solving the problems of easy loss of small target features and low recognition in existing technologies, and ensuring that the construction behavior of small targets can be accurately captured; Second, through the combined design of optical flow method and self-attention module, it distinguishes between dynamic construction behavior and static background of the scene, and filters out various static backgrounds and non-construction dynamic interference; Third, through the temporal feature gating fusion mechanism, it achieves effective fusion of continuous frame features, integrates construction behavior features in multiple frames, breaks through the limitation of single-frame extraction in existing technologies, and can completely represent continuous underwater pile foundation construction behavior, laying the foundation for the identification of complex safety hazards, such as unauthorized removal of life jackets and illegal crossing, and better meeting the actual monitoring needs of underwater pile foundation construction.
[0012] Furthermore, step A3 also includes: A31: Based on the characteristics of underwater pile foundation construction, key structures in the water transport scenario are identified through an attention-based semantic segmentation network, and the structural mask matrix of the water transport scenario is calculated. A32: Based on the temporal fusion features and the water transport scenario structure mask matrix, feature optimization is performed to obtain the physically constrained construction features of the water pile foundation; A33: Based on the feature gain weight of small targets and the construction characteristics of water pile foundations after physical constraints, the dimensions are adaptively adjusted, attention pooling operation is introduced to highlight key features, and safety hazard behavior features are output.
[0013] Furthermore, step A3 also includes: Based on the characteristics of underwater pile foundation construction, key structures in the water transport scene are identified through a Unet network with self-attention. The identification results are then processed by a binarization function to obtain the water transport scene structure mask matrix of the current frame image. Based on the temporal fusion features and the water transport scene structure mask matrix of the current frame image, the temporal fusion features are first processed by the Laplacian operator, and then the processing result is processed by a graph convolutional network. After that, the Hadamard product is performed with the water transport scene structure mask matrix of the current frame image and the result is negative to obtain the constrained optimization features of the current frame image. The temporal fusion features and the constrained optimization features are then added element-wise to obtain the physically constrained water pile foundation construction features of the current frame image. The construction features of the water-based pile foundation under physical constraints are subjected to attention pooling, and then Hadamard product is performed with the small target feature gain weights to output the safety hazard behavior features of the current frame image.
[0014] It needs further explanation that the construction scenario of underwater pile foundations has a clear division between critical operational structures and non-operational areas. Critical operational structures include the pile foundation itself, the work platform, and the construction channel, while non-operational areas include non-construction areas on the water surface and restricted areas outside the pile foundation. This scenario is also constrained by specific physical rules, meaning that construction personnel and operating equipment can only work within pre-defined areas such as the work platform and the pile foundation itself, and cannot penetrate the pile foundation or be suspended in the water without support. Traditional monitoring methods do not clearly define the critical structures in this scenario, nor do they incorporate the aforementioned physical rules into the feature extraction process. This leads to the misjudgment of irrelevant movements in non-operational areas, such as floating objects on the water surface or distant passing vessels, as operational activities in the construction area, resulting in distorted monitoring range. Existing technologies for feature optimization in the safety monitoring of underwater pile foundation construction generally adopt common feature processing methods. The core approach is to perform general optimization of the extracted basic features without considering the specific characteristics of the underwater pile foundation construction scenario. They neither distinguish between operational and non-operational areas nor incorporate on-site physical rules, and they do not specifically enhance the features of small-target potential hazards. In step A31 of this invention, based on the basic characteristics of underwater pile foundation construction, a Unet network with self-attention is used to identify key structures in the water transport scenario. Then, a structure mask matrix of the water transport scenario is obtained through binarization function processing. By utilizing the semantic segmentation capability of the Unet network with self-attention, key operational structures and non-operational areas in the underwater pile foundation construction scenario are distinguished, thereby achieving accurate division of the construction monitoring range. Secondly, step A32 combines the temporal fusion features of step A2 with the water transport scene structure mask matrix of step A31, and obtains the physically constrained construction features of the water pile foundation through the combined operation of the Laplacian operator and graph convolutional network. The Laplacian operator can capture the boundary information of the temporal fusion features and strengthen the feature boundaries of key structures such as the pile foundation body and the working platform. The graph convolutional network can strengthen the spatial correlation of features and conform to the physical distribution law of the on-site construction structure. Combined with the constraints of the water transport scene structure mask matrix, it can effectively filter invalid features in non-working areas and correct features that violate physical rules. Thus, the extracted construction features strictly conform to the physical rules of water pile foundation construction and avoid features that do not conform to the actual site conditions, such as personnel penetrating the pile foundation. Finally, step A33 combines the small target feature gain weights from step A2 with the water-based pile foundation construction features after physical constraints from step A32, and introduces an attention pooling operation to output safety hazard behavior features. The attention pooling operation is used to focus on key hazard behavior features during construction and filter out interference from scene-irrelevant features. The small target feature gain weights further strengthen the occluded and low-proportion small target hazard features, forming a synergistic effect with the attention pooling operation. Ultimately, this makes the hazard behavior features of workers and small machinery prominent, and avoids weakening key hazard features.
[0015] Furthermore, step A4 also includes: A41: Guided by the behavioral characteristics of safety hazards, extract initial text features and weighted text features based on the pre-processed waterway safety specification text data; A42: Cross-modal fusion of weighted text features and safety hazard behavior features is performed to obtain cross-modal fused features; The initial text features are obtained by processing the preprocessed water transport safety specification text data through a bidirectional pre-trained language model. The calculation process of the weighted text features includes: performing feature concatenation operation between the safety hazard behavior features and the initial text features; processing the concatenation result sequentially through a multilayer perceptron and a Sigmoid function; and performing a Hadamard product between the processing result and the initial text features to obtain the weighted text features. The calculation process of the cross-modal fusion feature includes: first, performing a feature concatenation operation between the safety hazard behavior feature and the weighted text feature, and then processing the concatenation result through an attention mechanism to obtain the cross-modal fusion feature.
[0016] It needs to be further explained that safety monitoring of underwater pile foundation construction should not only achieve video behavior recognition, but also determine whether the behavior complies with waterway safety regulations. Existing technologies for monitoring underwater pile foundation construction behavior generally adopt a single visual feature recognition mode. The core approach is to extract only the visual features of construction behavior to complete the identification of whether the behavior exists or not. It does not introduce textual data of waterway safety regulations, nor does it attempt to establish a connection between visual and textual data. Without the guidance of waterway safety regulations, it is impossible to judge the compliance of construction behavior. It can only complete basic behavior recognition and cannot meet the needs of timely identification of violations and prevention of safety hazards in on-site safety management. This invention first uses the safety hazard behavior features output from step A3 as a guide in step A41. Combined with preprocessed waterway safety specification text data, a bidirectional pre-trained language model is used to process the text data and extract initial text features. Then, the safety hazard behavior features are concatenated with the initial text features through feature concatenation. After being processed sequentially by a multilayer perceptron and a sigmoid function, the weighted text features are fused with the initial text features to obtain weighted text features. The bidirectional pre-trained language model is used to capture the semantic rules in the waterway safety specification text, ensuring that the initial text features fully conform to the specification requirements. The core design of the feature concatenation combined with the multilayer perceptron and sigmoid function is to establish a correlation between the initial text features and the visual safety hazard behavior features. Through dynamic weight adjustment, the specification rules related to the current construction behavior are highlighted, while irrelevant text information is weakened. This allows the weighted text features to accurately match the current construction behavior on site, providing targeted specification basis for subsequent compliance judgment and avoiding a disconnect between text features and construction behavior. Building upon this, step A42 concatenates the weighted text features obtained in step A41 with the safety hazard behavior features obtained in step A3, and then processes them through an attention mechanism to obtain cross-modal fusion features. The attention mechanism is used to focus on the associated regions of the safety hazard behavior features and the weighted text features, strengthening their synergistic effect and achieving deep integration of visual information of construction behavior and regulatory text rules. This breaks the limitation of the separation of vision and text in traditional monitoring, allowing the fused features to contain both the visual details of construction behavior and the corresponding regulatory requirements, laying a solid foundation for subsequent behavior compliance matching. Waterway safety regulations contain many specific and strict requirements, such as that workers must not remove their special protective equipment without authorization. The weighted text features in step A41 are used to extract the regulatory rules related to the current construction behavior. The cross-modal fusion features in step A42 can deeply bind the visual features of personnel approaching the restricted area with the text regulatory features, representing the correspondence between the current construction behavior and the regulatory requirements, thus solving the pain point of traditional monitoring that can identify behavior but cannot judge compliance.
[0017] Furthermore, step A5 also includes: A51: Enhance the characteristics of safety hazards to obtain enhanced safety hazard behavior characteristics; A52: Based on the enhanced behavioral characteristics of security risks and cross-modal fusion characteristics, the behavioral compliance matching degree is obtained by calculating cosine similarity and Euclidean distance; The calculation process of the enhanced safety hazard behavior features includes: performing multilayer perceptron processing on the safety hazard behavior features, processing the processing result with the Sigmoid function, and then performing a Hadamard product with the safety hazard behavior features to obtain the enhanced safety hazard behavior features. The calculation process for the behavior compliance matching degree includes: first, performing a Hadamard product between the enhanced safety hazard behavior features and the cross-modal fusion features; then, dividing the result by the L2 norm of the cross-modal fusion features to obtain the cosine matching degree; calculating the difference between the enhanced safety hazard behavior features and the cross-modal fusion features; taking the square of the L2 norm of this difference and processing it with an exponential function; subtracting the processing result from 1 to obtain the Euclidean distance matching degree; and adding the product of 0.6 and the cosine matching degree and the product of 0.4 and the Euclidean distance matching degree to obtain the behavior compliance matching degree.
[0018] It should be further explained that step A51 of this invention performs secondary feature enhancement processing on the safety hazard behavior features output by step A3. After the safety hazard behavior features are processed by a multilayer perceptron, they are normalized and adjusted by the Sigmoid function. The processing result is then fused with the original safety hazard behavior features to obtain enhanced safety hazard behavior features, amplifying the feature differences of minor violations. On this basis, step A52 combines the enhanced safety hazard behavior features obtained by step A51 with the cross-modal fusion features obtained by step A4. It adopts a dual-dimensional calculation method of cosine matching degree and Euclidean distance matching degree, and then obtains the behavior compliance matching degree through weighted fusion. This measures the similarity between the enhanced safety hazard behavior features and the cross-modal fusion features. The higher the similarity, the better the fit between the construction behavior and the waterway safety regulations. This quantifies the fit between the construction behavior and the waterway safety regulations, and realizes the gradient distinction between compliant operation, minor violation and serious violation.
[0019] Furthermore, step A6 also includes: A61: Based on the behavioral compliance matching degree, first subtract the security threshold from the behavioral compliance matching degree, and then perform binarization on the calculation result through a step function to obtain the behavioral compliance result; A62: If the compliance result is greater than 0, it is judged as compliant; otherwise, it is judged as non-compliant. A63: When the safety hazard behavior features are deemed compliant, no action is taken; when the safety hazard behavior features are deemed non-compliant, the cosine similarity between the enhanced safety hazard behavior features and the initial text features is calculated to obtain a text matching score. The highest text matching score is used as an index to extract the content from the corresponding pre-processed waterway safety specification text data, which serves as the basis for determining the non-compliance of waterway construction project safety hazards.
[0020] This invention also discloses an intelligent monitoring system for safety hazards in waterway construction projects based on AI vision, comprising: Video and text data acquisition module: Acquires video stream data of underwater pile foundation construction and text data of waterway safety regulations, and preprocesses them respectively to obtain preprocessed video stream data of underwater pile foundation construction and preprocessed text data of waterway safety regulations; Video temporal fusion module: Based on the preprocessed video stream data of underwater pile foundation construction, extract the foundation features and small target feature gain weights of underwater pile foundation construction, and then calculate the motion trajectory mask matrix and temporal fusion features in sequence; Safety hazard behavior identification module: Based on the foundation characteristics of underwater pile foundation construction, extract the structural mask matrix of the water transport scene and the construction characteristics of underwater pile foundation after physical constraints in sequence, and generate safety hazard behavior characteristics; Video and text fusion module: Guided by the behavioral characteristics of safety hazards, it extracts initial text features and weighted text features from the preprocessed waterway safety regulations text data, and then performs cross-modal fusion to obtain cross-modal fused features; Behavioral compliance matching degree calculation module: It performs feature enhancement on the characteristics of safety hazard behaviors and then combines cross-modal fusion features to calculate the behavioral compliance matching degree; Behavioral compliance result determination module: Calculates the behavioral compliance result based on the behavioral compliance matching degree, calculates the text matching score based on the behavioral compliance result, and extracts the content from the corresponding pre-processed waterway safety specification text data as the basis for determining non-compliance of waterway construction project safety hazards.
[0021] Compared with the prior art, the beneficial effects of the present invention are as follows: (1) This invention effectively solves the problems of poor accuracy in identifying safety hazards and inability to link with safety regulations in traditional AI visual monitoring technology; This invention focuses on targeted processing of construction video streams, strengthens the features of small targets such as workers, filters background interference such as static water surfaces and fixed pile foundations and non-construction dynamic interference, integrates continuous frame features to fully capture coherent construction behavior, and breaks the limitations of single frame extraction; On this basis, it links with water transport safety regulations texts to establish the relationship between construction visual behavior and regulations, and can also clarify the specific basis for violations of safety regulations, providing support for on-site rectification of safety hazards, and meeting the requirements of accuracy and compliance for safety monitoring of water pile foundation construction.
[0022] (2) In view of the problems of existing AI visual monitoring technology, which uses single-frame image feature extraction, small target features are easily lost, static background and non-construction dynamic interference are serious, and it is impossible to capture continuous behavior of water pile foundation construction and can only identify simple safety hazards, this invention combines the characteristics of water pile foundation construction scene to realize the accurate extraction and temporal fusion of video sequence features, effectively strengthen the features of small targets such as workers, and at the same time filter static backgrounds such as static water surface and fixed pile foundation, as well as non-construction dynamic interference such as floating objects on the water surface and unrelated ships in the distance, and capture the dynamic movement trajectory of personnel and equipment in the construction area; at the same time, it integrates multi-frame feature information to fully represent the continuous construction behavior process, breaks the limitation of single-frame extraction, realizes the analysis of continuous construction behavior, and improves the comprehensiveness and accuracy of hazard identification.
[0023] (3) In view of the problem that the existing monitoring technology for underwater pile foundation construction does not divide the working and non-working areas, resulting in distorted monitoring range, features that do not conform to the actual construction, and weakened key hidden danger features, this invention combines the characteristics of underwater pile foundation construction scenarios to construct a system that can automatically divide the key working structures and non-working areas under the physical constraints of the scenario, filter out interference from non-working areas such as floating objects on the water surface and distant passing ships, avoid distorted monitoring range, and ensure that the monitoring focuses on the actual working area; at the same time, it makes the extracted construction features conform to the physical rules of underwater pile foundation construction, avoids features that do not conform to the actual site conditions, such as personnel penetrating the pile foundation, and provides high-quality feature support for safety hazard identification and compliance judgment.
[0024] (4) In view of the problems that existing waterborne pile foundation construction monitoring technology adopts a single visual feature recognition mode, does not link with waterway safety regulations, can only identify construction behavior but cannot judge the compliance of the behavior, and lacks clear basis for violation, this invention constructs a visual-text cross-modal fusion link to realize the association between construction behavior and waterway safety regulations. It can accurately extract semantic rules in the waterway safety regulations text, and through dynamic weight adjustment, make the text features accurately match the current construction behavior, weaken irrelevant text information, and provide targeted regulatory basis for compliance judgment. At the same time, it realizes the deep integration of visual information of construction behavior and regulatory text rule information, and solves the pain point of traditional monitoring that can identify behavior but cannot judge compliance. Attached Figure Description
[0025] Figure 1 A flowchart illustrating an intelligent monitoring method for safety hazards in waterway construction projects based on AI vision, provided by this invention; Figure 2 Example 1 comparing the changes in attention weights of safety specification texts before and after the introduction of safety hazard behavior features, provided by this invention; Figure 3 Example 2 is a comparison of the changes in attention weight of safety specification texts before and after the introduction of safety hazard behavior features, which is provided by the present invention. Detailed Implementation
[0026] The present invention will be further described below with reference to the accompanying drawings, but this is not intended to limit the present invention in any way. Any modifications or substitutions made based on the teachings of the present invention shall fall within the protection scope of the present invention.
[0027] Example 1: An intelligent monitoring method for safety hazards in waterway construction projects based on AI vision, such as... Figure 1 As shown, it includes the following steps: A1: Collect video stream data of underwater pile foundation construction and text data of waterway safety regulations, and preprocess them separately to obtain preprocessed video stream data of underwater pile foundation construction and preprocessed text data of waterway safety regulations, including: A11: Acquire video stream data of underwater pile foundation construction using high-definition industrial cameras or drone equipment; the type of the underwater pile foundation construction video stream data is continuous frame image sequence data; A12: The video stream data of the underwater pile foundation construction is preprocessed by Gaussian filtering for noise reduction, frame rate linear interpolation normalization and histogram equalization to obtain the preprocessed video stream data of the underwater pile foundation construction. A13: Collect waterway safety regulations text data and waterway construction terminology data; after deduplicating the waterway construction terminology data, construct a custom word segmentation dictionary using terminology hierarchical annotation and word frequency weighting; preprocess the waterway safety regulations text data using forced matching word segmentation based on the custom word segmentation dictionary, invalid character removal using regular expressions, and stop word filtering to obtain preprocessed waterway safety regulations text data.
[0028] A2: Based on the preprocessed video stream data of underwater pile foundation construction, extract the foundation features and small target feature gain weights for underwater pile foundation construction, and then calculate the motion trajectory mask matrix and temporal fusion features in sequence, including: A21: Based on the preprocessed video stream data of underwater pile foundation construction, extract the foundation features and small target feature gain weights for underwater pile foundation construction. The calculation method is as follows: ; in, The image in frame s shows the foundation features of the underwater pile foundation construction. For convolutional layers with residual connections, This refers to a single frame image from the preprocessed video stream data of underwater pile foundation construction, where s is the frame index. The gain weights for small target features in the s-th frame image are... For the Sigmoid function, It is a multilayer perceptron. For max pooling; A22: Based on the characteristics of underwater pile foundation construction, inter-frame motion information is captured using the optical flow method, and then a self-attention module is introduced to filter static background interference. The motion trajectory mask matrix is calculated as follows: ; in, Let be the motion trajectory mask matrix of the s-th frame image. For self-attention mechanism, For optical flow calculation function, The foundation features of the underwater pile foundation construction in the (s-1)th frame image; A23: Based on the motion trajectory mask matrix and the foundation features of the underwater pile foundation construction, the foundation features of the underwater pile foundation construction in consecutive frames are fused temporally. The temporal fusion features are calculated using a temporal feature gating fusion mechanism. The calculation method of the temporal feature gating fusion mechanism is as follows: ; in, Let be the inter-frame transpose fusion feature of the s-th frame image. For transpose operation, For Hadama accumulation, Let be the temporal gating weight for the s-th frame. It is a gated multilayer sensor. For graph convolutional networks, Let be the temporal fusion feature of the s-th frame image, and ⊕ be the element-wise addition operation.
[0029] In this embodiment, the specific parameter settings for the neural network modules involved in steps A21, A22, and A23 are as follows: In step A21, the convolutional layer with residual connections consists of four convolutional sub-layers connected in series. Each convolutional sub-layer uses a 3×3 convolutional kernel. The number of convolutional kernels is 64, 128, 256, and 512 respectively. The stride is set to 1. The padding method is SAME padding. The activation function is ReLU. The residual connections use the shortcut connection method. The max pooling module uses a 2×2 pooling kernel with a step size of 2 and VALID filling method. The multilayer perceptrons in steps A21 and A22 both employ a 3-layer fully connected structure. In step A21, the MLP has 256 input neurons, 128 and 64 intermediate neurons respectively, and 64 output neurons. In step A22, the MLP has 512 input neurons, 256 and 128 intermediate neurons respectively, and 256 output neurons. The activation function for both MLP layers is ReLU. The self-attention module in step A22 adopts a multi-head self-attention structure with 8 attention heads and 512 input feature dimensions. The graph convolutional network in step A23 adopts a 2-layer graph convolutional structure with an input feature dimension of 512, a hidden layer unit number of 512, an output feature dimension of 512, a ReLU activation function, and an adjacency matrix generated adaptively using feature similarity. A gated multilayer perceptron consists of two fully connected layers and one gated unit. The number of neurons in each fully connected layer is 512, and the gated unit uses the Sigmoid activation function.
[0030] A3: Based on the foundation characteristics of underwater pile foundation construction, the structural mask matrix of the water transport scenario and the physically constrained underwater pile foundation construction characteristics are extracted sequentially, and safety hazard behavior characteristics are generated, including: A31: Based on the foundation characteristics of underwater pile foundation construction, key structures in the water transport scenario are identified using an attention-based semantic segmentation network. The structure mask matrix of the water transport scenario is calculated as follows: ; in, Let be the water transport scene structure mask matrix of the s-th frame image. For a self-attention Unet network, Binarization function; A32: Based on the temporal fusion features and the structural mask matrix of the water transport scene, feature optimization is performed to obtain the physically constrained construction features of the waterborne pile foundation. The calculation method is as follows: ; in, For the constrained optimization features of the s-th frame image, For the Laplace operator, The physical constraints of the s-th frame image define the construction features of the underwater pile foundation. A33: Based on the feature gain weights of small targets and the construction characteristics of underwater pile foundations after physical constraints, attention pooling is introduced to highlight key features and output safety hazard behavior features. The calculation method is as follows: ; in, The security risk behavior characteristics of the s-th frame image, This is an attention pooling operation.
[0031] Specifically, for scenarios involving multi-scale superposition and blurred edges of key structures in waterborne pile foundation construction, this invention also provides a method for calculating the structural mask matrix of water transport scenarios that integrates multi-scale feature enhancement and morphological optimization, replacing step A31. The calculation method is as follows: ; in, For morphological closing operations, a 3×3 structuring element is used to fill tiny holes in the mask matrix, smooth structural edges, and optimize the spatial integrity of the mask matrix.
[0032] A4: Guided by behavioral characteristics of safety hazards, based on the preprocessed waterway safety regulations text data, initial text features and weighted text features are extracted, and then cross-modal fusion is performed to obtain cross-modal fused features, including: A41: Guided by the behavioral characteristics of safety hazards, initial text features and weighted text features are extracted from the pre-processed waterway safety regulations text data. The calculation method is as follows: ; in, As initial text features, For bidirectional pre-trained language models, For the preprocessed text data of waterway safety regulations, For weighted text features, For feature splicing operations; A42: Cross-modal fusion is performed on weighted text features and security risk behavior features to obtain cross-modal fused features. The calculation method is as follows: ; in, Let be the cross-modal fusion feature of the s-th frame image.
[0033] For the safety compliance monitoring scenario of underwater pile foundation construction sites, this invention visualizes and verifies the weight allocation effect of the attention mechanism. The input text of the waterway safety regulations for this verification is the "Technical Specification for Safety Protection in Waterway Engineering Construction" which includes "5.1.5 Personnel entering the construction site must wear safety helmets. During operation, they must correctly wear and use personal protective equipment and tools" and "5.1.8 Safety protection facilities, signs, warning signs, etc. at the construction site shall not be removed or moved without authorization. If removal is necessary, it shall be approved by the person in charge of construction." Safety hazard behavior characteristics are also input. Without incorporating video features, the multimodal attention mechanism assigns weights based solely on text features. In this case, the attention weights for terms related to personnel protection, such as "correctly wearing and using personal protective equipment and tools," are at an average level, corresponding to a plain blue block in the visualization. However, when the video stream detects the dynamic behavior of construction workers moving from the work platform to the pile foundation, the multimodal attention mechanism automatically integrates and fuses the video dynamic features with the text features. At this point, the attention weights for core terms such as "correctly wearing and using personal protective equipment and tools" and "safety protection facilities" significantly increase, resulting in these terms being highlighted in orange in the visualization. Figure 2 , Figure 3 As shown; By introducing dynamic features from videos, it is possible to accurately focus on safety regulations and core terms related to the current construction activities, providing a clear weighting guide for subsequent compliance matching and effectively improving the accuracy and interpretability of compliance judgments.
[0034] A5: For behavioral characteristics of safety hazards, feature enhancement is performed, and then combined with cross-modal fusion features to calculate the behavioral compliance matching degree, including: A51: Enhance the behavioral characteristics of safety hazards to obtain the enhanced behavioral characteristics of safety hazards. The calculation method is as follows: ; in, The enhanced security risk behavior features of the s-th frame image; A52: Based on the enhanced behavioral characteristics of security risks and cross-modal fusion features, the behavioral compliance matching degree is obtained by calculating cosine similarity and Euclidean distance. The calculation method is as follows: ; Let be the cosine matching degree of the s-th frame image. Let be the Euclidean distance matching degree of the s-th frame image. Represents an exponential function. Let the behavior compliance matching degree of the image in frame s be , To obtain the L2 norm, To take the square of the L2 norm.
[0035] A6: Based on the behavioral compliance matching degree, calculate the behavioral compliance result, and based on the behavioral compliance result, calculate the text matching score. Extract the corresponding content from the preprocessed waterway safety specification text data as the basis for determining non-compliance in waterway construction project safety hazard assessments, including: A61: Based on the behavioral compliance matching degree, first subtract the security threshold from the behavioral compliance matching degree, and then perform binarization on the calculation result through a step function to obtain the behavioral compliance result; A62: If the compliance result is greater than 0, it is judged as compliant; otherwise, it is judged as non-compliant. A63: When the safety hazard behavior features are deemed compliant, no action is taken; when the safety hazard behavior features are deemed non-compliant, the cosine similarity between the enhanced safety hazard behavior features and the initial text features is calculated to obtain a text matching score. The highest text matching score is used as an index to extract the content from the corresponding pre-processed waterway safety specification text data, which serves as the basis for determining the non-compliance of waterway construction project safety hazards.
[0036] The neural network-related modules involved in this invention, such as convolutional layers, self-attention modules, graph convolutional networks, UNet networks with self-attention, multilayer perceptrons, attention pooling operations, and bidirectional pre-trained language models, are trained as follows: Based on preprocessed video stream data of underwater pile foundation construction and text data of waterway safety regulations, a training dataset with manual annotation was constructed. The compliance results of construction behavior corresponding to video frames and the matching waterway safety regulations text entries were accurately annotated. The ratio of training set, validation set and test set was 7:2:1. Among them, the bidirectional pre-trained language model was fine-tuned by transfer learning using BERT-base pre-trained weights, the Unet network with self-attention was fine-tuned by ImageNet pre-trained weights, and the remaining modules were trained by randomly initializing weights. The entire network adopted an end-to-end training method, with the deviation between the predicted value of behavior compliance matching degree and the actual value of manual annotation as the core training objective. During training, an adaptive moment estimation optimizer was used to iteratively update the network weights. The learning rate was dynamically adjusted using cosine annealing, with an initial learning rate of 1e-4, which gradually decreased to 1e-6 in the later stages of training. The mean squared error loss function was used as the loss function. To meet the multi-task training requirements of scene structure recognition, feature optimization, and compliance matching, a weighted fusion loss function was used to complete joint optimization. The weights of scene structure recognition loss, feature optimization loss, and compliance matching loss were 0.3, 0.3, and 0.4, respectively. Training iterations were performed using a mini-batch stochastic gradient descent method, with a batch size of 32. To avoid model overfitting, L2 regularization and random deactivation regularization strategies are used to constrain network weights. The regularization coefficient is set to 1e-5 and the deactivation probability is set to 0.2. During training, the accuracy of the construction behavior compliance judgment in the validation set is used as the model convergence evaluation index. Training is terminated when the improvement of the validation set index is less than 0.5% for 10 consecutive rounds, and the optimal weights of the network are saved.
[0037] Example 2: This invention also discloses an intelligent monitoring system for safety hazards in waterway construction projects based on AI vision, comprising: Video and text data acquisition module: Acquires video stream data of underwater pile foundation construction and text data of waterway safety regulations, and preprocesses them respectively to obtain preprocessed video stream data of underwater pile foundation construction and preprocessed text data of waterway safety regulations; Video temporal fusion module: Based on the preprocessed video stream data of underwater pile foundation construction, extract the foundation features and small target feature gain weights of underwater pile foundation construction, and then calculate the motion trajectory mask matrix and temporal fusion features in sequence; Safety hazard behavior identification module: Based on the foundation characteristics of underwater pile foundation construction, extract the structural mask matrix of the water transport scene and the construction characteristics of underwater pile foundation after physical constraints in sequence, and generate safety hazard behavior characteristics; Video and text fusion module: Guided by the behavioral characteristics of safety hazards, it extracts initial text features and weighted text features from the preprocessed waterway safety regulations text data, and then performs cross-modal fusion to obtain cross-modal fused features; Behavioral compliance matching degree calculation module: It performs feature enhancement on the characteristics of safety hazard behaviors and then combines cross-modal fusion features to calculate the behavioral compliance matching degree; Behavioral compliance result determination module: Calculates the behavioral compliance result based on the behavioral compliance matching degree, calculates the text matching score based on the behavioral compliance result, and extracts the content from the corresponding pre-processed waterway safety specification text data as the basis for determining non-compliance of waterway construction project safety hazards.
[0038] It should be noted that the sequence numbers of the above embodiments of the present invention are merely for descriptive purposes and do not represent the superiority or inferiority of the embodiments. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, apparatus, article, or method that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, apparatus, article, or method. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, apparatus, article, or method that includes that element.
[0039] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present invention.
[0040] The above are merely preferred embodiments of the present invention and do not limit the scope of the patent. Any equivalent structural or procedural transformations made based on the description and drawings of the present invention, or direct or indirect applications in other related technical fields, are similarly included within the scope of patent protection of the present invention.
Claims
1. A method for intelligent monitoring of safety hazards in waterway construction projects based on AI vision, characterized in that, Includes the following steps: A1: Collect video stream data of underwater pile foundation construction and text data of waterway safety specifications, and preprocess them respectively to obtain preprocessed video stream data of underwater pile foundation construction and preprocessed text data of waterway safety specifications; A2: Based on the preprocessed video stream data of underwater pile foundation construction, extract the foundation features and small target feature gain weights of underwater pile foundation construction, and then calculate the motion trajectory mask matrix and temporal fusion features in sequence; A3: Based on the basic characteristics of underwater pile foundation construction, extract the structural mask matrix of the water transport scene and the construction characteristics of underwater pile foundation after physical constraints in sequence, and generate safety hazard behavior characteristics; A4: Guided by the behavioral characteristics of safety hazards, based on the preprocessed waterway safety specification text data, initial text features and weighted text features are extracted, and then cross-modal fusion is performed to obtain cross-modal fused features; A5: Enhance the characteristics of safety hazard behaviors and then combine them with cross-modal fusion features to calculate the compliance matching degree of the behavior; A6: Calculate the behavioral compliance result based on the behavioral compliance matching degree, calculate the text matching score based on the behavioral compliance result, and extract the content from the corresponding pre-processed waterway safety specification text data as the basis for determining non-compliance of waterway construction project safety hazards.
2. The intelligent monitoring method for safety hazards in waterway construction projects based on AI vision as described in claim 1, characterized in that, Step A1 includes: A11: Acquire video stream data of underwater pile foundation construction using high-definition industrial cameras or drone equipment; the type of the underwater pile foundation construction video stream data is continuous frame image sequence data; A12: The video stream data of the underwater pile foundation construction is preprocessed by Gaussian filtering for noise reduction, frame rate linear interpolation normalization and histogram equalization to obtain the preprocessed video stream data of the underwater pile foundation construction. A13: Collect waterway safety standard text data and waterway construction terminology data; after deduplicating the waterway construction terminology data, construct a custom word segmentation dictionary through terminology hierarchical annotation and word frequency weighting; preprocess the waterway safety standard text data using forced matching word segmentation based on the custom word segmentation dictionary, invalid character removal using regular expressions, and stop word filtering to obtain preprocessed waterway safety standard text data.
3. The intelligent monitoring method for safety hazards in waterway construction projects based on AI vision as described in claim 2, characterized in that, Step A2 includes: A21: Based on the preprocessed video stream data of underwater pile foundation construction, extract the foundation features and small target feature gain weights of underwater pile foundation construction. A22: Based on the characteristics of underwater pile foundation construction, inter-frame motion information is captured by optical flow method, and then a self-attention module is introduced to filter static background interference, and the motion trajectory mask matrix is calculated. A23: Based on the motion trajectory mask matrix and the characteristics of the underwater pile foundation construction, the characteristics of the underwater pile foundation construction in consecutive frames are fused in a temporal sequence. The temporal fusion characteristics are calculated through a temporal feature gating fusion mechanism.
4. The intelligent monitoring method for safety hazards in waterway construction projects based on AI vision according to claim 3, characterized in that, Step A2 includes: First, the single-frame image in the preprocessed video stream data of underwater pile foundation construction is processed by a convolutional layer with residual connection to obtain the underwater pile foundation construction features of the current frame image. Then, the underwater pile foundation construction features are subjected to max pooling processing, and then processed by a multilayer perceptron and a sigmoid function to obtain the small target feature gain weight of the s-th frame image. Based on the water-based pile foundation construction features of the current frame image and the water-based pile foundation construction features of the previous frame image, the inter-frame motion information is captured by the optical flow calculation function, and then the inter-frame motion information is processed by the self-attention mechanism. After processing by the multilayer perceptron and the Sigmoid function, the motion trajectory mask matrix of the current frame image is obtained. The inter-frame transposed fusion features of the current frame image and the previous frame image are transposed and then multiplied by a Hadamard product to obtain the inter-frame transposed fusion features of the current frame image. The inter-frame transposed fusion features are processed by a graph convolutional network and then multiplied by a Hadamard product with the motion trajectory mask matrix of the current frame image. After processing by a gated multilayer perceptron and a Sigmoid function, the temporal gating weights of the current frame are obtained. The inter-frame transposed fusion features of the current frame image are multiplied by a Hadamard product with the temporal gating weights. At the same time, the inter-frame transposed fusion features of the previous frame image are multiplied by a Hadamard product with the result of 1 minus the temporal gating weights. The two results are then added element-wise to obtain the temporal fusion features of the current frame image.
5. The intelligent monitoring method for safety hazards in waterway construction projects based on AI vision as described in claim 4, characterized in that, Step A3 includes: A31: Based on the characteristics of underwater pile foundation construction, key structures in the water transport scenario are identified through an attention-based semantic segmentation network, and the structural mask matrix of the water transport scenario is calculated. A32: Based on the temporal fusion features and the water transport scenario structure mask matrix, feature optimization is performed to obtain the physically constrained construction features of the water pile foundation; A33: Based on the gain weight of small target features and the construction characteristics of underwater pile foundations after physical constraints, the dimensions are adaptively adjusted, attention pooling operation is introduced to highlight key features, and safety hazard behavior features are output.
6. The intelligent monitoring method for safety hazards in waterway construction projects based on AI vision as described in claim 5, characterized in that, Step A3 includes: Based on the characteristics of underwater pile foundation construction, key structures in the water transport scene are identified through a Unet network with self-attention. The identification results are then processed by a binarization function to obtain the water transport scene structure mask matrix of the current frame image. Based on the temporal fusion features and the water transport scene structure mask matrix of the current frame image, the temporal fusion features are first processed by the Laplacian operator, and then the processing result is processed by a graph convolutional network. After that, the Hadamard product is performed with the water transport scene structure mask matrix of the current frame image and the result is negative to obtain the constrained optimization features of the current frame image. The temporal fusion features and the constrained optimization features are then added element-wise to obtain the physically constrained water pile foundation construction features of the current frame image. Attention pooling is applied to the construction features of the underwater pile foundation after physical constraints in the current frame, and then the Hadamard product is performed with the small target feature gain weight to output the safety hazard behavior features of the current frame image.
7. The intelligent monitoring method for safety hazards in waterway construction projects based on AI vision as described in claim 6, characterized in that, The A4 step includes: A41: Guided by the behavioral characteristics of safety hazards, extract initial text features and weighted text features based on the pre-processed waterway safety specification text data; A42: Cross-modal fusion of weighted text features and safety hazard behavior features is performed to obtain cross-modal fused features; The initial text features are obtained by processing the preprocessed water transport safety specification text data through a bidirectional pre-trained language model. The calculation process of the weighted text features includes: performing feature concatenation operation between the safety hazard behavior features and the initial text features; processing the concatenation result sequentially through a multilayer perceptron and a Sigmoid function; and performing a Hadamard product between the processing result and the initial text features to obtain the weighted text features. The calculation process of the cross-modal fusion feature includes: first, performing a feature concatenation operation between the safety hazard behavior feature and the weighted text feature, and then processing the concatenation result through an attention mechanism to obtain the cross-modal fusion feature.
8. The intelligent monitoring method for safety hazards in waterway construction projects based on AI vision as described in claim 7, characterized in that, Step A5 includes: A51: Enhance the characteristics of safety hazard behaviors to obtain enhanced safety hazard behaviors; A52: Based on the enhanced behavioral characteristics of security risks and cross-modal fusion characteristics, the behavioral compliance matching degree is obtained by calculating cosine similarity and Euclidean distance; The calculation process of the enhanced safety hazard behavior features includes: performing multilayer perceptron processing on the safety hazard behavior features, processing the processing result with the Sigmoid function, and then performing a Hadamard product with the safety hazard behavior features to obtain the enhanced safety hazard behavior features. The calculation process for the behavior compliance matching degree includes: first, performing a Hadamard product between the enhanced safety hazard behavior features and the cross-modal fusion features; then, dividing the result by the L2 norm of the cross-modal fusion features to obtain the cosine matching degree; calculating the difference between the enhanced safety hazard behavior features and the cross-modal fusion features; taking the square of the L2 norm of this difference and processing it with an exponential function; subtracting the processing result from 1 to obtain the Euclidean distance matching degree; and adding the product of 0.6 and the cosine matching degree and the product of 0.4 and the Euclidean distance matching degree to obtain the behavior compliance matching degree.
9. The intelligent monitoring method for safety hazards in waterway construction projects based on AI vision according to claim 8, characterized in that, Step A6 includes: A61: Based on the behavioral compliance matching degree, first subtract the security threshold from the behavioral compliance matching degree, and then perform binarization on the calculation result through a step function to obtain the behavioral compliance result; A62: If the compliance result is greater than 0, it is judged as compliant; otherwise, it is judged as non-compliant. A63: When the safety hazard behavior features are deemed compliant, no action is taken; when the safety hazard behavior features are deemed non-compliant, the cosine similarity between the enhanced safety hazard behavior features and the initial text features is calculated to obtain a text matching score. The highest text matching score is used as an index to extract the content from the corresponding pre-processed waterway safety specification text data, which serves as the basis for determining the non-compliance of waterway construction project safety hazards.
10. An intelligent monitoring system for safety hazards in waterway construction projects based on AI vision, characterized in that, include: Video and text data acquisition module: Acquires video stream data of underwater pile foundation construction and text data of waterway safety regulations, and preprocesses them respectively to obtain preprocessed video stream data of underwater pile foundation construction and preprocessed text data of waterway safety regulations; Video temporal fusion module: Based on the preprocessed video stream data of underwater pile foundation construction, extract the foundation features and small target feature gain weights of underwater pile foundation construction, and then calculate the motion trajectory mask matrix and temporal fusion features in sequence; Safety hazard behavior identification module: Based on the foundation characteristics of underwater pile foundation construction, extract the structural mask matrix of the water transport scene and the construction characteristics of underwater pile foundation after physical constraints in sequence, and generate safety hazard behavior characteristics; Video and text fusion module: Guided by the behavioral characteristics of safety hazards, it extracts initial text features and weighted text features from the preprocessed waterway safety regulations text data, and then performs cross-modal fusion to obtain cross-modal fused features; Behavioral compliance matching degree calculation module: It performs feature enhancement on the characteristics of safety hazard behaviors and then combines cross-modal fusion features to calculate the behavioral compliance matching degree; Behavioral compliance result determination module: Calculates behavioral compliance result based on behavioral compliance matching degree, calculates text matching score based on behavioral compliance result, and extracts content from the corresponding preprocessed waterway safety specification text data as the basis for determining non-compliance of safety hazards in waterway construction projects; thereby realizing the intelligent monitoring method for safety hazards in waterway construction projects based on AI vision as described in any one of claims 1-9.