A foundation pit support monitoring method, device and equipment based on machine learning

By using machine learning methods, combined with visual pyramids and graph convolutional networks, the problems of low efficiency of manual inspection and high cost of sensors in foundation pit support detection are solved, realizing real-time and accurate monitoring of foundation pit support structures, and improving construction safety and efficiency.

CN119274140BActive Publication Date: 2025-10-28THREE GORGES HI TECH INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411403993.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-09
Publication Date
2025-10-28
Estimated Expiration
2044-10-09

AI Technical Summary

Technical Problem

Existing methods for detecting foundation pit support rely on manual inspections, which are inefficient and have poor accuracy and consistency. Sensor applications are costly and data processing is complex. They also fail to fully consider the positional relationship between the foundation pit and the support structure, resulting in insufficient monitoring and evaluation.

Method used

A machine learning-based approach is adopted, using the Visual Pyramid (PVT) network and Feature Pyramid (FPN) network to extract multi-scale features. Combined with Graph Convolutional Network (GCN) and attention mechanism, a dynamic adjacency matrix is ​​constructed, which integrates the semantic features of the foundation pit and the support structure to achieve real-time and accurate monitoring of the support structure.

Benefits of technology

It enables real-time and accurate monitoring of the foundation pit support structure, improves the safety and efficiency of the construction process, reduces human error and computational complexity, and enhances adaptability to complex environments and detection accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119274140B_ABST
    Figure CN119274140B_ABST
Patent Text Reader

Abstract

This invention relates to the field of artificial intelligence and discloses a machine learning-based method, apparatus, and equipment for monitoring foundation pit support. The method involves acquiring foundation pit image data; extracting multi-scale features from the image data using a Visual Pyramid (PVT) network; and fusing information from different scales using a Feature Pyramid (FPN) network to generate multi-scale feature vectors. These features are then flattened and global average pooling is performed to obtain a global semantic vector for each scale. Cross-scale semantic feature vectors are generated through channel concatenation and 1*1 convolution. Based on this, a dynamic adjacency matrix is ​​constructed, and a Graph Convolutional Network (GCN) is used to infer the semantic features of the relationship between the foundation pit and the support structure. Finally, an attention mechanism is used to generate enhanced features to determine whether the quantity and location of the support structure meet construction requirements, achieving real-time and accurate foundation pit support monitoring.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence, and in particular to a method, apparatus and equipment for monitoring foundation pit support based on machine learning. Background Technology

[0002] Excavation pit support, through structural and technical measures, ensures the stability and safety of the excavation pit or underground structure during construction, prevents soil collapse, and protects the surrounding environment. The stability of the support structure is crucial to worker safety and project progress. Traditional support inspection methods mainly rely on manual inspections; however, with technological advancements, more and more sensors are being introduced to detect the stability of the support structure.

[0003] However, existing support inspection methods have many problems. First, manual inspection relies on the subjective judgment of staff, which is not only inefficient but also makes it difficult to guarantee the accuracy and consistency of inspection results due to human interference. Second, manual inspection has poor real-time performance and cannot cope with rapidly changing construction environments. Furthermore, while the introduction of sensors has improved the automation of inspection, its application also faces challenges. Sensor installation and maintenance costs are high, data processing is complex, and comprehensive analysis combining data from multiple sensors is required. This multi-dimensional data processing demands sophisticated technical means to ensure the accuracy and usability of the data. Existing multi-target recognition algorithms often neglect the positional relationships between targets and fail to fully consider the positional relationship between the foundation pit and the support structure, resulting in deficiencies in the monitoring and evaluation of the support structure. Summary of the Invention

[0004] In view of this, this application provides a machine learning-based method for monitoring foundation pit support, which solves the technical problem that the existing technology does not consider the positional relationship between the foundation pit and the support structure, resulting in insufficient monitoring and evaluation of the support structure.

[0005] According to a first aspect of this application, a machine learning-based method for monitoring foundation pit support is provided, comprising:

[0006] Acquire foundation pit image data;

[0007] The Visual Pyramid (PVT) network extracts features at different scales from image data, and the Feature Pyramid (FPN) network fuses information at different scales to generate multi-scale feature vectors.

[0008] Multi-scale features are flattened and global average pooling is calculated to generate global semantic vectors for each scale. Through channel concatenation and 1*1 convolution operations, cross-scale semantic feature vectors are generated.

[0009] A dynamic adjacency matrix is ​​constructed based on multi-scale feature vectors and cross-scale semantic feature vectors, and the semantic features of the relationship between the foundation pit and the support structure are inferred through graph convolutional network (GCN).

[0010] The attention weights between the foundation pit and the support structure are calculated through an attention mechanism, and the semantic features of the relationship between the foundation pit and the support structure are fused with multi-scale features to generate enhanced features.

[0011] Based on the enhanced features, it is determined whether the quantity and location of the support structure meet the construction requirements.

[0012] According to a second aspect of this application, a machine learning-based foundation pit support monitoring device is provided, comprising:

[0013] The acquisition module is used to acquire foundation pit image data;

[0014] The multi-scale feature extraction module is used to extract features of different scales from image data through the Visual Pyramid (PVT) network and fuse information of different scales through the Feature Pyramid (FPN) network to generate multi-scale feature vectors.

[0015] The cross-scale semantic perception module is used to flatten features at multiple scales and calculate global average pooling to generate global semantic vectors for each scale. Through channel concatenation and 1*1 convolution operations, it generates cross-scale semantic feature vectors.

[0016] The dynamic relation graph reasoning module is used to construct a dynamic adjacency matrix based on multi-scale feature vectors and cross-scale semantic feature vectors, and to reason about the semantic features of the relationship between the foundation pit and the support structure through the graph convolutional network GCN.

[0017] The semantic attention fusion module is used to calculate the attention weight between the foundation pit and the support structure through the attention mechanism, and to fuse the semantic features of the relationship between the foundation pit and the support structure with multi-scale features to generate enhanced features;

[0018] The foundation pit support detection module is used to determine whether the quantity and location of the support structure meet the construction requirements based on the enhanced features.

[0019] According to a third aspect of this application, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described machine learning-based foundation pit support monitoring method.

[0020] By employing the above technical solutions, this application provides a machine learning-based method, device, equipment, and medium for monitoring foundation pit support. The method involves acquiring foundation pit image data; extracting features at different scales from the image data using a Visual Pyramid (PVT) network; fusing information from different scales using a Feature Pyramid (FPN) network to generate multi-scale feature vectors; flattening the multi-scale features and calculating global average pooling to generate global semantic vectors for each scale; generating cross-scale semantic feature vectors through channel concatenation and 1*1 convolution operations; constructing a dynamic adjacency matrix based on the multi-scale feature vectors and cross-scale semantic feature vectors; inferring the semantic features of the relationship between the foundation pit and the support structure using a Graph Convolutional Network (GCN); calculating the attention weights between the foundation pit and the support structure using an attention mechanism; fusing the semantic features of the relationship between the foundation pit and the support structure with the multi-scale features to generate enhanced features; and determining whether the quantity and location of the support structure meet construction requirements based on the enhanced features. This invention combines advanced technologies such as machine learning, cross-scale feature extraction, dynamic relationship reasoning, and semantic and visual fusion to achieve real-time and accurate monitoring of foundation pit support structures, effectively improving the safety and efficiency of the construction process.

[0021] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description

[0022] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0023] Figure 1 This illustration shows an application scenario diagram of a machine learning-based foundation pit support monitoring method provided in an embodiment of this application;

[0024] Figure 2 A flowchart illustrating a machine learning-based foundation pit support monitoring method provided in an embodiment of this application is shown.

[0025] Figure 3 The network structure of a machine learning-based foundation pit support monitoring method provided in an embodiment of this application is shown.

[0026] Figure 4 A semantic relationship diagram of the foundation pit support location provided in the embodiments of this application is shown;

[0027] Figure 5A schematic diagram of a machine learning-based foundation pit support monitoring device provided in an embodiment of this application is shown. Detailed Implementation

[0028] The specific implementation of this application will be described in detail below with reference to the accompanying drawings and embodiments. It should be noted that, unless otherwise specified, the embodiments and features in the embodiments of this application can be combined with each other.

[0029] The present invention provides a machine learning-based method for monitoring foundation pit support, which can be applied to, for example... Figure 1 In the scenario shown, a camera is installed at the construction site to monitor the foundation pit in real time, collect and store images of the pit, and input images. The server executes model training and detection processes, using a network to extract features and infer relationships, generating candidate regions and predicting the category and bounding box of the support structure. A loss function is used to measure the difference between the predicted results and the true annotations. Backpropagation is used to update the network weights by calculating gradients. During training, the network continuously optimizes until convergence. Finally, the model's performance is evaluated on a validation set. Through this end-to-end training process, the foundation pit support detection model in this embodiment can learn to detect objects at different scales and utilize the relationship information between objects to improve the accuracy and robustness of detection. The server executes a machine learning-based foundation pit support monitoring method. This invention acquires foundation pit image data; extracts features at different scales from the image data using a Visual Pyramid (PVT) network, and fuses information from different scales using a Feature Pyramid (FPN) network to generate multi-scale feature vectors; flattens the multi-scale features and calculates global average pooling to generate global semantic vectors for each scale; and generates cross-scale semantic feature vectors through channel concatenation and 1*1 convolution operations; constructs a dynamic adjacency matrix based on the multi-scale feature vectors and cross-scale semantic feature vectors, and infers the semantic features of the relationship between the foundation pit and the support structure using a Graph Convolutional Network (GCN); calculates the attention weights between the foundation pit and the support structure using an attention mechanism, and fuses the semantic features of the relationship between the foundation pit and the support structure with the multi-scale features to generate enhanced features; based on the enhanced features, it determines whether the quantity and location of the support structure meet the construction requirements. This embodiment of the invention combines advanced technologies such as machine learning, cross-scale feature extraction, dynamic relationship reasoning, and semantic and visual fusion to achieve real-time and accurate monitoring of foundation pit support structures, effectively improving the safety and efficiency of the construction process.

[0030] The present invention will now be described in detail through specific embodiments.

[0031] Example 1:

[0032] like Figure 2As shown in the figure, a machine learning-based method for monitoring foundation pit support is provided in an embodiment of the present invention, comprising:

[0033] Step 201: Obtain foundation pit image data;

[0034] Specifically, step 201 includes: acquiring video stream data of the foundation pit in real time, extracting static images from the video stream, and performing noise reduction and contrast enhancement operations on the static images.

[0035] Step 202: Extract features of different scales from image data through the Visual Pyramid (PVT) network, and fuse information of different scales through the Feature Pyramid (FPN) network to generate multi-scale feature vectors.

[0036] Step 202 uses PVT v2 (Pyramid Vision Transformer v2) as the backbone network to extract multi-scale feature maps through different stages of the PVT network. These feature maps capture information about the different sizes and resolutions of objects in the image. By using a Feature Pyramid Network (FPN), the multi-scale feature maps are further fused, ensuring that each scale contains rich semantic information. Multi-scale feature maps are collected through different stages of the PVT network, and at each stage, these features are passed to higher layers. High-level feature maps are upsampled and fused with low-level feature maps to generate multi-scale feature maps.

[0037] Step 202 specifically includes:

[0038] Step 202-1: Use the Visual Pyramid (PVT) network to extract multi-scale feature maps through multiple stages, where the feature maps output by each stage have different resolutions and number of channels.

[0039] Step 202-2: The Feature Pyramid Network (FPN) performs up-and-down path fusion on these feature maps to generate multi-scale feature maps with information at different scales.

[0040] Step 203: Flatten the multi-scale features and calculate global average pooling to generate global semantic vectors for each scale. Generate cross-scale semantic feature vectors through channel concatenation and 1*1 convolution operations.

[0041] Step 203 involves the interaction of key semantic information between feature maps at different scales. By enabling information interaction and integration between feature maps at different scales, cross-scale semantic features are generated. These features can enhance the correlation between objects in the image and help the network better understand the image content.

[0042] Step 203 specifically includes:

[0043] Step 203-1: Extract global semantic information from the multi-scale feature map and flatten the multi-scale feature map into a one-dimensional vector;

[0044] Step 203-2: Perform global average pooling on each flattened feature map to generate a global semantic vector for each scale;

[0045] Step 203-3: Concatenate the global semantic vectors of all scales along the channel dimension to generate a high-dimensional cross-scale semantic vector;

[0046] Step 203-4: Reduce the dimensionality of the concatenated semantic vectors using 1*1 convolution to generate cross-scale semantic feature vectors.

[0047] Step 204: Construct a dynamic adjacency matrix based on multi-scale feature vectors and cross-scale semantic feature vectors, and infer the semantic features of the relationship between the foundation pit and the support structure through a graph convolutional network (GCN).

[0048] Step 204 constructs a dynamic relationship graph. By generating a dynamic graph based on specific object relationships in the image, the network can more accurately infer the interaction relationships between objects. Through a Graph Convolutional Network (GCN), the node and edge information in the relationship graph can be further propagated and updated, improving the network's object detection capabilities.

[0049] Step 204 specifically includes:

[0050] Step 204-1: Generate transformed features and class activation maps based on multi-scale feature vectors to generate node representations;

[0051] Step 204-2: Through convolution operation, integrate multi-scale features and cross-scale semantic features to generate an adjacency matrix representing the dynamic relationship between objects;

[0052] Step 204-3: Using a graph convolutional network (GCN), edges are constructed based on the adjacency matrix propagation information and node representations to generate semantic features of the relationship between the foundation pit and the support structure.

[0053] Step 205: Calculate the attention weight between the foundation pit and the support structure through the attention mechanism, and fuse the semantic features of the relationship between the foundation pit and the support structure with multi-scale features to generate enhanced features;

[0054] Step 205 uses a semantic attention mechanism to combine previously extracted semantic features with visual features. By calculating the similarity between visual and semantic features, important image regions are assigned higher weights, thereby improving the accuracy of object detection. The fused feature map is then input into the subsequent detector for the classification and location regression of the support structure. Subsequently, a Region Proposal Network (RPN) is used to receive multi-scale features as input. These features contain image information at different scales. The RPN traverses the feature map using a sliding window and generates multiple candidate regions at each location. The candidate regions have different sizes and aspect ratios, which can cover objects of different sizes in the image. For each generated candidate box, the RPN predicts whether the region contains an object and generates a confidence score for each candidate box. In this way, the RPN can filter out high-confidence region proposals as input for subsequent object detection. In addition to predicting whether the candidate boxes contain objects, the RPN also performs boundary regression on these candidate boxes to further accurately locate the position of the support structure.

[0055] Step 205 specifically includes:

[0056] Step 205-1: Perform dimensionality reduction on the multi-scale features;

[0057] Step 205-2: Perform a reorganization operation on the semantic features and multi-scale features of the relationship between the foundation pit and the support structure to fuse the semantic features and multi-scale features of the relationship between the foundation pit and the support structure.

[0058] Step 205-3: Calculate the similarity between the semantic features of the relationship between the foundation pit and the support structure and the multi-scale features through the Hadamard product to obtain the attention weights;

[0059] The Hadamard product (also known as element-wise product or element-by-element multiplication) is an operation between matrices or tensors. It is calculated by multiplying corresponding elements of two matrices or tensors of the same dimension one by one to form a new matrix or tensor. If there are two matrices A and B of the same dimension, the Hadamard product is defined as follows: Among them, C ij =A ij *B ij Each element of matrix C is equal to the product of the corresponding elements of matrices A and B.

[0060] Step 205-4: Fuse the attention weights with the multi-scale features to generate enhanced features.

[0061] Step 206: Based on the enhanced features, determine whether the quantity and location of the support structure meet the construction requirements.

[0062] Step 206 specifically includes:

[0063] Step 206-1: Use the enhanced feature map to identify the support structure, determine its location in the foundation pit image, and add the bounding box of the support structure.

[0064] Step 206-2: Count the boundary frames of the support structure to determine whether the number of support structures meets the construction requirements;

[0065] Step 206-3: Evaluate the boundary frame of the support structure to determine whether it meets the location setting standards of the support structure in the construction requirements;

[0066] Step 206-4: Output the results showing whether the quantity and location of the support structure meet the construction requirements, and provide relevant testing information.

[0067] This invention provides a machine learning-based method for monitoring foundation pit support. The method involves acquiring foundation pit image data; extracting features at different scales from the image data using a Visual Pyramid (PVT) network; fusing information from different scales using a Feature Pyramid (FPN) network to generate multi-scale feature vectors; flattening the multi-scale features and calculating global average pooling to generate global semantic vectors for each scale; generating cross-scale semantic feature vectors through channel concatenation and 1*1 convolution operations; constructing a dynamic adjacency matrix based on the multi-scale feature vectors and cross-scale semantic feature vectors; and inferring the semantic features of the relationship between the foundation pit and the support structure using a Graph Convolutional Network (GCN); calculating the attention weights between the foundation pit and the support structure using an attention mechanism; and fusing the semantic features of the relationship between the foundation pit and the support structure with the multi-scale features to generate enhanced features; and determining whether the quantity and location of the support structure meet construction requirements based on the enhanced features. This invention combines machine learning, cross-scale feature extraction, dynamic relationship inference, and semantic and visual fusion to achieve real-time and accurate monitoring of foundation pit support structures, effectively improving the safety and efficiency of the construction process.

[0068] Example 2:

[0069] To better illustrate the foundation pit support detection principle and model training process of Embodiment 1 of the present invention, as follows: Figure 3 The diagram shows the network structure of a machine learning-based foundation pit support monitoring method. The network structure in Embodiment 2 of this invention includes multi-scale feature extraction (including PVT v2 network and Feature Pyramid Network (FPN), dynamic relation graph reasoning (including transformed features, class activation mapping, and Graph Convolutional Network (GCN), cross-scale semantic perception, semantic attention fusion, Region Proposal Network (RPN), and a detection head. The main steps of training this network model are as follows:

[0070] 1. Forward propagation:

[0071] In each training iteration, the input image passes through various modules of the network, undergoing forward propagation sequentially to generate object detection results. Multi-scale feature extraction: First, the input image is processed by PVT v2 (Pyramid VisionTransformer v2) or other backbone networks to extract multi-scale feature maps. These feature maps capture information about objects at different scales within the image. The cross-scale semantic perception module, based on the multi-scale feature maps, performs semantic interaction between features of different scales to generate cross-scale semantic features. The dynamic relationship graph inference module constructs a dynamic relationship graph based on the extracted multi-scale and semantic features. Through a graph convolutional network (GCN), information between nodes (semantic representations of the support structure and pit sidewalls) is propagated through an adjacency matrix, inferring the relationships between the support structure and pit sidewalls, further refining their semantic representations. The semantic attention fusion module (SAFM): Finally, the SAFM module combines the semantic representations of the support structure and pit sidewalls with visual features, utilizing an attention mechanism to enhance the feature representation of object detection. RPN generates candidate boxes: The RPN (Region Proposal Network) generates candidate regions for the support structure and pit sidewalls on multi-scale feature maps, and adjusts the boundaries of these candidate boxes through regression. Detection Head: Based on the features of the aforementioned modules, the network's detection head predicts the category and bounding box of the support structure and pit sidewalls. The detection head typically consists of a classification branch and a boundary regression branch. The classification branch is responsible for determining whether an object exists in the candidate box and its category, while the boundary regression branch is responsible for adjusting the position and size of the candidate box.

[0072] 2. Loss Calculation:

[0073] The network's loss function measures the error between the predicted result and the true label, and guides the network's weight updates. The loss function in the model network of this invention mainly includes two parts: Classification loss: used to measure the difference between the predicted class of an object in the candidate box and the true class label. Commonly used classification losses include cross-entropy loss or Focal Loss, which can handle class imbalance problems. Bounding box regression loss: used to measure the difference between the object's predicted bounding box and the ground truth bounding box, typically using Smooth L1 Loss or IoU loss (Intersection over Union Loss). Other losses (optional): In the semantic attention fusion module and the dynamic relation graph inference module, some additional loss functions may be introduced, such as losses related to graph convolutional networks, to ensure that semantic information can be effectively transmitted and used for detection.

[0074] 3. Backpropagation:

[0075] After the network computes the prediction and loss function through forward propagation, it uses a backpropagation algorithm (such as gradient descent) to update the network weights. By calculating the gradient of the loss function with respect to the network parameters, gradient descent gradually adjusts the network weights to reduce error. An optimizer (such as AdamW or SGD) is used to control the pace of weight updates, thereby accelerating convergence and preventing overfitting.

[0076] 4. Training data:

[0077] The network is trained on a detection dataset containing labeled support images, which indicate the support category and support structure bounding box information. To improve the network's generalization ability, data augmentation techniques such as random cropping, rotation, horizontal flipping, and color jitter are used during training to ensure that the model can adapt to various different scenarios.

[0078] Embodiment 2 of this invention combines multi-scale feature extraction and dynamic semantic reasoning to automatically capture the complex relationship between the foundation pit and the support structure. The feature extraction process is further optimized through a semantic attention mechanism, significantly improving the recognition accuracy of the support structure, especially in complex scenarios. Existing technologies often rely on periodic detection or manual intervention, resulting in delays and insufficient real-time performance, making it impossible to provide immediate feedback and adjustments during construction. This invention, based on the real-time reasoning capabilities of a deep learning model, enables real-time monitoring and automatic judgment of the foundation pit support structure, issuing timely warnings before problems occur, preventing construction risks, and improving overall construction safety and efficiency. Existing recognition methods are prone to false positives or false negatives when facing complex construction environments, varying lighting conditions, and obstructions. This invention, through the combination of PVT v2 and FPN networks, can extract multi-scale visual features and fuse them through a cross-scale semantic perception module. Even in complex environments (such as changes in lighting or partial obstruction by the support structure), it can accurately locate the support structure and determine whether its quantity and location meet construction requirements, exhibiting better environmental adaptability. This invention utilizes an automated machine learning model for the detection, location, quantity determination, and position assessment of support structures, significantly reducing the need for manual intervention and improving operational convenience. Simultaneously, the intelligent system can autonomously analyze whether the support structures meet construction requirements, reducing human error and judgment bias, and enhancing the level of automation in construction. A dynamic semantic relationship graph between support structures is constructed through a Dynamic Relationship Graph Reasoning (DRGR) module, and inference is performed using a Graph Convolutional Network (GCN), thereby more accurately analyzing the interaction between the foundation pit and the support structures. This dynamic reasoning process avoids the limitations of traditional methods that rely on fixed templates or features, and can flexibly respond to changes in different construction sites. By employing an attention mechanism, key areas of the foundation pit and support structure are weighted, effectively reducing the probability of false positives and false negatives. Through cross-scale semantic perception and visual feature fusion, the system can more comprehensively and accurately understand the overall layout and details of the foundation pit support structure. Utilizing the Pyramid Transformer structure improved by PVT v2, the computational complexity is significantly reduced while maintaining high accuracy, thus optimizing computational performance. Multi-scale feature fusion through Feature Pyramid Network (FPN) enables the system to efficiently process large-scale image data, making it suitable for real-time monitoring applications.

[0079] like Figure 4The diagram illustrates the semantic relationships between the foundation pit and the supporting structure. This semantic relationship graph helps the network understand the contextual information between the foundation pit and the supporting structure. For example, there is often a specific spatial relationship between the foundation pit and the supporting structure, with the supporting structure typically located at certain specific locations within the foundation pit. By incorporating this prior semantic knowledge, the network can more accurately predict the location and shape of the supporting structure. Semantic relationships also help the detector more robustly identify objects in complex environments. For instance, in noisy or complex images, supporting structures inferred from semantic relationships may be more easily and accurately detected.

[0080] The relative positions of the excavation pit and the supporting structure exhibit certain regularities. By incorporating this spatial information, the network can reduce false detections. For example, if a certain area is identified as an excavation pit, the probability of detecting a supporting structure in a specific area around the pit increases. Conversely, if a supporting structure is detected in an area far from the excavation pit, the network can use relational reasoning to determine that this is a false detection. Positional relationships can also reduce the number of candidate regions, focusing on areas where supporting structures are likely to be present, thereby improving detection accuracy.

[0081] Introducing semantic relationships helps the network better handle the detection tasks of foundation pits and support structures in different scenarios. For example, in new scenarios, even if the specific shape or size of the foundation pits and support structures changes, the network can still accurately detect these structures because it has learned the semantic and spatial relationships between them. The improved generalization ability is also reflected in the network's adaptation to different scales and complex environments. Semantic relationships enable the network to not only rely on visual features but also to reason using prior knowledge, thereby enhancing its robustness in different environments.

[0082] Furthermore, as Figures 2 to 3 In a specific implementation of the method, this invention provides a machine learning-based foundation pit support monitoring device, such as... Figure 4 As shown, the device includes:

[0083] Module 510 is used to acquire foundation pit image data;

[0084] The multi-scale feature extraction module 520 is used to extract features of different scales from image data through the visual pyramid PVT network and fuse information of different scales through the feature pyramid network FPN to generate multi-scale feature vectors.

[0085] The cross-scale semantic perception module 530 is used to flatten features at multiple scales and calculate global average pooling to generate global semantic vectors for each scale. Through channel concatenation and 1*1 convolution operations, it generates cross-scale semantic feature vectors.

[0086] The dynamic relation graph reasoning module 540 is used to construct a dynamic adjacency matrix based on multi-scale feature vectors and cross-scale semantic feature vectors, and to reason about the semantic features of the relationship between the foundation pit and the support structure through the graph convolutional network GCN.

[0087] The semantic attention fusion module 550 is used to calculate the attention weight between the foundation pit and the support structure through the attention mechanism, and to fuse the semantic features of the relationship between the foundation pit and the support structure with multi-scale features to generate enhanced features.

[0088] The foundation pit support detection module 560 is used to determine whether the quantity and location of the support structure meet the construction requirements based on the enhanced features.

[0089] This invention provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of a machine learning-based foundation pit support monitoring method, including:

[0090] Acquire foundation pit image data;

[0091] The Visual Pyramid (PVT) network extracts features at different scales from image data, and the Feature Pyramid (FPN) network fuses information at different scales to generate multi-scale feature vectors.

[0092] Multi-scale features are flattened and global average pooling is calculated to generate global semantic vectors for each scale. Through channel concatenation and 1*1 convolution operations, cross-scale semantic feature vectors are generated.

[0093] A dynamic adjacency matrix is ​​constructed based on multi-scale feature vectors and cross-scale semantic feature vectors, and the semantic features of the relationship between the foundation pit and the support structure are inferred through graph convolutional network (GCN).

[0094] The attention weights between the foundation pit and the support structure are calculated through an attention mechanism, and the semantic features of the relationship between the foundation pit and the support structure are fused with multi-scale features to generate enhanced features.

[0095] Based on the enhanced features, it is determined whether the quantity and location of the support structure meet the construction requirements.

[0096] It should be noted that the above embodiments are only used as examples of foundation pit support monitoring and identification to illustrate the principles and implementation steps of the present invention. They do not specifically limit the actual application scenarios, such as high formwork for buildings, ground piles, deep foundation pits, etc. For the functions or steps that can be implemented by computer-readable storage media or computer devices, please refer to the foregoing method embodiments. To avoid repetition, they will not be described one by one here.

[0097] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0098] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0099] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A machine learning-based method for monitoring foundation pit support, characterized in that, include: Real-time acquisition of video stream data of the foundation pit, and extraction of static images from the video stream; The static image is subjected to noise reduction and contrast enhancement operations. The Visual Pyramid (PVT) network is used to extract multi-scale feature maps through multiple stages, where the feature maps output by each stage have different resolutions and number of channels. The Feature Pyramid Network (FPN) then performs up-and-down path fusion on these feature maps to generate multi-scale feature maps with different scale information. Global semantic information is extracted from multi-scale feature maps and flattened into one-dimensional vectors. Global average pooling is performed on each flattened feature map to generate a global semantic vector for each scale. Global semantic vectors of all scales are concatenated along the channel dimension to generate a high-dimensional cross-scale semantic vector. The concatenated semantic vector is then dimensionality-reduced by 1*1 convolution to generate a cross-scale semantic feature vector. Based on the multi-scale feature vectors, transformation features and category activation maps are generated to generate node representations; through convolution operations, multi-scale features and cross-scale semantic features are integrated to generate an adjacency matrix representing the dynamic relationship between objects; through a graph convolutional network (GCN), edges are constructed based on the information propagated from the adjacency matrix and the node representations to generate semantic features representing the relationship between the foundation pit and the support structure. The multi-scale features are subjected to dimensionality reduction; the semantic features of the relationship between the foundation pit and the support structure and the multi-scale features are reorganized to fuse the semantic features of the relationship between the foundation pit and the support structure and the multi-scale features; the similarity between the semantic features of the relationship between the foundation pit and the support structure and the multi-scale features is calculated by Hadamard product to obtain attention weights; the attention weights are fused with the multi-scale features to generate enhanced features; Based on the enhanced features, it is determined whether the quantity and location of the support structure meet the construction requirements. Specifically, this includes identifying the support structure using the enhanced feature map, determining its location in the foundation pit image, and adding a bounding box to the support structure; counting the bounding boxes of the support structure to determine whether the quantity of the support structure meets the construction requirements; evaluating the bounding boxes of the support structure to determine whether they meet the location setting standards of the support structure in the construction requirements; outputting the result of whether the quantity and location of the support structure meet the construction requirements, and providing relevant detection information.

2. A machine learning-based foundation pit support monitoring device, characterized in that, include: The acquisition module is used to acquire video stream data of the foundation pit in real time and extract static images from the video stream; The static image is subjected to noise reduction and contrast enhancement operations. The multi-scale feature extraction module is used to extract multi-scale feature maps through multiple stages using the Visual Pyramid (PVT) network. Each stage outputs a feature map with different resolutions and number of channels. The Feature Pyramid Network (FPN) performs up-and-down path fusion on these feature maps to generate multi-scale feature maps with different scale information. The cross-scale semantic perception module is used to extract global semantic information from multi-scale feature maps, flatten the multi-scale feature maps into one-dimensional vectors; perform global average pooling on each flattened feature map to generate a global semantic vector for each scale; concatenate the global semantic vectors of all scales along the channel dimension to generate a high-dimensional cross-scale semantic vector; and reduce the dimensionality of the concatenated semantic vector through 1*1 convolution to generate a cross-scale semantic feature vector. The dynamic relationship graph reasoning module is used to generate transformation features and category activation maps to generate node representations based on the multi-scale feature vectors; through convolution operations, it integrates multi-scale features and cross-scale semantic features to generate an adjacency matrix representing the dynamic relationship between objects; through the graph convolutional network GCN, it constructs edges based on the information propagated by the adjacency matrix and the node representations to generate semantic features of the relationship between the foundation pit and the support structure. A semantic attention fusion module is used to perform dimensionality reduction on the multi-scale features; to perform a reshaping operation on the semantic features of the relationship between the foundation pit and the support structure and the multi-scale features, so as to fuse the semantic features of the relationship between the foundation pit and the support structure and the multi-scale features; to calculate the similarity between the semantic features of the relationship between the foundation pit and the support structure and the multi-scale features through the Hadamard product, and obtain attention weights; and to fuse the attention weights with the multi-scale features to generate enhanced features. The foundation pit support detection module is used to determine whether the quantity and location of the support structure meet the construction requirements based on the enhanced features.

3. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the machine learning-based foundation pit support monitoring method as described in claim 1.

4. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the machine learning-based foundation pit support monitoring method as described in claim 1.

Citation Information

Patent Citations

  • Lightweight medical image segmentation network, method and equipment based on multi-path pyramid

    CN117274607A

  • Large-area deep foundation pit holographic monitoring method based on image recognition technology

    CN118278251A