Building structure safety intelligent monitoring method, device and equipment and storage medium
By acquiring building data through multimodal sensors, performing multi-scale attention fusion and semantic segmentation, and combining tilt angle calculation, the problem of insufficient multi-parameter collaborative analysis in existing technologies is solved, and high-precision building structure safety monitoring is achieved.
Patent Information
- Application Number
- CN202610003603.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-05
- Publication Date
- 2026-02-03
AI Technical Summary
Existing methods for monitoring the safety of building structures suffer from a high rate of false alarms and a lack of ability to conduct collaborative analysis of multiple parameters such as tilt, vibration, and cracks, resulting in a high false alarm rate and an inability to effectively utilize the city's existing camera network.
Multimodal data of buildings are acquired using multimodal sensors. Vibration and crack features are identified through multi-scale attention fusion and semantic segmentation. Multi-parameter coupling calculation is performed by combining tilt angle to construct a safety monitoring model.
It improves the ability to identify and locate features of distant, small target buildings, significantly enhances monitoring reliability and environmental adaptability, and achieves high-precision multi-dimensional parameter monitoring.
Smart Images

Figure CN121459293A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of building safety, and particularly relates to a building structure safety intelligent monitoring method, device and equipment and a storage medium. BACKGROUND
[0002] At present, the health and safety monitoring of building structures mainly relies on three types of technologies: sensor-based methods (such as accelerometers, laser interferometers), machine vision methods (such as digital image correlation DIC, phase video motion magnification PVMM) and manual inspection methods (such as total station).
[0003] In the prior art, the DIC method needs to spray artificial speckles, and the on-site preparation is complicated; the PVMM method is easily disturbed by noise in actual application; the laser sensor is high in cost and difficult to scale; and the total station is low in efficiency. In addition, the existing methods are mostly single-parameter monitoring, and lack of multi-parameter collaborative analysis of inclination, vibration and cracks, resulting in high false alarm rate and inability to effectively utilize the existing camera network in the city. In summary, in the existing scheme, the multi-dimensional apparent characteristics such as inclination, settlement, vibration and crack are mostly analyzed comprehensively and collaboratively, resulting in high risk misjudgment rate. SUMMARY
[0004] The main purpose of the present application is to provide a building structure safety intelligent monitoring method, device, equipment and storage medium, i.e. program product, which aims to solve the technical problem of high risk misjudgment rate of the existing building structure safety intelligent monitoring method.
[0005] To achieve the above-mentioned purpose, the present application provides a building structure safety intelligent monitoring method, which comprises: obtaining multi-modal data of a building based on a multi-modal sensor; performing multi-scale attention fusion on the multi-modal data, and performing vibration feature recognition based on the fused features to obtain vibration features of the building; performing semantic segmentation on the multi-modal data to obtain crack features; determining an inclination angle of the building based on the multi-modal data; performing multi-parameter coupling calculation according to the vibration features, the crack features and the inclination angle to obtain a safety monitoring result of the building.
[0006] In an embodiment, the step of performing multi-scale attention fusion on the multi-modal data and performing vibration feature recognition based on the fused features to obtain the vibration features of the building comprises: performing multi-scale attention fusion on the multi-modal data to obtain a fused feature image; Sub-pixel positioning is performed based on the fused feature image to obtain a pixel displacement amount; The pixel displacement amount is converted into a vibration displacement amount in a world coordinate system based on a camera calibration model; Vibration feature recognition is performed according to the vibration displacement amount to obtain the vibration feature of the building.
[0007] In an embodiment, the step of performing multi-scale attention fusion on the multi-modal data to obtain a fused feature image comprises: The multi-modal data is processed through a channel attention path to obtain channel attention weights; The multi-modal data is processed through a spatial attention path to obtain spatial attention weights; The channel attention weights and the spatial attention weights are used for weighted fusion to obtain a fused feature image.
[0008] In an embodiment, the step of determining the inclination angle of the building based on the multi-modal data comprises: A pixel displacement correlation matrix and a thermal expansion coefficient of a main material of the building are obtained; The inclination angle of the building is determined according to the pixel displacement correlation matrix and the thermal expansion coefficient of the main material.
[0009] In an embodiment, the step of performing multi-parameter coupling calculation according to the vibration feature, the crack feature, and the inclination angle to obtain a safety monitoring result of the building comprises: Temperature data is determined in the multi-modal data; A crack evolution model is constructed according to the vibration feature, the inclination angle, and the temperature data; Crack evolution is performed according to the crack feature and the crack evolution model to obtain a crack change rate; The safety monitoring result of the building is determined according to the crack change rate.
[0010] In an embodiment, the method further comprises: A safety monitoring model is optimized through a hardware-aware neural architecture search compression technique, the safety monitoring model being used to obtain a safety monitoring result of the building according to the multi-modal data; The step of optimizing the safety monitoring model through the hardware-aware neural architecture search compression technique comprises: A multi-dimensional search space is constructed; Difference architecture search and hypernetwork training are performed based on the multi-dimensional search space to obtain an optimized hypernetwork; The optimized hypernetwork is processed through knowledge distillation to obtain a safety monitoring model.
[0011] In an embodiment, before the step of performing vibration feature recognition based on the fused features, the method comprises: constructing a fuzzy kernel model based on building vibration characteristics; based on the fuzzy kernel model and through a generalized intersection-over-union loss function, performing adversarial training to obtain a trained image feature extraction model; in the adversarial training process, self-adaptive weighting is performed on difficult samples.
[0012] In addition, to achieve the above-mentioned purpose, the present application also proposes a building structure safety intelligent monitoring device, the building structure safety intelligent monitoring device comprises: a data acquisition module configured to acquire multi-modal data of a building based on multi-modal sensors; a feature recognition module configured to perform multi-scale attention fusion on the multi-modal data, and perform vibration feature recognition based on the fused features to obtain vibration features of the building; a semantic segmentation module configured to perform semantic segmentation on the multi-modal data to obtain crack features; a tilt angle processing module configured to determine a tilt angle of the building based on the multi-modal data; an intelligent monitoring module configured to perform multi-parameter coupling calculation according to the vibration features, the crack features, and the tilt angle to obtain a safety monitoring result of the building.
[0013] In addition, to achieve the above-mentioned purpose, the present application also proposes a building structure safety intelligent monitoring device, the device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, the computer program being configured to implement the steps of the building structure safety intelligent monitoring method as described above.
[0014] In addition, to achieve the above-mentioned purpose, the present application also proposes a storage medium, which is a computer-readable storage medium, and the storage medium stores a computer program, and the computer program is executed by a processor to implement the steps of the building structure safety intelligent monitoring method as described above.
[0015] In addition, to achieve the above-mentioned purpose, the present application also provides a computer program product, which comprises a computer program, and the computer program is executed by a processor to implement the steps of the building structure safety intelligent monitoring method as described above.
[0016] The one or more technical solutions proposed in the present application have at least the following technical effects: The application obtains multi-modal data of a building based on a multi-modal sensor; multi-scale attention fusion is performed on the multi-modal data, and vibration feature recognition is performed based on the fused features to obtain vibration features of the building; semantic segmentation is performed on the multi-modal data to obtain crack features; the tilt angle of the building is determined based on the multi-modal data; multi-parameter coupling calculation is performed according to the vibration features, the crack features and the tilt angle to obtain the safety monitoring result of the building. Since multi-scale attention fusion is performed on the multi-modal data, the feature recognition and positioning capability for long-distance and small target buildings is improved; multi-parameter coupling calculation is performed on multiple parameters such as crack features, vibration features and tilt angles, which significantly improves the monitoring reliability and environmental adaptability. BRIEF DESCRIPTION OF DRAWINGS
[0017] The accompanying drawings, which are incorporated herein and form part of the specification, illustrate embodiments consistent with the present application and, together with the description, further serve to explain the principles of the application.
[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the accompanying drawings required to be used in the embodiments or prior art description will be briefly introduced. Obviously, for those skilled in the art, other drawings can also be obtained based on these drawings without creative labor.
[0019] Figure 1 An overall technical route schematic diagram in one implementation manner of the building structure safety intelligent monitoring method of the present application; Figure 2 A flowchart schematic diagram provided by the building structure safety intelligent monitoring method embodiment one of the present application; Figure 3 A flowchart schematic diagram provided by the building structure safety intelligent monitoring method embodiment two of the present application; Figure 4 A flowchart schematic diagram provided by the building structure safety intelligent monitoring method embodiment three of the present application; Figure 5 A model structure schematic diagram in one implementation manner of the building structure safety intelligent monitoring method of the present application; Figure 6 A module structure schematic diagram of the building structure safety intelligent monitoring device of the present application embodiment; Figure 7 A device structure schematic diagram of the hardware running environment involved in the building structure safety intelligent monitoring method in the present application embodiment.
[0020] The purpose implementation, functional features and advantages of the present application will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION
[0021] It should be understood that the specific embodiments described herein are merely intended to explain the technical solutions of the present application, and are not intended to limit the present application.
[0022] In order to better understand the technical solutions of the present application, the following will be described in detail in combination with the drawings of the specification and specific embodiments.
[0023] In some embodiments of the present application, the health monitoring of the building structure can be performed by the Digital Image Correlation (DIC) method, which requires spraying artificial speckle on the surface of the measured object, and calculating the vibration displacement through a correlation algorithm. Although the accuracy of about 0.2mm can be achieved, the on-site preparation work is complicated, and it is difficult to apply it on a large scale in existing buildings.
[0024] In some embodiments of the present application, the micro-motion can be extracted by phase amplification. Although the correlation can reach 0.996 in a laboratory environment, noise and image defects are easily introduced in actual application, resulting in a decrease in motion estimation accuracy.
[0025] In some embodiments of the present application, the health monitoring can be realized by laser displacement sensors, total station instruments and other sensors. Although the laser displacement sensor has high accuracy, it is expensive; the total station instrument has the disadvantage of low efficiency, and is a single-point contact measurement with limited coverage.
[0026] In some embodiments of the present application, a three-layer intelligent architecture of perception-processing-decision can be used to solve the monitoring problems in complex environments through spatio-temporal fusion of multi-source data. Specifically, as shown in Figure 1 , Figure 1 is a schematic diagram of the overall technical route in one implementation of the building structure safety intelligent monitoring method of the present application.
[0027] Among them, the perception layer can comprehensively utilize the existing optical sensors, infrared thermal imagers in the city, and can construct a multi-modal perception network through the laser ranging point cloud data of the laser radar. The optical sensor is used to obtain high-resolution images, identify and fuse feature images and cracks; the infrared thermal imager is used to detect external wall hollowing, internal defects; the laser radar is used for three-dimensional point cloud modeling, and assists in tilt and settlement analysis.
[0028] The processing layer can include an edge computing node and a cloud analysis center. The edge node can deploy a lightweight model to implement multi-scale attention fusion (MAF) feature enhancement, motion blur adversarial training (MBAT), and sub-pixel localization, etc., and is responsible for real-time data processing and preliminary feature extraction; the cloud center performs large-scale data fusion, model training and deep analysis.
[0029] It should be noted that the lightweight model described above can be a model optimized by neural architecture search for compression (NASC) technology.
[0030] It should be noted that the decision layer performs displacement calculation, tilt analysis, crack evaluation, risk fusion, etc. on the data uploaded by the processing layer based on a multi-parameter coupled solution model, and outputs the building safety risk level, early warning information and maintenance strategy.
[0031] The main solution of the embodiment of the application is: obtaining multi-modal data of a building based on a multi-modal sensor; performing multi-scale attention fusion and vibration feature recognition on the multi-modal data to obtain vibration features of the building; performing semantic segmentation on the multi-modal data to obtain crack features; determining an inclination angle of the building based on the multi-modal data; and performing multi-parameter coupled solution according to the vibration features, the crack features and the inclination angle to obtain a safety monitoring result of the building. By performing multi-scale attention fusion on the multi-modal data, the recognition difficulty of fused feature images in long-distance monitoring is solved; by coupling and solving multiple parameters such as vibration features, crack features and inclination angles of the building, the monitoring reliability and environmental adaptability are significantly improved.
[0032] It should be noted that the execution subject of the embodiment can be a computing service device with data processing, network communication and program running functions, such as a computer, a server, etc., or an electronic device, a virtual device, etc. that can realize the above functions. The embodiment and the following embodiments will be described below with the building structure safety intelligent monitoring device (referred to as monitoring device) as an example.
[0033] Based on this, the embodiment of the application provides a building structure safety intelligent monitoring method, which is described below with reference to Figure 2 , Figure 2 The flowchart provided by the first embodiment of the building structure safety intelligent monitoring method of the application is shown in the figure. Figure 2The data processing flow shown can include: data acquisition and preprocessing (multimodal sensor synchronous acquisition, denoising, distortion correction, space-time alignment), feature extraction and identification (MAF positioning feature marking, semantic segmentation identifying cracks, point cloud calculating inclination angle), sub-pixel positioning and vibration calculation (gradient vector optimization sub-pixel positioning, camera calibration converting physical displacement), multi-parameter fusion and decision (coupling vibration, inclination, crack, temperature data, output risk index).
[0034] In this embodiment, the building structure safety intelligent monitoring method includes steps S10-S40: Step S10, acquiring multimodal data of the building based on a multimodal sensor; Step S20, performing multiscale attention fusion on the multimodal data, and performing vibration feature identification based on the fused features to obtain vibration features of the building.
[0035] It can be understood that through the perception layer of the embodiment of the application, multimodal data can be obtained based on existing optical sensors, infrared thermographs, laser radar point cloud data, etc. in the city.
[0036] It should be noted that the multiscale attention fusion algorithm can be used to realize multiscale attention fusion of the multimodal data. In the embodiment of the application, the multiscale attention fusion algorithm is an algorithm that can solve the problem of fusion feature image identification in building long-distance monitoring. In the building safety monitoring scene, the fusion feature image accounts for only 0.1%-0.3% of the image area at 30 meters away, and the recall rate of the traditional detection method is less than 85% in such small target identification. The MAF algorithm simulates the selective attention mechanism of the human visual system, and enhances the key features through a double-path weight distribution strategy. The channel attention path focuses on the channel dimension of the feature map, and filters the high-frequency texture features sensitive to vibration. The spatial attention path enhances the key spatial positions such as corner points and edges according to the geometric structure characteristics of the building. The algorithm is expected to improve the recognition rate of the fusion feature image at a distance of 200 meters to more than 97%, and reduce the false detection rate caused by strong light interference by 40%, thereby providing high-precision input for subsequent sub-pixel positioning.
[0037] It should be explained that through the multiscale attention fusion algorithm of the embodiment of the application, multiscale attention fusion can be performed on the collected multimodal data. Through vibration feature identification based on the fused features, vibration features for representing the vibration condition of the building are obtained. Through the vibration features, the vibration stress, vibration amplitude, vibration displacement and other parameters of the building can be determined.
[0038] In its specific implementation, the monitoring device in this application acquires multimodal data of buildings based on multimodal sensors; it performs multi-scale attention fusion on the multimodal data and identifies vibration features from the fused features to obtain the vibration characteristics of the buildings. Because it acquires multimodal data of buildings by comprehensively utilizing existing multimodal sensors in the city, it solves the problems of isolated and difficult-to-coordinate data from different sources through spatiotemporal alignment and feature-level fusion technology. This achieves high-precision synchronous perception and comprehensive analysis of multi-dimensional parameters such as building micro-vibration, structural tilt, cracks, and external wall hollowing, significantly improving the comprehensive understanding of building safety status and monitoring reliability in complex environments, providing a complete underlying technical framework for intelligent operation and maintenance of city-level building clusters.
[0039] Step S30: Perform semantic segmentation on the multimodal data to obtain crack features; Step S40: Determine the tilt angle of the building based on the multimodal data; Step S50: Perform multi-parameter coupled solution based on the vibration characteristics, crack characteristics, and tilt angle to obtain the safety monitoring results of the building.
[0040] It is understood that semantic segmentation is a pixel-level classification task in computer vision, which assigns a category label to each pixel of an image in multimodal data. The semantic segmentation method used in the embodiments of this application may be semantic segmentation based on convolutional neural networks, semantic segmentation based on the Transformer architecture, or other semantic segmentation methods. The embodiments of this application do not limit this.
[0041] It should be understood that the aforementioned tilt angle is the offset angle of the entire or partial building structure relative to the calibrated reference plane. The aforementioned reference plane can be the ground or other reference planes, and the embodiments of this application do not limit this.
[0042] In some embodiments of this application, the tilt angle can be determined based on point cloud data in multimodal data. Specifically, a three-dimensional model can be obtained by performing three-dimensional modeling based on point cloud data collected by LiDAR; the tilt angle of the building can then be obtained by performing audit analysis based on the three-dimensional building model.
[0043] In its specific implementation, the monitoring device of this application performs semantic segmentation on multimodal data to obtain crack features; determines the building's tilt angle based on the multimodal data; and performs multi-parameter coupled calculation based on vibration features, crack features, and tilt angle to obtain the building's safety monitoring results. Because it achieves high-precision, low-cost, and intelligent monitoring of buildings by analyzing multiple parameters such as building micro-vibration, structural tilt, and crack features, it significantly improves monitoring reliability and environmental adaptability.
[0044] This application embodiment acquires multimodal data of a building based on multimodal sensors; performs multi-scale attention fusion on the multimodal data, and identifies vibration features based on the fused features to obtain the building's vibration characteristics; performs semantic segmentation on the multimodal data to obtain crack features; determines the building's tilt angle based on the multimodal data; and performs multi-parameter coupled calculation based on the vibration features, crack features, and tilt angle to obtain the building's safety monitoring results. Because multi-scale attention fusion is performed on the multimodal data, the ability to identify and locate features of distant, small-target buildings is improved; and the multi-parameter coupled calculation using multiple parameters such as crack features, vibration features, and tilt angle significantly improves monitoring reliability and environmental adaptability.
[0045] Based on the first embodiment of this application, in the second embodiment of this application, the content that is the same as or similar to that in the first embodiment described above can be referred to the above description, and will not be repeated hereafter. Based on this, please refer to... Figure 3 , Figure 3 This is a flowchart illustrating Embodiment 2 of the intelligent monitoring method for building structural safety in this application.
[0046] like Figure 2 As shown in the embodiment of this application, the step of performing multi-scale attention fusion on the multimodal data and identifying vibration features based on the fused features to obtain the vibration features of the building includes: Step S21: Perform multi-scale attention fusion on the multimodal data to obtain a fused feature image; Step S22: Perform sub-pixel localization based on the fused feature image to obtain the pixel displacement.
[0047] The multi-scale attention fusion in this embodiment can be implemented based on the MAF algorithm, a core algorithm that enhances the ability to identify and locate features of small, long-distance targets through a dual-path weight allocation mechanism. The dual paths are the channel attention path and the spatial attention path. Specifically, the step of performing multi-scale attention fusion on the multimodal data to obtain a fused feature image includes: performing feature processing on the multimodal data through the channel attention path to obtain channel attention weights; performing feature processing on the multimodal data through the spatial attention path to obtain spatial attention weights; and performing weighted fusion based on the channel attention weights and the spatial attention weights to obtain the fused feature image.
[0048] It should be noted that the channel attention path in this embodiment is used to determine the importance of each channel in the feature map of multimodal data, and to strengthen channels that are useful for identifying vibration-sensitive features (such as high-frequency textures) while suppressing useless channels. These channels can be understood as different feature filters, some responsible for identifying textures and others for identifying edges. Specifically, the feature processing of the channel attention path in this embodiment can be exemplified as follows: ; in, This represents a feature map obtained by performing average pooling on multimodal data, which captures the overall distribution of features and global contextual information; This represents the feature map obtained through the MaxPooling operation, which highlights the most significant and unique response points among the features. This represents the ReLU activation function, which has the following functional form: This is used to introduce nonlinear transformations to enhance the expressive power of the model. The Sigmoid activation function has the following form: This is used to generate a weighted graph in a spatial dimension, where the closer the value is to 1, the more important the spatial location is. , Represents two learnable weight matrices used in a multilayer perceptron (MLP) to perform linear transformation and dimensionality reduction on pooled features in order to learn the complex relationships between channels. This represents the channel attention weights obtained based on feature F.
[0049] It should be noted that the spatial attention path in this embodiment is used to determine the importance of each spatial location of the feature map (i.e., each region of the image) and to strengthen those locations that contain key structural information (such as building corners and edges) while suppressing irrelevant regions such as the background.
[0050] In some embodiments of this application, the spatial attention path can use a 7×7 large convolution kernel to capture the macroscopic spatial relationships of building edges and compensate for long-distance perspective distortion. Specifically, it can be as follows: ; in, This indicates that a 7x7 kernel is used for the convolution operation. Choosing a larger kernel is to capture a wider range of macroscopic spatial context information and structural relationships (such as the long edges of buildings) in the feature map, which is crucial for perceiving geometric features such as building edges and corners. This means concatenating the two spatial feature maps obtained from average pooling and max pooling along the channel dimension to form a richer feature map containing both types of information. This represents the spatial attention weights obtained based on feature F.
[0051] It should be noted that when obtaining the channel attention weights and spatial attention weights, these two attention features can be weighted and fused to obtain a fused feature image. Specifically, the attention weights calculated from the two paths can be applied to the original features while preserving the original information to prevent gradient vanishing, providing high-precision input for sub-pixel localization, as shown in the following formula: ; in, The feature map representing the input is the common input of the channel attention path and the spatial attention path; Indicated based on feature map The obtained channel attention weights, Indicated based on feature map The obtained spatial attention weights; This represents the final output feature map after optimization by the multi-scale attention fusion mechanism, i.e., the fused feature image. This represents a residual connection. By adding the weighted features to the original input features F, it ensures that the original feature information is not lost even under the attention mechanism, which is beneficial for gradient propagation and stable model training. This indicates element-wise multiplication, which multiplies the generated channel attention weight map with the original input feature map to weight the features of different channels.
[0052] It should be noted that the aforementioned feature maps can be obtained based on image sequences in multimodal data and using an image feature extraction model. In this embodiment, to achieve feature extraction, a motion fuzzy adversarial training method is also designed to realize dynamic monitoring of buildings under strong wind loads.
[0053] It is understandable that when wind speeds exceed 8 m / s, building vibrations can cause directional blurring in image sequences, reducing the cross-union ratio (CUI) of traditional detection methods to below 0.7. This application's embodiment innovatively establishes a motion blur model based on physical laws through motion blur adversarial training. By using fuzzy kernel parameterization and adversarial training strategies, the system can maintain a positioning accuracy above 0.85 even in typhoon scenarios. Specifically, before the step of vibration feature recognition based on fused features, the process includes: constructing a fuzzy kernel model based on building vibration characteristics; conducting adversarial training based on the fuzzy kernel model using a generalized CUI loss function to obtain a trained image feature extraction model; and adaptively weighting difficult samples during adversarial training.
[0054] In this embodiment, the motion blur effect caused by building vibration can be accurately simulated using a mathematical model. The specific blur kernel model can be as follows: ; in, This represents the generated motion blur kernel, which defines how each pixel in the image is blended with its neighboring pixels due to motion. This indicates the direction angle of the main vibration of a building caused by wind-induced vibration (unit: degrees). Under severe weather conditions such as typhoons, this angle is mainly distributed in the range of 30° to 60°, and it determines the direction of motion ambiguity. This represents the summation index, with values ranging from 0 to... , indicating the fuzzy length above generated Discrete points are used to simulate continuous motion blur. Blur length. It reflects the degree of motion blur in the image and is positively correlated with wind speed, according to an empirical formula. , Wind speed (unit: m / s). The Dirac Delta Function is used to represent each discrete fuzzy point in a specific direction.
[0055] It should be noted that the fuzzy kernel mathematical model can simulate the image blurring effect caused by building vibration under specific wind conditions, thereby generating a batch of realistic training images with motion blur. These training images can then be used to train the image feature extraction model. During the training process of the image feature extraction model, this embodiment uses Generalized Intersection over Union (GIoU) for loss optimization. Even if the predicted bounding box and the ground truth bounding box do not overlap in the blurred image, it can still provide an effective gradient for model optimization, solving the problem of traditional IoU loss failing in this scenario. Specifically, it can be shown below: ; in, The smallest enclosing box (GIoU) represents the region containing both the predicted and ground truth bounding boxes. GIoU loss introduces... This effectively solves the problem that the traditional IoU loss gradient is zero and cannot be optimized when two boxes do not overlap. This represents the bounding box predicted by the model. Represents the actual bounding box. Let C represent the area of the smallest closure region. This represents the area of the minimum closure region C minus the predicted bounding box. With real frame The area of union, or GIoU, which is the area of the "blank" region between the two bounding boxes, guides model optimization by penalizing this blank region. IOU, or Intersection over Union Ratio, is used to represent the ratio of the intersection area to the union area between the predicted bounding box and the ground truth bounding box.
[0056] Understandably, the generalized intersection-union ratio (OUNR) is a function used to evaluate the similarity between two predicted bounding boxes and the ground truth bounding boxes.
[0057] In some embodiments of this application, adaptive sample weighting can also be applied to the training process. Specifically, higher weights are assigned to difficult samples (such as blurred images under strong winds) to enhance the model's learning ability under extreme conditions. This technique enables the system to keep the localization error of the fused feature image within 3 pixels even under level 12 wind load conditions. Specifically, it can be as follows: .
[0058] It should be noted that the aforementioned pixel displacement refers to the displacement of the building within the captured image / video. To determine the pixel displacement, this embodiment of the application may affix feature markers to the building. These feature markers can be high-contrast geometric pattern targets, serving as reference points for visual recognition and sub-pixel localization. By searching for and comparing the positions of these feature markers in the fused feature map, the pixel displacement of the building can be determined.
[0059] Step S23: Based on the camera calibration model, convert the pixel displacement into vibration displacement in the world coordinate system; Step S24: Based on the vibration displacement, vibration characteristics are identified to obtain the vibration characteristics of the building.
[0060] It should be noted that the above-mentioned camera calibration model can be a preset calibration model constructed based on the camera's intrinsic and extrinsic parameters to convert between pixel size and actual size. The specific conversion ratio can be set according to the needs of actual application, and this application embodiment does not limit it.
[0061] It is understandable that, through camera calibration models, the pixel displacement of a building image can be converted into the vibration displacement of the building in the real world coordinate system using a vibration displacement equation. By performing vibration feature identification based on this vibration displacement, the vibration characteristics of the building can be obtained. Specifically, the vibration displacement equation in this embodiment can be as follows: ; in, This represents the micro-vibration displacement of the building calculated at time t (unit: mm), i.e., the vibration displacement amount. N This represents the total number of feature tags that were successfully identified and tracked from the fused feature map. Indicates the first i Each feature marker is a unique vector (in pixels) in the pixel coordinate system of the image. It can be a two-dimensional vector (Δx, Δy) that represents how many pixels the point has moved in the x and y directions from the previous frame to the current frame. Let L2 norm be the norm used to calculate the magnitude of this three-dimensional displacement vector, i.e., the total displacement. The formula is: The L2 norm ignores the direction of displacement and only focuses on the amplitude of vibration. It can effectively suppress the interference of outliers at individual points on the overall result, thus improving robustness. H represents the homography transformation matrix, used to transform pixel coordinates to the world coordinate system.
[0062] In some embodiments of this application, in order to determine the tilt angle of a building, the step of determining the tilt angle of the building based on the multimodal data includes: obtaining a pixel displacement correlation matrix and the thermal expansion coefficient of the main material of the building; and determining the tilt angle of the building based on the pixel displacement correlation matrix and the thermal expansion coefficient of the main material.
[0063] It should be noted that the aforementioned coefficient of thermal expansion of the main building material refers to the coefficient of thermal expansion of the building's main materials. This coefficient can be used to quantify the expansion or contraction effect of the main materials due to temperature changes. For example, the typical coefficient of thermal expansion of concrete is approximately 12 × 10⁻⁶. -6 / ℃.
[0064] It should also be noted that the aforementioned pixel displacement correlation matrix can be a matrix used to transform image pixel coordinates to real-world coordinates. In this embodiment, the aforementioned pixel displacement correlation matrix can specifically be a camera calibration matrix or a homography transformation matrix. Specifically, in this embodiment, the tilt angle of a building can be determined based on the parameters in the camera calibration matrix or homography transformation matrix and the thermal expansion coefficient of the main material, as shown below: ; in, This indicates the final, temperature-compensated building tilt angle. This represents the pixel displacement of the feature marker in the vertical y-direction in the image coordinate system, which can be obtained through sub-pixel localization. This represents the pixel displacement of the feature marker in the horizontal x-direction in the image coordinate system, which can be obtained through sub-pixel localization. This represents a parameter in the camera calibration matrix or homography transformation matrix, which is typically related to the camera's intrinsic parameters and shooting angle, and correlates pixel changes in the x-direction with physical changes in the real world. This represents another parameter in the camera calibration matrix or homography transformation matrix. This parameter is usually related to the camera's perspective transformation and shooting angle, and is used to correct the distortion caused by perspective, thereby obtaining accurate physical quantities. This indicates the coefficient of thermal expansion of the building's main materials. This represents the temperature effect compensation term (unit: same as angle). The expansion or contraction of a material due to temperature changes can be misinterpreted as tilt by the visual system. This term quantifies this error angle and adds it back (or subtracts it) from the measurement value, thereby offsetting the temperature interference and obtaining the true structural tilt.
[0065] This application's embodiments achieve multi-scale attention fusion of multimodal data to obtain a fused feature image; sub-pixel localization is performed based on the fused feature image to obtain pixel displacement; the pixel displacement is converted into vibration displacement in the world coordinate system based on a camera calibration model; and vibration feature identification is performed based on the vibration displacement to obtain the building's vibration characteristics. Because multi-scale attention fusion of multimodal data is used, it solves industry pain points such as low accuracy in recognizing long-distance and small targets and sensitivity to environmental interference; by converting pixel displacement into vibration displacement in the world coordinate system, it achieves the conversion from image signals to physical quantities, providing a direct and scientific basis for risk assessment.
[0066] Based on the first and / or second embodiments of this application, in the third embodiment of this application, the content that is the same as or similar to the first and / or second embodiments described above can be referred to the above description and will not be repeated hereafter. Based on this, please refer to... Figure 4 ,Figure 4 This is a flowchart illustrating Embodiment 3 of the intelligent monitoring method for building structural safety in this application.
[0067] In this embodiment of the application, the step of obtaining the safety monitoring results of the building by performing multi-parameter coupled calculation based on the vibration characteristics, the crack characteristics, and the tilt angle includes: Step S51: Determine the temperature data from the multimodal data; Step S52: Construct a crack evolution model based on the vibration characteristics, the tilt angle, and the temperature data; Step S53: Based on the crack characteristics and the crack evolution model, perform crack evolution to obtain the crack change rate; Step S54: Determine the safety monitoring results of the building based on the rate of change of the cracks.
[0068] It should be noted that the crack evolution module described above can be an empirical or semi-empirical physical model used to predict the rate of change of crack width over time. This crack evolution model shows that the crack propagation rate is not determined by a single factor, but is driven by the linear superposition of three physical effects: vibration stress, structural tilt, and temperature.
[0069] It is understandable that the temperature data mentioned above can be used to represent the temperature of the environment in which the building is located.
[0070] In some embodiments of this application, the crack evolution model of this application can be as follows: ; in, The weighting coefficients representing the vibration stress term are parameters obtained through data or experimental fitting, and are used to quantify vibration stress. The degree of influence on the crack propagation rate can be determined based on the vibration characteristics, and the specific method of obtaining it is not limited in the embodiments of this application. This represents the weighting coefficient for the tilt angle term of a building. This represents the weighting coefficient for the temperature term (unit: e.g., mm / (°C·time)). This is a fitting parameter used to quantify temperature. The extent of its influence on crack propagation rate. This refers to the building's tilt angle. Tilt can cause uneven distribution of structural loads, resulting in additional bending moments. This indicates that the effect is non-linear.
[0071] This application, based on the mechanical properties of materials, establishes for the first time a coupled relationship model of displacement-tilt-crack, achieving accurate risk assessment through four steps: 1) constructing a temperature compensation model for vibration displacement; 2) developing a material expansion compensation algorithm for tilt angle; 3) establishing a time-varying evolution equation for crack width; and 4) designing a decision function that integrates multiple physical quantities. This technology addresses three major industry pain points: 1) eliminating false alarms caused by temperature changes (e.g., the thermal expansion coefficient of concrete is 12 × 10⁻⁶). 6 / ℃); 2) Quantify the accelerating effect of vibration on crack propagation; 3) Achieve continuous dynamic assessment of risk values.
[0072] In some embodiments of this application, a Neural Architecture Search-Compression (NASC) technique is proposed, which deploys image feature extraction models on edge computing nodes to obtain lightweight models. This NASC technique integrates the initial models of core algorithms such as MAF for automatic search and compression, producing a neural network that is optimal in terms of accuracy, speed, and model size. This lightweight model optimized by NASC is deployed on edge devices to process video streams in real time and perform efficient feature tag recognition tasks, which is crucial for supporting the large-scale, low-cost operation of the system. Specifically, the NASC technique in this application innovatively combines hardware-aware Neural Architecture Search (NAS) with Model Compression (MC) techniques, using automated machine learning to search for and train the optimal network architecture that combines high accuracy and high efficiency. Specifically, the method further includes: optimizing a safety monitoring model using hardware-aware neural architecture search and compression technology, wherein the safety monitoring model is used to obtain the safety monitoring results of the building based on the multimodal data; the step of optimizing the safety monitoring model using hardware-aware neural architecture search and compression technology includes: constructing a multidimensional search space; performing differential architecture search and supernet training based on the multidimensional search space to obtain an optimized supernet; and processing the optimized supernet through knowledge distillation to obtain a safety monitoring model.
[0073] It is understandable that this security monitoring model is the same lightweight model mentioned above used for deployment on edge nodes.
[0074] In this embodiment, the search space of NASC covers three dimensions: operation type, channel dimension, and attention module insertion position, enabling fine-grained control over the network architecture. Specifically, the operation type can include candidate operations such as standard convolution, depthwise separable convolution, inverse residual structure, and skip connections; the channel dimension can include the number of channels in each layer, which can be searched in a preset set (e.g., {16, 24, 32, 40, 48}) to optimize computational cost; the attention module insertion position can determine the specific insertion position of the MAF attention module in the network to maximize the small target feature extraction capability with minimal computational overhead.
[0075] In some embodiments of this application, a super network can be constructed based on paths in a multi-objective search space, and the architecture parameters of the super network can be optimized using a Pareto optimal multi-objective optimization strategy. Specifically, the multi-objective loss function of NASC in this application simultaneously minimizes model error and hardware deployment costs. The multi-objective loss function L can be expressed as follows: ; in, This indicates the specific architectural parameters to be searched, namely the operation type, channel dimension, and the insertion position of the attention module. Denotes the loss function on the validation set. This represents the optimal value of the weights after training for the current architecture. This represents the hardware-aware loss term, specifically including: latency loss (Llatency): at the target edge device (… Actual measurements were obtained from NASC ( Model inference latency (in milliseconds). Parameter loss Lparams: Number of model parameters (in megabytes). Computational loss LFLOPs: Forward inference computation of the model (in gigallopian spectroscopy). This represents the tradeoff coefficient.
[0076] It should be noted that, in this embodiment, a gradient-based differential architecture search method can be used to achieve architecture search and supernetwork training. Specifically, this embodiment can first train a supernetwork containing all possible paths in a multi-dimensional search space, and introduce the aforementioned architecture parameter α to characterize the weights of the paths in the supernetwork. During training, the network weights w and architecture parameter α can be alternately optimized to update the model: ; in, ξ Indicates the learning rate; wThe network weights represent all trainable parameters of the supernetwork itself, including all possible paths, such as the kernel parameters of all candidate 3x3 convolutional layers, the kernel parameters of all candidate 5x5 convolutional layers, and the weights of all fully connected layers. This represents the loss function on the training set. This represents the partial derivative of the architecture parameter α.
[0077] In some embodiments of this application, knowledge distillation can be used to train the final model. Specifically, the rich knowledge learned by the pre-trained original large teacher model (such as an image feature extraction model) can be transferred to the lightweight student network (i.e., the super network) obtained above through knowledge distillation. This can further stabilize the training and improve the performance of the small model, enabling further processing of the optimized super network. The loss function for knowledge distillation in this application embodiment can be as follows: ; in, This represents the loss from a typical task, such as cross-entropy loss. This represents the KL divergence loss, used to bring the output logic layer of the teacher model closer together. With student network output logic layer The distribution of , where T is the temperature parameter and β is the distillation weight. Used to represent the true labels of training data during the training process. This represents the predicted output obtained by the student model during the training process when it makes predictions based on the input training data.
[0078] In one example of this application's embodiments, the model parameter size is compressed from 32.6M to 8.2M using NASC technology, achieving a compression ratio of 74.8%. Inference latency on edge devices is reduced from over 15ms to 9.3ms, meeting the real-time processing requirements of over 100 frames per second. Simultaneously, the model accuracy retention rate reaches 99.2% (mAP@0.5), a decrease of only 0.8 percentage points. Furthermore, the optimized model power consumption is reduced by over 60%, significantly extending the continuous operating time of edge devices outdoors. NASC technology successfully solves the core bottleneck problem of large-scale, embedded deployment of high-precision AI models in urban building safety monitoring scenarios.
[0079] In some embodiments of this application, the lightweight model finally obtained by this application can be as follows: Figure 5 As shown, Figure 5 This is a schematic diagram of a model structure in one implementation of the intelligent monitoring method for building structural safety in this application. Figure 5As shown, the lightweight model in this embodiment can realize functions such as data acquisition and preprocessing, feature extraction and recognition, sub-pixel localization and vibration calculation, and multi-parameter fusion and decision-making. Specifically, data acquisition and preprocessing can be achieved through synchronous acquisition by multi-modal sensors (optical / infrared / laser power supply) and data denoising and correction (distortion correction, spatiotemporal alignment). Feature extraction and recognition can include semantic segmentation network identification of cracks, MAF algorithm localization of feature markers, and calculation of tilt angle from point cloud data. Sub-pixel localization and vibration calculation can include: sub-pixel localization (gradient vector optimization), pixel-uniqueness to physical uniqueness (camera calibration model), and calculation of micro-vibration response. Multi-parameter fusion and decision-making can include: inputting multi-dimensional data (vibration / tilt / crack / temperature, etc.), multi-parameter coupled solution model, calculating risk index, and triggering early warning mechanism.
[0080] This application's embodiments determine temperature data from multimodal data; construct a crack evolution model based on vibration characteristics, tilt angle, and temperature data; perform crack evolution based on crack characteristics and the crack evolution model to obtain the crack change rate; and determine the building's safety monitoring results based on the crack change rate. Because it uses multi-dimensional data such as vibration characteristics, tilt angle, and temperature data for multi-parameter coupling, it improves the comprehensiveness and scientific rigor of risk assessment. Simultaneously, it enables trend prediction of crack evolution through the crack evolution model, achieving proactive early warning.
[0081] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the intelligent monitoring method for building structure safety of this application. Any simple modifications based on this technical concept are within the protection scope of this application.
[0082] This application also provides an intelligent monitoring device for building structural safety; please refer to [reference needed]. Figure 6 , Figure 6 This is a schematic diagram of the module structure of the intelligent building structure safety monitoring device according to an embodiment of this application. The intelligent building structure safety monitoring device includes: Data acquisition module 10 is used to acquire multimodal data of the building based on multimodal sensors; The feature recognition module 20 is used to perform multi-scale attention fusion on the multimodal data and to perform vibration feature recognition based on the fused features to obtain the vibration features of the building. Semantic segmentation module 30 is used to perform semantic segmentation on the multimodal data to obtain crack features; The tilt angle processing module 40 is used to determine the tilt angle of the building based on the multimodal data; The intelligent monitoring module 50 is used to perform multi-parameter coupled calculations based on the vibration characteristics, crack characteristics, and tilt angle to obtain the building's safety monitoring results. The intelligent building structure safety monitoring device provided in this application, employing the intelligent building structure safety monitoring method described in the above embodiments, can solve the technical problem of high risk misjudgment rates in existing intelligent building structure safety monitoring methods. Compared with the prior art, the beneficial effects of the intelligent building structure safety monitoring device provided in this application are the same as those of the intelligent building structure safety monitoring method provided in the above embodiments, and other technical features in the intelligent building structure safety monitoring device are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.
[0083] This application provides an intelligent monitoring device for building structure safety, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, which are executed by the at least one processor to enable the at least one processor to perform the intelligent monitoring method for building structure safety in Embodiment 1 described above.
[0084] The following is for reference. Figure 7 The diagram illustrates a structural schematic suitable for implementing the intelligent monitoring device for building structure safety in the embodiments of this application. The intelligent monitoring device for building structure safety in the embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Description), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 7 The intelligent monitoring device for building structural safety shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.
[0085] like Figure 7As shown, the intelligent monitoring device for building structural safety may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM) 1004. The RAM 1004 also stores various programs and data required for the operation of the intelligent monitoring device for building structural safety. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to the I / O interface 1006: input devices 1007 including, for example, a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1003 including, for example, magnetic tape, hard disk, etc.; and communication devices 1009. Communication device 1009 allows the intelligent building structure safety monitoring device to communicate wirelessly or wiredly with other devices to exchange data. Although the figure shows intelligent building structure safety monitoring devices with various systems, it should be understood that implementation or possession of all the systems shown is not required. More or fewer systems may be implemented alternatively.
[0086] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.
[0087] The intelligent monitoring device for building structure safety provided in this application, employing the intelligent monitoring method for building structure safety in the above embodiments, can solve the technical problem of high risk misjudgment rate in existing intelligent monitoring methods for building structure safety. Compared with the prior art, the beneficial effects of the intelligent monitoring device for building structure safety provided in this application are the same as those of the intelligent monitoring method for building structure safety provided in the above embodiments, and other technical features in this intelligent monitoring device are the same as those disclosed in the previous embodiment method, and will not be repeated here.
[0088] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.
[0089] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0090] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, which are used to execute the intelligent monitoring method for building structure safety described in the above embodiments.
[0091] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), or flash memory, optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.
[0092] The aforementioned computer-readable storage medium may be included in the intelligent monitoring equipment for building structural safety; or it may exist independently and not be assembled into the intelligent monitoring equipment for building structural safety.
[0093] The aforementioned computer-readable storage medium carries one or more programs, which, when executed by the intelligent monitoring device for building structural safety, cause the intelligent monitoring device for building structural safety to: Acquiring multimodal data of buildings based on multimodal sensors; Multi-scale attention fusion is performed on the multimodal data, and vibration feature identification is performed based on the fused features to obtain the vibration features of the building; Semantic segmentation is performed on the multimodal data to obtain crack features; The tilt angle of the building is determined based on the multimodal data; The safety monitoring results of the building are obtained by performing multi-parameter coupled calculations based on the vibration characteristics, crack characteristics, and tilt angle.
[0094] Computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0095] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0096] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.
[0097] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., computer programs) for executing the above-described intelligent monitoring method for building structure safety. This addresses the technical problem of high false alarm rates in existing intelligent monitoring methods for building structure safety. Compared to the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the intelligent monitoring method for building structure safety provided in the above embodiments, and will not be elaborated upon here.
[0098] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the intelligent monitoring method for building structure safety described above.
[0099] The computer program product provided in this application can solve the technical problem of high risk misjudgment rate in existing intelligent monitoring methods for building structure safety. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the intelligent monitoring method for building structure safety provided in the above embodiments, and will not be repeated here.
[0100] The above description is only a part of the embodiments of this application and does not limit the scope of protection of this application. All equivalent structural transformations made under the technical concept of this application and using the content of this application specification and drawings, or direct / indirect applications in other related technical fields, are included in the scope of protection of this application.
Claims
1. A method for intelligent monitoring of building structural safety, characterized in that, The method includes: Acquiring multimodal data of buildings based on multimodal sensors; Multi-scale attention fusion is performed on the multimodal data, and vibration feature identification is performed based on the fused features to obtain the vibration features of the building; Semantic segmentation is performed on the multimodal data to obtain crack features; The tilt angle of the building is determined based on the multimodal data; The safety monitoring results of the building are obtained by performing multi-parameter coupled calculations based on the vibration characteristics, crack characteristics, and tilt angle.
2. The intelligent monitoring method for building structural safety as described in claim 1, characterized in that, The step of performing multi-scale attention fusion on the multimodal data and identifying vibration features based on the fused features to obtain the vibration characteristics of the building includes: Multi-scale attention fusion is performed on the multimodal data to obtain a fused feature image; Sub-pixel localization is performed based on the fused feature image to obtain the pixel displacement. Based on the camera calibration model, the pixel displacement is converted into vibration displacement in the world coordinate system; Vibration characteristics of the building are obtained by identifying vibration features based on the vibration displacement.
3. The intelligent monitoring method for building structural safety as described in claim 2, characterized in that, The step of performing multi-scale attention fusion on the multimodal data to obtain a fused feature image includes: The multimodal data is processed using channel attention paths to obtain channel attention weights; The spatial attention weights are obtained by performing feature processing on the multimodal data through a spatial attention path. The fused feature image is obtained by weighted fusion based on the channel attention weight and the spatial attention weight.
4. The intelligent monitoring method for building structural safety as described in claim 1, characterized in that, The step of determining the tilt angle of the building based on the multimodal data includes: Obtain the pixel displacement correlation matrix and the thermal expansion coefficient of the main material of the building; The tilt angle of the building is determined based on the pixel displacement correlation matrix and the thermal expansion coefficient of the main material.
5. The intelligent monitoring method for building structural safety as described in claim 1, characterized in that, The step of obtaining the building's safety monitoring results by performing multi-parameter coupled calculations based on the vibration characteristics, crack characteristics, and tilt angle includes: Temperature data was determined from the multimodal data; A crack evolution model is constructed based on the vibration characteristics, the tilt angle, and the temperature data; The crack evolution is performed based on the crack characteristics and the crack evolution model to obtain the crack change rate; The safety monitoring results of the building are determined based on the rate of change of the cracks.
6. The intelligent monitoring method for building structural safety as described in claim 1, characterized in that, The method further includes: The safety monitoring model is optimized using hardware-aware neural architecture search compression technology. The safety monitoring model is used to obtain the safety monitoring results of the building based on the multimodal data. The steps for optimizing the security monitoring model using hardware-aware neural architecture search compression technology include: Constructing a multidimensional search space; Based on the multidimensional search space, differential architecture search and supernet training are performed to obtain an optimized supernet. The optimized supernetwork is processed by knowledge distillation to obtain a security monitoring model.
7. The intelligent monitoring method for building structural safety as described in claim 1, characterized in that, Before the step of vibration feature recognition based on the fused features, the following steps are included: Construct a fuzzy kernel model based on building vibration characteristics; Based on the fuzzy kernel model and through adversarial training using the generalized intersection-union loss function, a trained image feature extraction model is obtained. During adversarial training, difficult samples are adaptively weighted.
8. A building structure safety intelligent monitoring device, characterized in that, The intelligent monitoring device for building structural safety includes: The data acquisition module is used to acquire multimodal data of the building based on multimodal sensors; The feature recognition module is used to perform multi-scale attention fusion on the multimodal data and to perform vibration feature recognition based on the fused features to obtain the vibration features of the building. A semantic segmentation module is used to perform semantic segmentation on the multimodal data to obtain crack features; A tilt angle processing module is used to determine the tilt angle of the building based on the multimodal data; The intelligent monitoring module is used to perform multi-parameter coupled calculations based on the vibration characteristics, crack characteristics, and tilt angle to obtain the safety monitoring results of the building.
9. An intelligent monitoring device for building structural safety, characterized in that, The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the intelligent monitoring method for building structural safety as described in any one of claims 1 to 7.
10. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the intelligent monitoring method for building structure safety as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Super-resolution deblurring method of a generated antagonistic network
CN109377459A
Complex environment power transmission line foreign matter detection method based on deep learning
CN118506263A
Dam safety perception fusion association method based on multi-modal space-time diagram neural network
CN121167580A
House structure safety edge visual monitoring method and device, equipment and storage medium
CN121236608A
System and method for supporting emergency recovery based on multi-sensors
KR102781742B1