AI intelligent visual disaster monitoring and alarming method and system
Through AI intelligent visual disaster monitoring methods, using visible light and infrared image feature fusion and time series feature extraction, the problems of delayed response and high false alarm rate of traditional monitoring methods are solved, and timely and accurate monitoring of sudden disasters such as bridges and tunnels is achieved.
Patent Information
- Application Number
- CN202510747216.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-09-12
AI Technical Summary
Traditional monitoring methods have problems with response lag and high false alarm rate in monitoring sudden disasters such as bridge collapse, rockfall at tunnel entrances, rockfall on slopes, landslides, and mudslides.
An AI intelligent visual disaster monitoring method is adopted to obtain visible light and infrared images, perform feature fusion and time series feature extraction, use the state assessment network to evaluate the health status of the monitored target, and combine the long-short-term memory subnetwork and the attention subnetwork to improve monitoring accuracy.
It improves the timeliness and accuracy of disaster monitoring, reduces the false alarm rate, and realizes accurate monitoring of disasters and sensitivity assessment of early failures.
Smart Images

Figure CN120636095A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of visual detection technology, and more specifically, relates to an AI intelligent visual disaster monitoring and alarm method and system. Background Art
[0002] Sudden disasters such as bridge collapse, rockfall at tunnel entrances, rockfall on slopes, landslides, and mudslides are highly concealed, evolve rapidly, and have great destructive power. Accurate monitoring of the above-mentioned sudden disasters can ensure the safety of infrastructure and life and property.
[0003] Traditional monitoring methods mainly rely on manual inspections or the deployment of sensor monitoring, which has pain points such as delayed response and high false alarm rates. Summary of the Invention
[0004] The purpose of this application is to provide an AI intelligent visual disaster monitoring and alarm method and system to improve the timeliness and accuracy of disaster monitoring.
[0005] A first aspect of an embodiment of the present application provides an AI intelligent visual disaster monitoring and alarm method, comprising: Acquire multiple continuous target visible light images and multiple continuous target infrared images collected within a preset time period, each of the target visible light images and each of the target infrared images contains the monitored target; each continuous target visible light image and each continuous target infrared image have a one-to-one correspondence in terms of collection time; each collection time corresponds to an image pair, and each image pair includes one target visible light image and one target infrared image; For each acquisition time, feature extraction is performed on the target visible light image and target infrared image in the image pair corresponding to the acquisition time, and the extracted features are fused to obtain a fused feature map; The health status of the monitored target is obtained by performing the following operations on the multiple fused feature maps through the state assessment network: Extracting time series features from the multiple fused feature maps to obtain a first time series feature of the monitoring target; The health status of the monitoring target is evaluated based on the first time series feature; the health status of the monitoring target is used to indicate whether a disaster occurs to the monitoring target.
[0006] A second aspect of the embodiments of the present application provides an AI intelligent visual disaster monitoring and alarm device, comprising: An image acquisition module is configured to acquire a plurality of continuous visible light images and a plurality of continuous infrared images of a target acquired within a preset time period, wherein each of the visible light images and the infrared images of the target contain the monitored target; each of the continuous visible light images and the continuous infrared images of the target correspond to each other in terms of acquisition time; each acquisition time corresponds to an image pair, and each image pair includes one visible light image and one infrared image of the target; A feature fusion module is used to extract features from the target visible light image and the target infrared image in the image pair corresponding to each acquisition time, and fuse the extracted features to obtain a fused feature map; The state assessment module is used to perform the following operations on multiple fused feature maps through the state assessment network to obtain the health status of the monitoring target: Extracting time series features from the multiple fused feature maps to obtain a first time series feature of the monitoring target; The health status of the monitoring target is evaluated based on the first time series feature; the health status of the monitoring target is used to indicate whether a disaster occurs to the monitoring target.
[0007] The third aspect of the embodiments of the present application provides an AI intelligent visual disaster monitoring and alarm system, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, the steps of the above-mentioned AI intelligent visual disaster monitoring and alarm method are implemented.
[0008] In a fourth aspect of an embodiment of the present application, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned AI intelligent visual disaster monitoring and alarm method are implemented.
[0009] The AI intelligent visual disaster monitoring and alarm method and system provided by the embodiments of the present application have the following beneficial effects: The embodiment of the present application can enhance the feature expression capability of the monitored target by fusing the features of the target visible light image and the target infrared image to obtain a fused feature map. By performing time series extraction on the fused feature map to obtain the first time series feature, the dynamic law of disaster development of the monitored target can be captured, thereby achieving accurate assessment of the health status of the monitored target, that is, achieving accurate monitoring of disasters.
[0010] Compared with the existing human patrol method, the method of this embodiment can improve the timeliness of disaster monitoring; compared with the existing sensor monitoring method, the method of this embodiment can improve the accuracy of disaster monitoring. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments or descriptions of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0012] Figure 1 This is a schematic block diagram of an AI intelligent visual disaster monitoring and alarm system provided in one embodiment of the present application; Figure 2 A flowchart of an AI intelligent visual disaster monitoring and alarm method provided in one embodiment of the present application; Figure 3 This is a structural block diagram of the AI intelligent visual disaster monitoring and alarm device provided in one embodiment of the present application; Figure 4 This is a schematic block diagram of an on-site AI analysis and early warning module provided in one embodiment of the present application. DETAILED DESCRIPTION
[0013] In the following description, specific details such as specific system structures and techniques are provided for purposes of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present application with unnecessary detail.
[0014] In order to make the purpose, technical solutions and advantages of this application clearer, specific embodiments will be described below with reference to the accompanying drawings.
[0015] Please refer to Figure 1 The method of this embodiment can be applied to an AI intelligent visual disaster monitoring and alarm system. This system includes on-site cameras and an on-site AI analysis and warning module. The on-site cameras may include visible light cameras and infrared cameras deployed at key road sections. The visible light cameras are used to capture visible light images of targets containing monitored targets, and the infrared cameras are used to capture infrared images of targets containing monitored targets. The on-site AI analysis and warning module is used to identify the visible light and infrared images of the targets to determine their health status.
[0016] In addition, the AI intelligent visual disaster monitoring and alarm system can also include a data acquisition, processing and control terminal, which can collect the recognition results of the on-site AI analysis and early warning module and send them to the cloud server. In this way, staff can download and view the recognition results of the on-site AI analysis and early warning module from the cloud server on a PC or mobile terminal, monitor the health status of the target in real time, and detect disasters in a timely manner.
[0017] Please refer to Figure 2 , Figure 2 This is a flowchart of an AI intelligent visual disaster monitoring and alarm method provided in one embodiment of the present application. The method may be executed by an on-site AI analysis and early warning module in an AI intelligent visual disaster monitoring and alarm system. The method may include: S101: Acquire multiple continuous target visible light images and multiple continuous target infrared images collected within a preset time period, each target visible light image and each target infrared image contains a monitored target; each continuous target visible light image and each continuous target infrared image has a one-to-one correspondence in terms of collection time; one collection time corresponds to one image pair, and one image pair includes one target visible light image and one target infrared image.
[0018] In this embodiment, the monitoring target can be a bridge, tunnel entrance, slope, landslide, etc. By deploying visible light cameras at multiple locations of the monitoring target, a target visible light image containing the monitoring target can be obtained. Correspondingly, by deploying infrared cameras at multiple locations of the monitoring target, a target infrared image containing the monitoring target can be obtained.
[0019] S102: For each acquisition time, feature extraction is performed on the target visible light image and the target infrared image in the image pair corresponding to the acquisition time, and the extracted features are fused to obtain a fused feature map.
[0020] In this embodiment, feature extraction from the target's visible light image can reveal image features such as texture and color, thereby revealing structural details of the monitored target, such as cracks in bridges and rock mass morphology on slopes. Feature extraction from the target's infrared image can also capture the temperature distribution on the target's surface (e.g., abnormal structural heating and temperature gradient changes in the rock mass). Visible light images are easily affected by lighting and weather (e.g., blurry images at night or in rainy and foggy conditions), while infrared images can partially penetrate clouds and fog and are insensitive to lighting. Therefore, by fusing the features extracted from the target's visible light image with those from its infrared image, a fused feature map containing both visual and thermal information can be generated, compensating for the environmental adaptability limitations of a single modality.
[0021] S103: Multiple fused feature maps are processed through the state assessment network to perform the following operations to obtain the health status of the monitored target: Extract time series features from multiple fused feature maps to obtain the first time series features of the monitored target; The health status of the monitoring target is evaluated based on the first time series feature; the health status of the monitoring target is used to characterize whether a disaster occurs to the monitoring target.
[0022] In this embodiment, considering that there are usually time series characteristics from the initiation of disaster hazards to the occurrence of disasters (such as the gradual increase in slope displacement rate and abnormal fluctuations in bridge vibration frequency), by extracting time series features from multiple fused feature maps, the evolution pattern of features over time (such as the continuous increase in temperature gradient in a certain area and the periodic expansion of crack width) can be captured to avoid misjudgment of single-frame images.
[0023] On the basis of obtaining the first time series feature, the first time series feature can be input into a classifier, and the health state of the monitored target is determined to be a normal state or an abnormal state according to the output of the classifier.
[0024] From the above, it can be concluded that this embodiment can enhance the feature expression ability of the monitored target by fusing the features of the target visible light image and the target infrared image to obtain a fused feature map. By performing time series extraction on the fused feature map to obtain the first time series feature, the dynamic law of disaster development of the monitored target can be captured, thereby achieving accurate assessment of the health status of the monitored target, that is, achieving accurate monitoring of disasters.
[0025] Compared with the existing human patrol method, the method of this embodiment can improve the timeliness of disaster monitoring; compared with the existing sensor monitoring method, the method of this embodiment can improve the accuracy of disaster monitoring.
[0026] In one embodiment of the present application, time series feature extraction is performed on multiple fused feature maps to obtain a first time series feature of the monitoring target, including: The fused feature map is input into the long short-term memory sub-network to obtain the second time series features corresponding to each time step in multiple time steps; The attention sub-network performs weighted summation on the second time series features corresponding to each time step to obtain the first time series features of the monitored target.
[0027] In this embodiment, a long short-term memory (LSTM) subnetwork can be used to implement preliminary encoding of temporal features, compressing the spatial information (such as crack locations and hotspot areas) and temporal information (such as feature values at each moment) of the fused feature map into a hidden state (second temporal feature) that contains temporal dependencies. On this basis, the weight distribution mechanism of the attention subnetwork highlights key disaster-related time steps, forming a more compact and discriminative first temporal feature, which facilitates accurate judgment of the disaster status.
[0028] In one embodiment of the present application, weighted summation is performed on the second time series features corresponding to each time step to obtain the first time series features of the monitoring target, including: Determine the weights corresponding to each time step; Based on the weights corresponding to the respective time steps, the second time series features corresponding to the respective time steps are weighted summed to obtain the first time series features of the monitoring target; The method of determining the weight corresponding to each time step includes: Calculate the reference value of the anomaly score at this time step; The actual displacement of the monitoring target is obtained based on a displacement sensor provided on the monitoring target, the theoretical displacement of the monitoring target is determined based on the ambient temperature of the monitoring target, and the displacement error of the monitoring target is determined based on the actual displacement and the theoretical displacement; The actual strain data of the monitoring target is obtained based on the strain sensor set on the monitoring target, the theoretical strain information of the monitoring target is determined based on the structural mechanics model of the monitoring target, and the strain error of the monitoring target is determined based on the actual strain data and the theoretical strain data; Adjusting the reference value of the anomaly score of the time step based on the displacement error and the strain error to obtain the anomaly score of the time step; The weight corresponding to the time step is determined based on the positive correlation between the anomaly score and the weight of the time step.
[0029] In this embodiment, the weights corresponding to each time step can be calculated based on the attention module, and then the second time series features corresponding to each time step can be weighted and summed to obtain the first time series features of the monitored target.
[0030] Specifically, we can first calculate the anomaly score of each time step. The higher the anomaly score of the time step, the more significant the deviation of the corresponding second time series feature from the normal state, which may contain more key information about potential faults or anomalies. By assigning higher weights to these time steps, we can pay more attention to the characteristic patterns under abnormal conditions and improve the sensitivity to early faults. The specific calculation process of the anomaly score is as follows: (1) Calculate the reference value of the anomaly score: Assume that the output of the LSTM module is: ;in, represents the input data at the tth time step, represents the hidden state of the t-th time step, which incorporates historical sequence information.
[0031] Through attention calculation, the attention weight of each time step is obtained: ; in, represents the feature importance at the t-th time step, represents the normalized attention weight, 、 , are the parameter matrices of LSTM module and attention module, represents the attention weight matrix, represents the hidden state weight matrix, represents the attention bias term, and T represents the total number of time steps.
[0032] The hidden states of the key time steps are aggregated through the above attention weights to form a context vector Q containing the timing focus, as shown below: ; The above context vector Q is activated using the Sigmoid function to obtain the probability that the t-th time step is an abnormal event, that is, the abnormality score, which is specifically: ; in, represents the anomaly score at the t-th time step, represents the decoder state weight matrix, Represents the decoder bias term.
[0033] (2) Adjust the reference value of the anomaly score based on physical constraints: In this embodiment, considering that in the health monitoring of structures such as bridges and tunnels, the structural response (such as displacement and strain) is physically related to environmental factors (temperature and load), prior knowledge such as thermal expansion and contraction formulas and structural mechanics models can be incorporated into the calculation of anomaly scores to reduce false positives.
[0034] Taking a bridge as an example, for displacement constraints, the thermal expansion coefficient of the bridge material can be set to , the current temperature is , the length of the bridge component is L, and the theoretical displacement of the bridge component is obtained according to the thermal expansion and contraction formula: ; in, represents the theoretical displacement, Indicates the reference value of temperature.
[0035] Comparing the above theoretical displacement with the actual displacement actually monitored by the displacement sensor, the displacement error can be obtained as: ; in, represents the displacement error at the tth time step, represents the actual displacement at the tth time step.
[0036] To address structural mechanics constraints, finite element analysis can be performed based on the structural mechanics model of the bridge (e.g., a finite element model) to obtain theoretical strain data for the bridge structure. The theoretical strain data is then compared with the actual strain data obtained by displacement sensor monitoring to obtain the strain error: ; in, represents the strain error at the t-th time step, represents the actual strain data at the t-th time step, represents theoretical strain data.
[0037] The reference value of the anomaly score is adjusted based on the above displacement error and strain error to obtain the anomaly score at the t-th time step: ; in, represents the anomaly score at the t-th time step, 、 These are all preset scale parameters.
[0038] Based on the abnormal scores of each time step, the abnormal scores of each time step can be normalized to obtain the weight of each time step. : .
[0039] From the above, it can be concluded that this embodiment comprehensively considers the calculation results of the attention module and physical constraints to determine the weight of each time step, thereby achieving accurate assessment of abnormal time points (critical time points) of the monitoring target and improving sensitivity to early faults.
[0040] In one embodiment of the present application, feature extraction is performed on each of the target visible light image and the target infrared image in the image pair corresponding to the acquisition time, and the extracted features are fused to obtain a fused feature map, including: Based on the first convolutional neural network, features of the target visible light image and the target infrared image are extracted respectively, and the features extracted from the target visible light image and the features extracted from the target infrared image are fused; Among them, the first convolutional neural network is trained in the following way: Acquire a plurality of first sample image pairs, each first sample image pair comprising a visible light sample image and a corresponding infrared sample image; Performing sample expansion on the plurality of first sample image pairs to obtain a plurality of second sample image pairs; Based on the plurality of first sample image pairs and the plurality of second sample image pairs, an initial convolutional neural network is trained to obtain a first convolutional neural network.
[0041] In this embodiment, the hierarchical characteristics of the first convolutional neural network can be used to extract features at different levels in the target visible light image and the target infrared image in stages to obtain a multi-scale feature map. Among them, shallow features (1~2 layers) are used to capture local details, such as texture, edges, light and dark changes, etc., and deep features (3~5 layers) are used to capture semantic information, such as vehicles, bridge cracks, etc. Specifically, two branches can be used to extract features of the target visible light image and the target infrared image respectively, and the shared weights of the two branches can reduce parameter redundancy and force cross-modal features to be aligned in the same semantic space. Finally, after the shallow features and deep features are aligned in size through upsampling / downsampling, they are cross-layer splicing or addition to obtain a fused feature map.
[0042] Considering that the training process of the first convolutional neural network requires a large amount of paired annotated data, cross-modal image datasets are scarce in scenes such as bridges, tunnels, and slopes, especially lacking annotated samples containing damage, resulting in insufficient model generalization. To address this issue, this embodiment performs sample expansion on the existing first sample image pairs to obtain second sample image pairs, and then trains the first convolutional neural network based on the first and second sample image pairs.
[0043] Specifically, the first sample image pair and the second sample image pair can be input into the initial convolutional neural network, the health status of the monitored target can be used as the true label of the pre-trained state evaluation network, and the initial convolutional neural network and the pre-trained state evaluation network can be jointly trained to obtain the trained first convolutional neural network and state evaluation network.
[0044] From the above, it can be concluded that this embodiment trains the first convolutional neural network based on the expanded sample data, which is beneficial to improving the adaptability of the first convolutional neural network to complex scenarios and improving the robustness and generalization ability of the first convolutional neural network.
[0045] In one embodiment of the present application, performing sample expansion on a plurality of first sample image pairs to obtain a plurality of second sample image pairs includes: The adversarial generative network is used to perform sample expansion on the first sample image pair to obtain the second sample image pair; the loss function of the adversarial generative network is: ; in, represents the loss function, represents the cross-modal adversarial loss, represents the semantic alignment loss, represents the cycle consistency loss, represents the adversarial feature matching loss, , , , All are weight coefficients; ; in, Represents the real visible light image judged by the discriminator Belongs to the real data distribution The probability of Represents the generated visible light image obtained by the discriminator authenticity, represents the visible light discriminator; Represents the real visible light-infrared image pair judged by the discriminator Belongs to the real paired data distribution The probability of Represents the real visible light-generated infrared image pair judged by the discriminator The rationality of pairing, represents the cross-modal pairing discriminator; ; in, represents the semantic encoding of any category c in the visible light image, represents the indicator function, represents the infrared feature extractor, represents the feature mean of any category c in the real infrared image; ; in, represents the inverse generator, which is used to generate visible light images based on infrared images. represents the expected value of the infrared image y; ; Among them, the adversarial feature matching loss The calculation is obtained by using the pre-trained second convolutional neural network to extract the visible light image and generate the high-level semantic feature map of the infrared image. Indicates that the feature map is The height of the layer, Indicates that the feature map is The width of the layer, Indicates that the feature map is The number of channels of the layer, Represents the pre-trained CNN network The feature extraction function of the layer.
[0046] In this embodiment, a generative adversarial network can be used to generate image pairs of visible light images and infrared images. For the scenario of cross-modal image generation, the loss function of the generative adversarial network includes the following four parts: (1) Cross-modal adversarial loss
[0047] For cross-modal image generation, a dual-branch structure can be introduced into the discriminator to determine the unimodal authenticity and cross-modal pairing rationality of visible-infrared image pairs. The unimodal branch determines whether the generated infrared image matches the spectral distribution of the real infrared image (e.g., temperature correlation) and whether the visible image is authentic (either retaining the original input or enhancing its authenticity). The cross-modal branch determines whether the visible-infrared image pair matches in semantic regions (e.g., the location of a bridge crack in the visible light image aligns with the temperature anomaly in the infrared image).
[0048] Specifically, cross-modal adversarial loss The calculation formula is as follows: ; in, Represents the real visible light image judged by the discriminator Belongs to the real data distribution The probability of Represents the generated visible light image obtained by the discriminator authenticity, represents the visible light discriminator; Represents the real visible light-infrared image pair judged by the discriminator Belongs to the real paired data distribution The probability of Represents the real visible light-generated infrared image pair judged by the discriminator The rationality of pairing, represents the cross-modal paired discriminator.
[0049] (2) Semantic alignment loss
[0050] In this implementation, semantic alignment loss is calculated by introducing semantic segmentation constraints. Specifically, a pre-trained semantic segmentation model (such as DeepLab) can be used to extract semantic labels (such as "crack", "concrete", and "rebar") of visible light images, forcing the generated infrared images to have signals with specific physical meanings in the same semantic areas (such as crack areas corresponding to temperature anomalies in infrared).
[0051] Based on this, we get the semantic alignment loss The calculation formula is as follows: ; in, represents the semantic encoding of any category c in the visible light image, represents the indicator function, represents the infrared feature extractor, Represents the feature mean of any category c in the real infrared image.
[0052] (3) Cycle consistency loss
[0053] In this embodiment, an inverse generator (for generating a visible light image based on an infrared image) may be added, and the generated infrared image is required to restore the key features (such as texture details) of the original visible light image after passing through the inverse generator to avoid information loss.
[0054] Based on this, we get the cycle consistency loss The calculation formula is: ; in, represents the inverse generator, which is used to generate visible light images based on infrared images. Represents the expected value of the infrared image y.
[0055] (4) Resistance feature matching loss
[0056] In this embodiment, a pre-trained second convolutional neural network (such as VGG) can be used to extract high-level semantic features of visible light and generate infrared images, and the feature distributions of the two are forced to be close in the loss function to ensure that the features of cross-modal images at the semantic level (such as "damage" and "structural abnormality") are consistent.
[0057] Based on this, we get the resistance feature matching loss The calculation formula is: ; in, Indicates that the feature map is The height of the layer, Indicates that the feature map is The width of the layer, Indicates that the feature map is The number of channels of the layer, Represents the pre-trained CNN network The feature extraction function of the layer.
[0058] From the above, it can be concluded that this embodiment, based on the scenario of cross-module image generation, introduces a cross-modal adversarial loss to force the cross-modal image output by the generator to be as close as possible to the real image distribution of the target modality; by introducing a semantic alignment loss, it ensures that the cross-modal images remain consistent in high-level semantics (such as object category, structure, and damage location); by introducing a cycle consistency loss, it ensures the reversibility of cross-modal conversion, that is, the "generation-restoration" process can restore the original image and avoid mode collapse (insufficient diversity of generated images); by introducing an adversarial feature matching loss, it compares the features extracted by the intermediate layer of the discriminator (rather than only the true and false discrimination results of the output layer), forcing the generated image to be close to the real image in multi-level features (such as edges, textures, and semantics), which helps to address the limitation of traditional adversarial losses that only focus on low-level visual similarity, and improves the semantic rationality and detail richness of the generated image.
[0059] In one embodiment of the present application, the AI intelligent visual disaster monitoring and alarm method further includes: Set the weight of the semantic alignment loss when the number of training rounds of the adversarial network is less than or equal to the set number are respectively greater than the weight of the cross-modal adversarial loss , the weight of cycle consistency loss and the weight of the adversarial feature matching loss ; Set the weight of the cross-modal adversarial loss when the number of training rounds of the adversarial network exceeds the set number are greater than the weight of the semantic alignment loss and the weight of the cycle consistency loss , and the weight of the adversarial feature matching loss are greater than the weight of the semantic alignment loss and the weight of the cycle consistency loss .
[0060] In this embodiment, considering that in the early stage of adversarial generative network training, there is insufficient understanding of the cross-modal mapping relationship, it is necessary to prioritize ensuring the high-level semantic consistency between the generated image and the source image (such as key information such as disaster location and type). At this time, the weight of the semantic alignment loss can be increased to force the model to focus on the correct mapping at the semantic level, avoiding semantic deviation (such as incorrect generation of the bridge crack position) due to excessive pursuit of visual realism (cross-modal adversarial loss) or feature details (adversarial feature matching loss).
[0061] Correspondingly, in the later stages of GAN training, semantic alignment has been initially mastered, and the visual realism and feature space consistency of the generated images need to be further improved. At this time, the weight of the cross-modal adversarial loss can be increased, and the adversarial game can be used to force the generated images to be close to the real data in terms of underlying visual features such as color and texture, thereby enhancing the model's ability to fit the target modal distribution. At the same time, by increasing the weight of the adversarial feature matching loss, starting from the discriminator's intermediate layer features, the generated images can be constrained to be consistent with the real images in multi-level features (such as edges and structures), thereby solving the limitation of traditional GANs that only focus on low-level visual similarities.
[0062] From the above, it can be concluded that this embodiment increases the weight of the semantic alignment loss in the early stages of GAN training to ensure that key disaster information (such as cracks and deformations) is not lost or misplaced during cross-modal conversion. In the later stages of GAN training, the cross-modal adversarial loss and adversarial feature matching loss are enhanced to make the generated images visually closer to the real data.
[0063] In one embodiment of the present application, the health status of the monitored target includes multiple health status levels, and the AI intelligent visual disaster monitoring and alarm method further includes: If the health status level of the monitored target is a specified health status level, a first alarm signal is output.
[0064] In this embodiment, multiple classifiers can be set up in the output layer of the state assessment network to classify the health status of the monitored target into multiple levels, such as normal, warning, and fault, so that a graded response can be made based on the health status level. For example, when the health status level is fault, a first alarm signal can be output to remind staff to carry out emergency repairs in a timely manner; when the health status level is warning, staff can be reminded to carry out maintenance in a timely manner.
[0065] refer to Figure 1 In one embodiment of the present application, when the health status level is fault, the on-site AI analysis and early warning module outputs a first alarm signal to the data acquisition, processing and control terminal, and the data acquisition, processing and control terminal controls the on-site sound and light alarm to sound an alarm, reminding passing vehicles to detour to ensure vehicle safety.
[0066] In one embodiment of the present application, the on-site AI analysis and early warning module can only output the health status level to the data acquisition, processing and control terminal, which determines whether the monitored target is faulty. When the monitored target is judged to be in a faulty state, the data acquisition, processing and control terminal controls the on-site sound and light alarm to sound an alarm, reminding passing vehicles to detour to ensure vehicle safety.
[0067] Corresponding to the AI intelligent visual disaster monitoring and alarm method of the above embodiment, Figure 3This is a structural block diagram of an AI intelligent visual disaster monitoring and alarm device provided in one embodiment of the present application. For ease of explanation, only the parts related to the embodiment of the present application are shown. Figure 3 The AI intelligent visual disaster monitoring and alarm device 20 includes: an image acquisition module 21, a feature fusion module 22 and a state assessment module. The image acquisition module 21 is configured to acquire a plurality of continuous visible light images and a plurality of continuous infrared images of a target captured within a preset time period, wherein each visible light image and each infrared image contains the monitored target; each continuous visible light image and each continuous infrared image of the target have a one-to-one correspondence in terms of acquisition time; each acquisition time corresponds to an image pair, and each image pair includes one visible light image and one infrared image of the target; A feature fusion module 22 is configured to extract features from the target visible light image and the target infrared image in the image pair corresponding to each acquisition time, and fuse the extracted features to obtain a fused feature map; The state evaluation module 23 is used to perform the following operations on the multiple fused feature maps through the state evaluation network to obtain the health status of the monitored target: Extract time series features from multiple fused feature maps to obtain the first time series features of the monitored target; The health status of the monitoring target is evaluated based on the first time series feature; the health status of the monitoring target is used to characterize whether a disaster occurs to the monitoring target.
[0068] In one embodiment of the present application, the status assessment module 23 is specifically configured to: The fused feature map is input into the long short-term memory sub-network to obtain the second time series features corresponding to each time step in multiple time steps; The attention sub-network performs weighted summation on the second time series features corresponding to each time step to obtain the first time series features of the monitored target.
[0069] In one embodiment of the present application, the status assessment module 23 is further configured to: Determine the weights corresponding to each time step; Based on the weights corresponding to the respective time steps, the second time series features corresponding to the respective time steps are weighted summed to obtain the first time series features of the monitoring target; The method of determining the weight corresponding to each time step includes: Calculate the reference value of the anomaly score at this time step; The actual displacement of the monitoring target is obtained based on a displacement sensor provided on the monitoring target, the theoretical displacement of the monitoring target is determined based on the ambient temperature of the monitoring target, and the displacement error of the monitoring target is determined based on the actual displacement and the theoretical displacement; The actual strain data of the monitoring target is obtained based on the strain sensor set on the monitoring target, the theoretical strain information of the monitoring target is determined based on the structural mechanics model of the monitoring target, and the strain error of the monitoring target is determined based on the actual strain data and the theoretical strain data; Adjusting the reference value of the anomaly score of the time step based on the displacement error and the strain error to obtain the anomaly score of the time step; The weight corresponding to the time step is determined based on the positive correlation between the anomaly score and the weight of the time step.
[0070] In one embodiment of the present application, the feature fusion module 22 is specifically configured to: Based on the first convolutional neural network, features of the target visible light image and the target infrared image are extracted respectively, and the features extracted from the target visible light image and the features extracted from the target infrared image are fused; Among them, the first convolutional neural network is trained in the following way: Acquire a plurality of first sample image pairs, each first sample image pair comprising a visible light sample image and a corresponding infrared sample image; Performing sample expansion on the plurality of first sample image pairs to obtain a plurality of second sample image pairs; Based on the plurality of first sample image pairs and the plurality of second sample image pairs, an initial convolutional neural network is trained to obtain a first convolutional neural network.
[0071] In one embodiment of the present application, the feature fusion module 22 is further configured to: The adversarial generative network is used to perform sample expansion on the first sample image pair to obtain the second sample image pair; the loss function of the adversarial generative network is: ; in, represents the loss function, represents the cross-modal adversarial loss, represents the semantic alignment loss, represents the cycle consistency loss, represents the adversarial feature matching loss, , , , All are weight coefficients; ; in, Represents the real visible light image judged by the discriminator Belongs to the real data distribution The probability of Represents the generated visible light image obtained by the discriminator authenticity, represents the visible light discriminator; Represents the real visible light-infrared image pair judged by the discriminator Belongs to the real paired data distribution The probability of Represents the real visible light-generated infrared image pair judged by the discriminator The rationality of pairing, represents the cross-modal pairing discriminator; ; in, represents the semantic encoding of any category c in the visible light image, represents the indicator function, represents the infrared feature extractor, represents the feature mean of any category c in the real infrared image; ; in, represents the inverse generator, which is used to generate visible light images based on infrared images. represents the expected value of the infrared image y; ; Among them, the adversarial feature matching loss The calculation is obtained by using the pre-trained second convolutional neural network to extract the visible light image and generate the high-level semantic feature map of the infrared image. Indicates that the feature map is The height of the layer, Indicates that the feature map is The width of the layer, Indicates that the feature map is The number of channels of the layer, Represents the pre-trained CNN network The feature extraction function of the layer.
[0072] In one embodiment of the present application, the feature fusion module 22 is further configured to: Set the weight of the semantic alignment loss when the number of training rounds of the adversarial network is less than or equal to the set number are respectively greater than the weight of the cross-modal adversarial loss , the weight of cycle consistency loss and the weight of the adversarial feature matching loss ; Set the weight of the cross-modal adversarial loss when the number of training rounds of the adversarial network exceeds the set number are greater than the weight of the semantic alignment loss and the weight of the cycle consistency loss , and the weight of the adversarial feature matching loss are greater than the weight of the semantic alignment loss and the weight of the cycle consistency loss .
[0073] In one embodiment of the present application, the health status of the monitored target includes multiple health status levels, and the status assessment module 23 is specifically configured to: If the health status level of the monitored target is a specified health status level, a first alarm signal is output.
[0074] See also Figure 4 , Figure 4 This is a schematic block diagram of an AI intelligent visual disaster monitoring and alarm system provided in one embodiment of the present application. Figure 4 The AI intelligent visual disaster monitoring and alarm system 300 in the embodiment shown may include: one or more processors 301, one or more input devices 302, one or more output devices 303 and one or more memories 304. The above-mentioned processors 301, input devices 302, output devices 303 and memories 304 communicate with each other through a communication bus 305. The memory 304 is used to store computer programs, which include program instructions. The processor 301 is used to execute the program instructions stored in the memory 304. Among them, the processor 301 is configured to call the program instructions to execute the functions of each module / unit in the above-mentioned device embodiments, such as Figure 3 The functions of the image acquisition module 21, the feature fusion module 22 and the state assessment module 23 are shown.
[0075] It should be understood that in the embodiment of the present application, the processor 301 may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.
[0076] The input device 302 may include a touchpad, a fingerprint collection sensor (for collecting user fingerprint information and fingerprint direction information), a microphone, etc. The output device 303 may include a display (LCD, etc.), a speaker, etc.
[0077] The memory 304 may include a read-only memory and a random access memory, and provides instructions and data to the processor 301. A portion of the memory 304 may also include a non-volatile random access memory. For example, the memory 304 may also store preset constants such as a first threshold value and a first step length.
[0078] In the specific implementation, the processor 301, input device 302, and output device 303 described in the embodiment of this application can execute the implementation method described in the AI intelligent visual disaster monitoring and alarm method provided in the embodiment of this application, and can also execute the implementation method of the AI intelligent visual disaster monitoring and alarm system described in the embodiment of this application, which will not be repeated here.
[0079] In another embodiment of the present application, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program. The computer program includes program instructions. When the program instructions are executed by a processor, all or part of the process of the method in the above embodiment is implemented. The computer program can also be used to instruct related hardware to complete the process. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor, the steps of each of the above method embodiments are implemented. The computer program includes computer program code, which can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium can include: any entity or device capable of carrying computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal and software distribution medium.
[0080] The computer-readable storage medium can be the internal storage unit of the AI intelligent visual disaster monitoring and alarm system in any of the aforementioned embodiments, such as the hard disk or memory of the AI intelligent visual disaster monitoring and alarm system. The computer-readable storage medium can also be an external storage device of the AI intelligent visual disaster monitoring and alarm system, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a flash memory card, etc. equipped with the AI intelligent visual disaster monitoring and alarm system. Furthermore, the computer-readable storage medium can include both the internal storage unit of the AI intelligent visual disaster monitoring and alarm system and an external storage device. The computer-readable storage medium is used to store computer programs and other programs and data required by the AI intelligent visual disaster monitoring and alarm system. The computer-readable storage medium can also be used to temporarily store data that has been output or is about to be output.
[0081] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0082] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the AI intelligent visual disaster monitoring and alarm system and unit described above can refer to the corresponding process in the aforementioned method embodiment and will not be repeated here.
[0083] In the several embodiments provided in this application, it should be understood that the disclosed AI intelligent visual disaster monitoring and alarm system and method can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of units is only a logical function division. There may be other division methods in actual implementation. For example, modules can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces or units, or it can be an electrical, mechanical or other form of connection.
[0084] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the embodiments of the present application.
[0085] In addition, the functional modules in the various embodiments of the present application may be integrated into a processing module, or each module may exist physically separately, or two or more modules may be integrated into a single module. The above-mentioned integrated modules may be implemented in the form of hardware or software functional modules.
[0086] The above are only specific embodiments of the present application, but the scope of protection of the present application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and such modifications or substitutions should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. An AI intelligent visual disaster monitoring and alarm method, characterized in that: include: Acquire multiple continuous target visible light images and multiple continuous target infrared images collected within a preset time period, each of the target visible light images and each of the target infrared images contains the monitored target; each continuous target visible light image and each continuous target infrared image have a one-to-one correspondence in terms of collection time; each collection time corresponds to an image pair, and each image pair includes one target visible light image and one target infrared image; For each acquisition time, feature extraction is performed on the target visible light image and target infrared image in the image pair corresponding to the acquisition time, and the extracted features are fused to obtain a fused feature map; The health status of the monitored target is obtained by performing the following operations on the multiple fused feature maps through the state assessment network: Extracting time series features from the multiple fused feature maps to obtain a first time series feature of the monitoring target; The health status of the monitoring target is evaluated based on the first time series feature; the health status of the monitoring target is used to indicate whether a disaster occurs to the monitoring target.
2. The AI intelligent visual disaster monitoring and alarm method according to claim 1, characterized in that: The extracting time series features from the multiple fused feature maps to obtain a first time series feature of the monitoring target includes: Inputting the fused feature map into the long short-term memory sub-network to obtain a second time series feature corresponding to each time step in the multiple time steps; The attention sub-network performs weighted summation on the second time series features corresponding to each time step to obtain the first time series features of the monitoring target.
3. The AI intelligent visual disaster monitoring and alarm method according to claim 2, characterized in that: The weighted summing of the second time series features corresponding to the respective time steps to obtain the first time series features of the monitoring target includes: Determine the weights corresponding to each time step; Performing weighted summation on the second time series features corresponding to each time step based on the weight corresponding to each time step, to obtain the first time series features of the monitoring target; The method of determining the weight corresponding to each time step includes: Calculate the reference value of the anomaly score at this time step; obtaining an actual displacement of the monitoring target based on a displacement sensor provided on the monitoring target, determining a theoretical displacement of the monitoring target based on an ambient temperature of the monitoring target, and determining a displacement error of the monitoring target based on the actual displacement and the theoretical displacement; obtaining actual strain data of the monitored object based on a strain sensor provided on the monitored object, determining theoretical strain information of the monitored object based on a structural mechanics model of the monitored object, and determining a strain error of the monitored object based on the actual strain data and the theoretical strain data; Adjusting a reference value of the anomaly score of the time step based on the displacement error and the strain error to obtain an anomaly score of the time step; The weight corresponding to the time step is determined based on the positive correlation between the anomaly score and the weight of the time step.
4. The AI intelligent visual disaster monitoring and alarm method according to claim 1, characterized in that: The target visible light image and the target infrared image in the image pair corresponding to the acquisition time are respectively subjected to feature extraction, and the extracted features are fused to obtain a fused feature map, including: extracting features of the target visible light image and the target infrared image based on a first convolutional neural network, and fusing the features extracted from the target visible light image and the features extracted from the target infrared image; The first convolutional neural network is trained in the following way: Acquire a plurality of first sample image pairs, each first sample image pair comprising a visible light sample image and a corresponding infrared sample image; Performing sample expansion on the plurality of first sample image pairs to obtain a plurality of second sample image pairs; Based on the multiple first sample image pairs and the multiple second sample image pairs, an initial convolutional neural network is trained to obtain the first convolutional neural network.
5. The AI intelligent visual disaster monitoring and alarm method according to claim 4, characterized in that: The performing sample expansion on the plurality of first sample image pairs to obtain a plurality of second sample image pairs includes: The first sample image pair is sample expanded using a generative adversarial network to obtain the second sample image pair; the loss function of the generative adversarial network is: in, represents the loss function, represents the cross-modal adversarial loss, represents the semantic alignment loss, represents the cycle consistency loss, represents the adversarial feature matching loss, , , , All are weight coefficients; in, Represents the real visible light image judged by the discriminator Belongs to the real data distribution The probability of Represents the generated visible light image obtained by the discriminator authenticity, represents the visible light discriminator; Represents the real visible light-infrared image pair judged by the discriminator Belongs to the real paired data distribution The probability of Represents the real visible light-generated infrared image pair judged by the discriminator The rationality of pairing, represents the cross-modal pairing discriminator; in, represents the semantic encoding of any category c in the visible light image, represents the indicator function, represents the infrared feature extractor, represents the feature mean of any category c in the real infrared image; in, represents the inverse generator, which is used to generate visible light images based on infrared images. represents the expected value of the infrared image y; Among them, the adversarial feature matching loss The calculation is obtained by using the pre-trained second convolutional neural network to extract the visible light image and generate the high-level semantic feature map of the infrared image. Indicates that the feature map is The height of the layer, Indicates that the feature map is The width of the layer, Indicates that the feature map is The number of channels of the layer, Represents the pre-trained CNN network The feature extraction function of the layer.
6. The AI intelligent visual disaster monitoring and alarm method according to claim 5, characterized in that: Also includes: Set the weight of the semantic alignment loss when the number of training rounds of the adversarial network is less than or equal to the set number are respectively greater than the weight of the cross-modal adversarial loss , the weight of cycle consistency loss and the weight of the adversarial feature matching loss ; Set the weight of the cross-modal adversarial loss when the number of training rounds of the adversarial network exceeds the set number are greater than the weight of the semantic alignment loss and the weight of the cycle consistency loss , and the weight of the adversarial feature matching loss are greater than the weight of the semantic alignment loss and the weight of the cycle consistency loss .
7. The AI intelligent visual disaster monitoring and alarm method according to claim 1, characterized in that: The health status of the monitored target includes multiple health status levels, and the AI intelligent visual disaster monitoring and alarm method further includes: If the health status level of the monitored target is a specified health status level, a first alarm signal is output.
8. An AI intelligent visual disaster monitoring and alarm device, characterized in that: include: An image acquisition module is configured to acquire a plurality of continuous visible light images and a plurality of continuous infrared images of a target acquired within a preset time period, wherein each of the visible light images and the infrared images of the target contain the monitored target; each of the continuous visible light images and the continuous infrared images of the target correspond to each other in terms of acquisition time; each acquisition time corresponds to an image pair, and each image pair includes one visible light image and one infrared image of the target; A feature fusion module is used to extract features from the target visible light image and the target infrared image in the image pair corresponding to each acquisition time, and fuse the extracted features to obtain a fused feature map; The state assessment module is used to perform the following operations on multiple fused feature maps through the state assessment network to obtain the health status of the monitoring target: Extracting time series features from the multiple fused feature maps to obtain a first time series feature of the monitoring target; The health status of the monitoring target is evaluated based on the first time series feature; the health status of the monitoring target is used to indicate whether a disaster occurs to the monitoring target.
9. An AI intelligent visual disaster monitoring and alarm system, characterized in that: include: A first image acquisition device is used to acquire a target visible light image containing a monitoring target; The second image acquisition device is used to acquire an infrared image of a target including a monitoring target; An on-site AI analysis and early warning module comprises a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor implements the steps of the method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Embedded landslide real-time early warning method based on spatial-temporal feature decoupling
CN121034058A