Intelligent assessment method for earthquake damage of buildings based on multimodal large model
By constructing a large multimodal earthquake damage assessment model and comprehensively utilizing building earthquake damage images and language descriptions, the problems of insufficient comprehensiveness and low safety of building earthquake damage assessment in existing technologies have been solved, achieving a more accurate and efficient building earthquake damage assessment.
Patent Information
- Application Number
- CN202311278623.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-07
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2043-10-07
AI Technical Summary
Existing building earthquake damage assessment methods rely on images and digital features and cannot provide comprehensive and digital assessments. Traditional methods also have the problems of high professional requirements and low on-site work efficiency.
An intelligent assessment method for earthquake damage to buildings based on a multimodal large model is adopted. By comprehensively utilizing information sources such as multimodal damage images and damage language descriptions, combined with deep learning technology, a multimodal earthquake damage assessment large model is constructed and integrated into the drone or inspection vehicle system to collect and evaluate earthquake damage data in real time.
It achieves a more professional and accurate assessment of the extent of building damage, provides a detailed description of the damage type and extent, improves the scientific nature and accuracy of the assessment, reduces the risks of on-site professionals, and improves the efficiency and safety of the assessment.
Smart Images

Figure CN117292230B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of earthquake engineering, and in particular relates to an intelligent assessment method for earthquake damage to buildings based on a multi-modal large model. Background Art
[0002] When natural disasters such as earthquakes strike, buildings often suffer varying degrees of damage and destruction. Traditional building damage assessment methods rely primarily on on-site observations and empirical judgment by civil engineering professionals. However, these methods require high professional expertise, have low on-site efficiency, and pose uncontrollable risks to personal safety.
[0003] Furthermore, most existing automated earthquake damage assessment technologies based on deep learning methods focus on damage image classification tasks based on convolutional neural network models. For example, based on convolutional neural networks or machine learning models, classification and detection tasks are performed on regional earthquake damage satellite images, damage images of individual buildings, and their image feature parameters. However, the limitation of this type of assessment method is that it relies solely on the use of "damage images" or "digital feature parameter" datasets to classify earthquake damage images or locate damage, and cannot provide other types of information such as text descriptions of damage or historical data. Therefore, methods that rely on images and digital features may be limited by the amount and quality of data and may not provide a comprehensive digital and textual assessment of earthquake damage.
[0004] Therefore, in order to accurately assess the extent of earthquake damage to buildings and provide a scientific basis to guide emergency management and post-disaster reconstruction, it is urgent to develop a building damage assessment system based on a multimodal large model, which can effectively improve the intelligence and automation level of building damage assessment. Summary of the Invention
[0005] The purpose of the present invention is to overcome the problems existing in the prior art and provide an intelligent assessment method for earthquake damage to buildings based on a multimodal large model. By comprehensively utilizing multiple information sources such as multimodal damage images and damage language descriptions, combined with deep learning technology, a more professional and accurate assessment of the degree of damage to buildings can be achieved, which can provide important technical support for earthquake disaster emergency rescue work.
[0006] To achieve the above objectives, the present invention adopts the following technical solution: an intelligent earthquake damage assessment method for buildings based on a multimodal large model, comprising the following steps:
[0007] S1. Collect earthquake damage images of buildings to build a dataset:
[0008] Through web crawling, field surveys, post-disaster images, and survey reports containing damage images from different regions, building types, and earthquake magnitudes, the collected damage images were resized to a uniform size and annotated with text descriptions according to the "Building Earthquake Damage Level Classification Standard" to construct a multimodal earthquake damage assessment dataset.
[0009] S2. Construction of a large multi-modal earthquake damage assessment model:
[0010] S21. A multimodal earthquake damage assessment model is constructed by two image encoders and two text encoders. The inputs of the four encoders are thermal images, RGB earthquake damage images, damage text descriptions, and lidar data. Each encoder uses a self-attention mechanism and forward propagation to gradually extract and encode the features of different modal data. After processing by each encoder, a fixed-length feature representation of the corresponding input data is obtained. The fixed-length feature representation includes the correlation and important features between different modal data, and parameters are shared between two adjacent encoders.
[0011] S22. Optimize the multimodal features based on the minimization objective of the image-text contrast loss function. The image-text contrast loss function is used to measure the similarity or correlation between images and texts. It consists of three parts: image-text matching loss, which is used to measure the degree of match between an image and its corresponding text description; text generation loss, which is used to measure the similarity between the text description generated by the model and the real text description; damage information classification loss, which is used to measure the accuracy between the damage information predicted by the model and the real damage information;
[0012] S23. Construct a multimodal encoder by self-attention mechanism, forward propagation, multi-head self-attention mechanism and multi-head cross-modal attention mechanism, and output a new multimodal vector representation, each vector contains a fusion of visual, textual and position information;
[0013] S3. Training a large multi-modal earthquake damage assessment model:
[0014] Based on the preparation of the multimodal earthquake damage assessment dataset in step S1 and the construction of the multimodal earthquake damage assessment large model in step S2, model training is performed to obtain the model's training weights. The training process includes two stages: pre-training and fine-tuning. In the pre-training stage, a dataset of large-scale building damage images and text descriptions corresponding to the damage phenomena crawled from the Internet is used to train the model's general understanding of image features. In the fine-tuning stage, the more refined dataset selected in step S1 is used to improve the model's accuracy and generalization ability.
[0015] In practical applications, the multimodal earthquake damage assessment model loads pre-trained weights and, after collecting real-time images of building damage at a given site, provides guidance on damage classification, damage level assessment, and post-earthquake repair and reinforcement in various scenarios.
[0016] S4. Multimodal earthquake damage assessment large model algorithm integrated into drones or inspection vehicles:
[0017] Integrate the multimodal earthquake damage assessment model loaded with pre-trained weights in step S3 into a drone or inspection vehicle system. In an actual disaster scenario, the inspection vehicle or drone equipped with the multimodal model will take real-time pictures of the earthquake damage site, record the geographic location of each picture, and perform earthquake damage analysis on the photographed buildings using the multimodal earthquake damage assessment model.
[0018] S5. Assess the damage to buildings within the target disaster area:
[0019] When conducting earthquake damage surveys for regional building clusters, drones or inspection vehicles integrated with the multimodal earthquake damage assessment model will plan their routes in advance and determine the order in which to visit each building in the target area to be scanned. The drones or inspection vehicles will use cameras, thermal imagers, and lidar equipment to collect image, temperature, and geometric dimension data for each building in the disaster area. The collected earthquake damage image data for the buildings in the disaster area will be input into the multimodal earthquake damage assessment model constructed and trained in steps S2 and S3, where visual feature extraction, text encoding, multimodal encoding, and loss function optimization will be performed to assess the extent of damage to each building.
[0020] S6. Generate earthquake damage assessment reports for each building within the target affected area:
[0021] Based on the earthquake damage assessment results of each building obtained by the drone and inspection vehicle system integrated with the multimodal earthquake damage assessment model in step S5, the graph neural network model is used to evaluate the damage of similar buildings in the vicinity of a given geographical location. The target disaster area is divided into grids, each grid is divided into damage levels, and the earthquake economic losses of each grid are calculated. According to the damage level and loss amount indicators, an overall assessment of the disaster area is conducted and compiled into a normative earthquake damage assessment report with geographical location information. The analysis results will be transmitted to the local interface in real time, and the detailed earthquake damage and loss distribution will be displayed on the map.
[0022] In step S1, building types include masonry houses, reinforced concrete frame structures, and wooden houses; building damage levels are divided into basically intact, slightly damaged, moderately damaged, severely damaged, and destroyed; and the damage images are adjusted to the same size of 1024*1024.
[0023] In step S2, the data processing process inside the image encoder is as follows: the damaged image is divided into smaller, non-overlapping blocks through the visual feature extraction end inside the image encoder, a linear projection layer is applied to each block to generate a plane vector, a position code is added to each plane vector so that the self-attention mechanism can distinguish blocks at different positions, and finally the plane vector is used as the query, key and value vector of the attention mechanism, the correlation between each block and other blocks is calculated through the multi-head self-attention mechanism, and a new vector representation is output, each vector containing the visual information and spatial information in the damaged image.
[0024] In step S2, the data processing process inside the text encoder is as follows: the text information description of the building earthquake damage assessment and the lidar data are tokenized through text encoding and converted into a tag sequence, a fixed-size vector representation is generated for each tag, and a position code is added to each vector so that the self-attention mechanism can maintain the context and order of the input text. Finally, these vector representations are used as the query, key, and value vectors of the attention mechanism, and the correlation between each tag and other tags is calculated through the multi-head self-attention mechanism, and a new vector representation is output, each vector containing the text information and position information in the text description and the lidar data.
[0025] In step S2, assuming that a damaged image I and its corresponding text description T are input, the cosine similarity is used to measure the matching degree between the image and the text. The matching degree expression is:
[0026] L match =1-Sim(I, T),
[0027] Where Sim represents the cosine similarity between image and text, L match represents the image-text matching loss;
[0028] Assume that the text description generated by the model is T′, and use cross entropy loss to measure the similarity between the generated text description and the real text description. The similarity expression is:
[0029] L generate =-∑(T*log(T′)+(1-T)*log(1-T′), where T represents the real text description, T′ represents the text description generated by the model, and L generate represents the text generation loss;
[0030] Assume that the prediction result of the model for damage information classification is P, and the actual damage information category is D. The cross entropy loss is used to measure the accuracy between the predicted damage information and the actual damage information. The accuracy expression is:
[0031] L classification=-∑(D*log(P)+(1-D)*log(1-*P),
[0032] Where D represents the one-hot vector of the real damage information, P represents the prediction result of the model, and L classification Indicates the damage information classification loss.
[0033] In step S6, the specific steps of using the graph neural network for evaluation include:
[0034] S61. When adjacent buildings are similar in terms of structural system, geometric dimensions, and dynamic characteristics, the extent of damage to these buildings after experiencing the same earthquake can be determined through building similarity relationships. Therefore, by using a multimodal earthquake damage assessment model to assess damage to representative individual buildings at multiple locations, damage images, text descriptions, and damage information are obtained for each building in each location scenario.
[0035] S62, representing each location scene as a node in a target assessment area graph according to its geographical location, structural system, building age, and geometric dimension attributes, and connecting edges based on similarity or distance between the nodes to construct a graph structure;
[0036] S63. Use the graph structure to disseminate and update information for multiple locations within the region, so that each building node that has not been inspected and assessed can learn about damage information from its neighboring nodes that have received building damage assessment results and update its own damage information.
[0037] S64. Based on the results of the multimodal earthquake damage assessment model and the graph neural network output, an economic loss assessment is conducted on the damage to each building in the target disaster area, and a regional earthquake damage and loss assessment report is generated.
[0038] The beneficial effects of the present invention are:
[0039] 1) The assessment method of the present invention achieves a more professional and accurate assessment of the degree of building damage by comprehensively utilizing multiple information sources such as multimodal damage images and damage language descriptions, combined with deep learning technology. The damage language description can provide an overall and local description of the earthquake damage image, which helps to more accurately classify and locate the type and degree of building damage. It has high scientificity and accuracy, and can provide important technical support for earthquake disaster emergency management and post-disaster rescue work.
[0040] 2) Through information fusion and analysis, the assessment method of this invention provides different sources of information: verbal damage descriptions and images. These complementary sources provide a more comprehensive description and understanding of the earthquake damage. Verbal damage descriptions provide intuitive semantic information, while images provide more specific and fine-grained visual information. Furthermore, by combining data with textual descriptions, natural language processing techniques can be used to perform semantic analysis on the textual information and fuse it with image features. This method can process high-resolution images and multiple data modalities (such as RGB, LiDAR, and thermal imaging) to improve the accuracy of damage assessment. Compared with traditional inspection methods, this method can provide a more detailed and accurate understanding of the nature and extent of damage.
[0041] 3) In the assessment method of the present invention, building earthquake damage images and text descriptions are used as training data sets for training, which can provide richer and more comprehensive information, help to assess the type, extent and location of earthquake damage, and thus achieve a more accurate and comprehensive earthquake damage assessment.
[0042] 4) In the assessment method of the present invention, by integrating a large multimodal model into drones or inspection vehicles, drones and inspection vehicles can effectively cover large areas, and can complete building earthquake damage assessment and decision-making faster and more widely; drones or inspection vehicles can be used to collect building earthquake damage image data in difficult terrain while quickly completing damage identification and assessment, which can reduce the need for manual inspections in dangerous areas, thereby minimizing the risk to the lives of on-site professionals.
[0043] 5) The assessment method of the present invention can realize real-time processing of building earthquake damage data and rapid acquisition of analysis results, thereby improving the timeliness of disaster relief work; the method can be seamlessly integrated with GIS tools to provide visualization and spatial analysis functions, supporting effective decision-making and resource prioritization in emergency situations; compared with traditional methods, this method can significantly improve the speed, accuracy and safety of earthquake loss assessment, providing targeted guidance for emergency rescue and repair work. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 Schematic diagram of the process of the evaluation method of the present invention;
[0045] Figure 2 Schematic diagram of the principle of the multi-modal earthquake damage assessment model constructed in the assessment method of the present invention;
[0046] Figure 3 This is a diagram showing the damage assessment and positioning results obtained in an embodiment of the present invention. DETAILED DESCRIPTION
[0047] The present invention will be further explained below with reference to the accompanying drawings and specific embodiments.
[0048] Example: Figure 1-3 As shown, the present invention provides a method for intelligent assessment of building earthquake damage based on a multimodal large model, comprising the following steps:
[0049] S1. Collect earthquake damage images of buildings to build a dataset:
[0050] Damage images from different regions, building types, and earthquake magnitudes were collected through web crawlers, field surveys, post-disaster images, and expert reports. The collected damage images were resized to a uniform size of 1024*1024, and all collected damage images were annotated with text descriptions according to the "Standard for Classification of Earthquake Damage Levels of Buildings" to construct a multimodal earthquake damage assessment dataset.
[0051] Building types include masonry houses, reinforced concrete frame structures and wooden houses; building damage levels are divided into basically intact, slightly damaged, moderately damaged, severely damaged and destroyed.
[0052] The specific labeling method and specifications are as follows: {"Image Number": "001", "Text Description": Building type (masonry building, reinforced concrete frame structure, wooden structure, etc.), Building damage level classification (IV, basically intact, slightly damaged, moderately damaged, severely damaged, and destroyed), "Damage Location": The spatial location of the damage in the image, "Repair Suggestions": Detailed and guiding repair suggestions are given for the damage in each image.}
[0053] S2. Construction of a large multi-modal earthquake damage assessment model:
[0054] The multi-modal earthquake damage assessment model built is as follows Figure 2 As shown in the figure, "multimodal" means that the model can process and understand multiple types of data inputs, such as text, images, sounds, etc., so that it can learn and reason across multiple perceptual modes; "large model" refers to a deep learning model with a large number of parameters, which usually requires a large amount of data for training and has strong representation learning and generalization capabilities.
[0055] The specific steps are as follows: S21, a multimodal earthquake damage assessment model is constructed by two image encoders and two text encoders. The corresponding inputs of the four encoders are thermal imaging images, RGB earthquake damage pictures, damage text information descriptions and lidar data. Each encoder uses self-attention mechanism and forward propagation to gradually extract and encode the features of different modal data, and obtains the fixed-length feature representation of the corresponding input data after processing by each encoder. The fixed-length feature representation includes the correlation and important features between different modal data, and parameters are shared between two adjacent encoders.
[0056] For the image encoder, the data processing process inside the image encoder is as follows: the damaged image is divided into smaller, non-overlapping blocks, such as a 3x3 grid, through the visual feature extraction end inside the image encoder to obtain 9 blocks; a linear projection layer is applied to each block to generate a plane vector, such as a vector of length 64; position encoding is added to each plane vector so that the self-attention mechanism can distinguish blocks at different positions; finally, the plane vector is used as the query, key and value vectors of the attention mechanism, and the correlation between each block and other blocks is calculated through the multi-head self-attention mechanism, and a new vector representation is output. For example, 8 heads are used, each head outputs a vector of length 8, and then these vectors are spliced together to obtain a vector of length 64 as the final output. In this way, 9 new vector representations are obtained, each vector contains visual information and spatial information in the damaged image.
[0057] For the text encoder, the data processing process inside the text encoder is as follows: the text information description of the building earthquake damage assessment and the lidar data are tokenized through text encoding and converted into a token sequence; for example, byte pair encoding (BPE) is used to divide the text description into subword units and the lidar data is converted into digital tokens; a fixed-size vector representation is generated for each token, and position encoding is added to each vector so that the self-attention mechanism can capture the context and order of the input text; finally, these vector representations are used as the query, key and value vectors of the attention mechanism, and the correlation between each token and other tokens is calculated through the multi-head self-attention mechanism, and a new vector representation is output, each vector containing the text information and position information in the text description and lidar data.
[0058] S22. Optimize multimodal features according to the minimization objective of the image-text contrast loss function. The image-text contrast loss function is used to measure the similarity or correlation between images and texts. It consists of three parts: image-text matching loss, which is used to measure the degree of matching between an image and its corresponding text description; text generation loss, which is used to measure the similarity between the text description generated by the model and the real text description; damage information classification loss, which is used to measure the accuracy between the damage information predicted by the model and the real damage information.
[0059] Assume that we input a damaged image I and its corresponding text description T, and use cosine similarity to measure the degree of match between the image and the text. The expression of the matching degree is:
[0060] L match =1-Sim(I, T),
[0061] Where Sim represents the cosine similarity between image and text, L match represents the image-text matching loss;
[0062] Assume that the text description generated by the model is T′, and use cross entropy loss to measure the similarity between the generated text description and the real text description. The similarity expression is:
[0063] L generate =-∑(T*log(T′)+(1-T)*log(1-T′),
[0064] In the formula, T represents the real text description, T′ represents the text description generated by the model, and L generate represents the text generation loss;
[0065] Assume that the prediction result of the model for damage information classification is P, and the actual damage information category is D. The cross entropy loss is used to measure the accuracy between the predicted damage information and the actual damage information. The accuracy expression is:
[0066] L classification =-∑(D*log(P)+(1-D)*log(1-P),
[0067] Where D represents the one-hot vector of the real damage information, P represents the prediction result of the model, and L classification Indicates the damage information classification loss.
[0068] The image-text contrast loss function is used to measure the similarity or correlation between images and text, thereby prompting the model to learn better image-text matching and alignment features. After obtaining excellent multimodal alignment features, it is necessary to encode their multimodal features.
[0069] S23. A multimodal encoder is constructed by self-attention mechanism, forward propagation, multi-head self-attention mechanism and multi-head cross-modal attention mechanism, and outputs a new multimodal vector representation, each of which contains a fusion of visual, textual and position information.
[0070] The visual feature extraction and text encoder communicate and share parameters in a multimodal encoder, allowing them to learn from each other and adjust their parameters. Two sublayers are then added to the multimodal encoder: a multi-head self-attention mechanism to further enhance the representation within each modality; and a multi-head cross-modal attention mechanism to compute correlations between different modalities and output a new multimodal vector representation, each of which incorporates a fusion of visual, textual, and location information. The multimodal encoder maps data from different modalities into a shared low-dimensional representation, transforming the multimodal data into a unified representation. By integrating multiple modalities, the model can effectively leverage the complementary information provided by LiDAR, thermal imaging, and visual RGB data, enabling it to provide more accurate and comprehensive damage assessments during earthquake response and recovery operations.
[0071] S3. Training a large multi-modal earthquake damage assessment model:
[0072] Based on the preparation of the multimodal earthquake damage assessment dataset in step S1 and the construction of the multimodal earthquake damage assessment large model in step S2, model training is performed to obtain the model's training weights. The training process includes two stages: pre-training and fine-tuning. In the pre-training stage, a dataset of large-scale building damage images and text descriptions corresponding to the damage phenomena crawled from the Internet is used to train the model's general understanding of image features. In the fine-tuning stage, the more refined dataset selected in step S1 is used to improve the model's accuracy and generalization ability.
[0073] In practical applications, the multimodal earthquake damage assessment model loads pre-trained weights and, after collecting real-time images of building damage at a given site, provides guidance on damage classification, damage level assessment, and post-earthquake repair and reinforcement in various scenarios.
[0074] S4. Multimodal earthquake damage assessment large model algorithm integrated into drones or inspection vehicles:
[0075] The multimodal earthquake damage assessment model loaded with pre-trained weights in step S3 is integrated into the drone or inspection vehicle system. In the actual disaster scenario, the inspection vehicle or drone equipment integrated with the multimodal large model will take real-time pictures of the earthquake damage site, record the geographic location information of each picture, and perform earthquake damage analysis on the photographed buildings through the multimodal earthquake damage assessment large model.
[0076] S5. Assess the damage to buildings within the affected target area:
[0077] When conducting earthquake damage surveys for regional building complexes, drones or inspection vehicles integrated with a large multimodal earthquake damage assessment model will plan routes in advance and determine the order of visits to each building in the target area to be scanned. Specifically, the disaster area is divided into several grids, and priority scores are assigned based on the attributes of the grids (such as population density, building type, terrain, etc.) to reflect the urgency of their assessment. Specifically, each grid can be represented by a two-tuple (x, y), where x represents the grid number and y represents the priority score. The access order is represented by a list S, in which each element corresponds to the number of a grid. The grid with a higher priority score is positioned higher in the list.
[0078] Assume there are n grids, and the grid priority scores are stored in an n-dimensional vector P, where P[i] represents the priority score of the i-th grid. Sorting by priority score yields a permutation P', where P'[i] represents the number of the i-th grid after the permutation. The access order list S can be defined as follows:
[0079] S=[P′[1],P′[2],...,P′[n]],
[0080] Through the formula, priority scores are assigned according to grid attributes and the access order is calculated.
[0081] Use cameras, thermal imagers, and lidar equipment mounted on drones or patrol vehicles to collect images, temperature, and geometric dimension data of the disaster area. Specifically, have drones and patrol vehicles fly along a planned route and take two types of images on each grid: one is a regular RGB image, used to show the appearance and structure of the building; the other is a thermal imaging image, used to show the temperature distribution and thermal anomalies of the building. At the same time, have drones and patrol vehicles use lidar to scan each grid and record the distance data of each point to show the shape and height of the building.
[0082] The collected multimodal earthquake damage data is input into the multimodal earthquake damage assessment model constructed in step S2, and visual feature extraction, text encoding, multimodal encoding and loss function optimization operations are performed to obtain the damage information of each building damage image in each grid.
[0083] S6. Generate earthquake damage assessment reports for each building within the target affected area:
[0084] Based on the earthquake damage assessment results of each building obtained by the drone and inspection vehicle system integrated with the multimodal earthquake damage assessment model in step S5, the graph neural network model is used to evaluate the damage of similar buildings in the vicinity of a given geographical location. The target disaster area is divided into grids, each grid is divided into damage levels, and the earthquake economic losses of each grid are calculated. According to the damage level and loss amount indicators, an overall assessment of the disaster area is conducted and compiled into a normative earthquake damage assessment report with geographical location information. The analysis results will be transmitted to the local interface in real time, and the detailed earthquake damage and loss distribution will be displayed on the map.
[0085] like Figure 3 As shown in the figure, after using a multimodal large model to assess the damage of four locations A, B, C, and D, a graph neural network is used to assess the damage of similar buildings in the surrounding area. The specific steps are as follows:
[0086] S61. When adjacent buildings are similar in terms of structural system, geometric dimensions, and dynamic characteristics, the extent of damage to adjacent buildings after experiencing the same earthquake disaster can be determined through building similarity relationships. Therefore, after conducting damage assessments for the four location scenarios A, B, C, and D using the multimodal damage assessment model, damage images, text descriptions, and damage information for each location scenario were obtained.
[0087] S62. Represent each location scene as a node in a target assessment area graph based on its attributes (e.g., geographic location, structural system, building age, and geometric dimensions), and connect edges based on similarity or distance between nodes to construct a graph structure.
[0088] S63. Use graph neural networks to propagate and update information on the graph structure, so that each building node that has not been inspected and assessed can learn about the damage situation from its neighboring nodes that have received building damage assessment results and update its own damage information.
[0089] S64. Based on the results of the multimodal earthquake damage assessment model and the graph neural network output, an economic loss assessment is conducted on the damage to each building in the target disaster area, and a regional earthquake damage and loss assessment report is generated.
[0090] Through the above steps, graph neural networks can be used to assess the damage of similar buildings in neighboring areas, thereby improving the efficiency and accuracy of earthquake damage assessment.
[0091] Furthermore, according to the aforementioned implementation plan, as images of earthquake damage are continuously collected, the model automatically learns and optimizes, thereby improving the accuracy and efficiency of earthquake damage assessments. In summary, this specific implementation involves collecting earthquake damage images in real time, preprocessing them, and training a large multimodal model. The model then outputs damage classification, damage level assessment, and damage geolocation information, all while undergoing self-learning and optimization.
[0092] Unlike the current method of using only building earthquake damage pictures or text to evaluate earthquake damage, the present invention proposes an intelligent building earthquake damage assessment system based on a multimodal large model. It uses a variety of pictures and text descriptions containing damage information as training data, aiming to solve the problems of low on-site work efficiency of professionals and difficulties in data collection and quantitative analysis in traditional earthquake damage assessment methods.
[0093] The present invention collects images of building damage caused by earthquakes, including crawling images from social media platforms, and inputs the images into a multimodal large model for identification. The multimodal large model is integrated into an unmanned aerial vehicle system or patrol vehicle to take images of building damage in real time and record geographic location information. This assessment method, which associates damage information with geographic location information, can record regional damage information of the area to be assessed in real time. In addition, non-professionals can obtain professional judgments on the degree of damage to their own homes by taking photos and inputting them into the large model. Professionals can use the large model to classify earthquake damage and conduct quantitative analysis of the classification results, such as the crack size and spalling area of concrete components.
[0094] The present invention simplifies the building earthquake damage assessment process and improves the accuracy and efficiency of the assessment; it provides non-professionals with a professional assessment method for damage to their own homes without the need for expert participation; and it provides professionals with an integrated solution for classification and quantitative analysis; it can adapt to different scenarios and needs and can be continuously updated and iterated as data increases.
[0095] Based on the earthquake damage assessment results of each building, the present invention can use graph neural networks to provide more accurate and comprehensive results for damage deduction of buildings in the target disaster area; by accurately capturing structural similarities and spatial relationships, the integrated graph neural network can significantly improve the prediction accuracy of earthquake damage distribution in building groups.
[0096] The above description is only used to illustrate the technical solution of the present invention and is not intended to limit it. Other modifications or equivalent substitutions made to the technical solution of the present invention by ordinary technicians in this field should be included in the scope of the claims of the present invention as long as they do not depart from the spirit and scope of the technical solution of the present invention.
Claims
1. An intelligent building earthquake damage assessment method based on a multimodal large model, characterized by: The following steps are involved: S1. Collect earthquake damage images of buildings to build a dataset: Through web crawling, field surveys, post-disaster images, and expert reports, we collected damage images from different regions, building types, and earthquake magnitudes. We resized the images to a uniform size and annotated them with text descriptions according to the "Building Earthquake Damage Classification Standard" to construct a multimodal earthquake damage assessment dataset. S2. Construction of a large multi-modal earthquake damage assessment model: S21. A multimodal earthquake damage assessment model is constructed by two image encoders and two text encoders. The inputs of the four encoders are thermal images, RGB earthquake damage images, damage text descriptions, and lidar data. Each encoder uses a self-attention mechanism and forward propagation to gradually extract and encode the features of different modal data. After processing by each encoder, a fixed-length feature representation of the corresponding input data is obtained. The fixed-length feature representation includes the correlation and important features between different modal data, and parameters are shared between two adjacent encoders. S22. Optimize the multimodal features based on the minimization objective of the image-text contrast loss function. The image-text contrast loss function is used to measure the similarity or correlation between images and texts. It consists of three parts: image-text matching loss, which is used to measure the degree of match between an image and its corresponding text description; text generation loss, which is used to measure the similarity between the text description generated by the model and the real text description; damage information classification loss, which is used to measure the accuracy between the damage information predicted by the model and the real damage information; S23. Construct a multimodal encoder by self-attention mechanism, forward propagation, multi-head self-attention mechanism and multi-head cross-modal attention mechanism, and output a new multimodal vector representation, each vector contains a fusion of visual, textual and position information; S3. Training a large multi-modal earthquake damage assessment model: Based on the preparation of the multimodal earthquake damage assessment dataset in step S1 and the construction of the multimodal earthquake damage assessment large model in step S2, model training is performed to obtain the model's training weights. The training process includes two stages: pre-training and fine-tuning. In the pre-training stage, a dataset of large-scale building damage images and text descriptions corresponding to the damage phenomena crawled from the Internet is used to train the model's general understanding of image features. In the fine-tuning stage, the more refined dataset selected in step S1 is used to improve the model's accuracy and generalization ability. In practical applications, the multimodal earthquake damage assessment model loads pre-trained weights and, after collecting real-time images of building damage at a given site, provides guidance on damage classification, damage level assessment, and post-earthquake repair and reinforcement in various scenarios. S4. Multimodal earthquake damage assessment large model algorithm integrated into drones or inspection vehicles: Integrate the multimodal earthquake damage assessment model loaded with pre-trained weights in step S3 into a drone or inspection vehicle system. In an actual disaster scenario, the inspection vehicle or drone equipped with the multimodal model takes real-time pictures of the earthquake damage site, records the geographic location information of each picture, and uses the multimodal earthquake damage assessment model to perform earthquake damage analysis on the photographed buildings. S5. Assess the damage to buildings within the target disaster area: When conducting earthquake damage surveys for regional building clusters, drones or inspection vehicles integrated with the multimodal earthquake damage assessment model will plan their routes in advance and determine the order in which to visit each building in the target area to be scanned. The drones or inspection vehicles will use cameras, thermal imagers, and lidar equipment to collect image, temperature, and geometric dimension data for each building in the disaster area. The collected earthquake damage image data for the buildings in the disaster area will be input into the multimodal earthquake damage assessment model constructed and trained in steps S2 and S3, where visual feature extraction, text encoding, multimodal encoding, and loss function optimization will be performed to assess the extent of damage to each building. S6. Generate earthquake damage assessment reports for each building within the target affected area: Based on the earthquake damage assessment results of each building obtained by the drone and inspection vehicle system integrated with the multimodal earthquake damage assessment model in step S5, the graph neural network model is used to evaluate the damage of similar buildings in the vicinity of a given geographical location. The target disaster area is divided into grids, each grid is divided into damage levels, and the earthquake economic losses of each grid are calculated. According to the damage level and loss amount indicators, an overall assessment of the disaster area is conducted and compiled into a normative earthquake damage assessment report with geographical location information. The analysis results will be transmitted to the local interface in real time, and the detailed earthquake damage and loss distribution will be displayed on the map.
2. The intelligent building earthquake damage assessment method based on a multimodal large model according to claim 1 is characterized by: In step S1, the building types include masonry houses, reinforced concrete frame structures and wooden houses; The building damage levels are divided into basically intact, slightly damaged, moderately damaged, severely damaged and destroyed; the damage images are adjusted to the same size of 1024*1024.
3. The intelligent building earthquake damage assessment method based on a multimodal large model according to claim 1 is characterized by: In step S2, the data processing process inside the image encoder is as follows: the damaged image is divided into smaller, non-overlapping blocks through the visual feature extraction end inside the image encoder, a linear projection layer is applied to each block to generate a plane vector, a position code is added to each plane vector so that the self-attention mechanism can distinguish blocks at different positions, and finally the plane vector is used as the query, key and value vector of the attention mechanism, the correlation between each block and other blocks is calculated through the multi-head self-attention mechanism, and a new vector representation is output, each vector containing the visual information and spatial information in the damaged image.
4. The intelligent building earthquake damage assessment method based on a multimodal large model according to claim 1 is characterized by: In step S2, the data processing process inside the text encoder is as follows: the text information description of the building earthquake damage assessment and the lidar data are tokenized through text encoding and converted into a tag sequence, a fixed-size vector representation is generated for each tag, and a position code is added to each vector so that the self-attention mechanism can maintain the context and order of the input text. Finally, these vector representations are used as the query, key, and value vectors of the attention mechanism, and the correlation between each tag and other tags is calculated through the multi-head self-attention mechanism, and a new vector representation is output, each vector containing the text information and position information in the text description and the lidar data.
5. The intelligent building earthquake damage assessment method based on a multimodal large model according to claim 1 is characterized by: In step S2, assuming that a damaged image I and its corresponding text description T are input, the cosine similarity is used to measure the matching degree between the image and the text. The matching degree expression is: L match =1-Sim(I,T), Where Sim represents the cosine similarity between image and text, L match represents the image-text matching loss; Assume that the text description generated by the model is T′, and use cross entropy loss to measure the similarity between the generated text description and the real text description. The similarity expression is: L generate =-∑[T*log(T')+(1-T)*log(1-T')], In the formula, T represents the real text description, T′ represents the text description generated by the model, and L generate represents the text generation loss; Assume that the prediction result of the model for damage information classification is P, and the one-hot vector of the true damage information is D. The cross entropy loss is used to measure the accuracy between the predicted damage information and the true damage information. The accuracy expression is: L classification =-∑[D*log(P)+(1-D)*log(1-P)], Where D represents the one-hot vector of the real damage information, P represents the prediction result of the model, and L classification Indicates the damage information classification loss.
6. The intelligent building earthquake damage assessment method based on a multimodal large model according to claim 1 is characterized by: In step S6, the specific steps of using the graph neural network for evaluation include: S61. When adjacent buildings are similar in terms of structural system, geometric dimensions, and dynamic characteristics, the extent of damage to these buildings after experiencing the same earthquake can be determined through building similarity relationships. Therefore, by using a multimodal earthquake damage assessment model to assess damage to representative individual buildings at multiple locations, damage images, text descriptions, and damage information are obtained for each building in each location scenario. S62, representing each location scene as a node in a target assessment area graph according to its geographical location, structural system, building age, and geometric dimension attributes, and connecting edges based on similarity or distance between the nodes to construct a graph structure; S63. Use the graph structure to disseminate and update information for multiple locations within the region, so that each building node that has not been inspected and assessed can learn about damage information from its neighboring nodes that have received building damage assessment results and update its own damage information. S64. Based on the results of the multimodal earthquake damage assessment model and the graph neural network output, an economic loss assessment is conducted on the damage to each building in the target disaster area, and a regional earthquake damage and loss assessment report is generated.
Citation Information
Patent Citations
Lycium barbarum insect pest recognition method based on image-text multi-modal feature fusion
CN116563707A
Aero-engine fault diagnosis method and system based on multi-modal deep learning
CN116842423A