Damage assessment method and device, equipment and storage medium
By extracting target object and area information from remote sensing images and generating structured reports using a knowledge base, the inefficiencies and adaptability of existing damage assessment methods are addressed, achieving efficient and accurate damage assessment.
Patent Information
- Application Number
- CN202511495171.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-17
- Publication Date
- 2026-02-13
AI Technical Summary
Existing damage assessment methods rely on manual interpretation, which is inefficient and difficult to meet the needs of large-scale image processing. Furthermore, existing automated models lack generalization ability and scene adaptability when dealing with complex damage types and diverse ground targets.
The system extracts the attribute information of the target object and the global description information of the region from remote sensing images, constructs the input content by combining it with a preset knowledge base, and generates a structured damage assessment report using the target model. Information is extracted and fused through a fine-grained target perception model and a global description model, and a professional assessment report is generated by calling an external knowledge base using RAG technology.
It enables efficient and accurate damage assessment analysis, applicable to various scenarios such as natural disasters, emergencies, and battlefield situations, thus improving the adaptability and accuracy of damage assessment.
Smart Images

Figure CN121525834A_ABST
Abstract
Description
Technical Field
[0001] This application relates to, but is not limited to, the field of image processing technology, and in particular to a damage assessment method, apparatus, device, and storage medium. Background Technology
[0002] Supported by modern remote sensing technology, satellites and drones can acquire high-resolution image data, which is widely used in disaster assessment, environmental monitoring, and security defense. Image damage assessment, as a key component, aims to determine the extent of damage to surface targets by analyzing changes in their condition, providing a basis for emergency response and decision-making. Current damage assessment methods mostly rely on manual visual interpretation or automated processing based on a single model. While manual interpretation is accurate, it is inefficient and struggles to meet the demands of large-scale image processing; existing automated models, on the other hand, are limited by their generalization ability and scene adaptability, especially when faced with complex damage types and diverse ground features. Summary of the Invention
[0003] In view of this, embodiments of this application provide at least one damage assessment method, apparatus, device, and storage medium.
[0004] The technical solution of this application embodiment is implemented as follows: In a first aspect, embodiments of this application provide a damage assessment method, comprising: extracting global descriptive information of the area to be assessed and attribute information of each target object in the area to be assessed from a remote sensing image of the area to be assessed, wherein the attribute information includes the category, location, and preliminary damage rating of the target object; the global descriptive information includes at least the attribute characteristics of each type of land parcel in the area to be assessed; constructing first input content for guiding a target model based on the attribute information, the global descriptive information, and a preset knowledge base; and generating a global damage assessment report for the area to be assessed based on the first input content and the target model.
[0005] Secondly, embodiments of this application provide a damage assessment device, comprising: an extraction module, configured to extract global descriptive information of the area to be assessed and attribute information of each target object in the area to be assessed from a remote sensing image of the area to be assessed, wherein the attribute information includes the category, location, and preliminary damage rating of the target object; the global descriptive information includes at least the attribute characteristics of each type of land parcel in the area to be assessed; a construction module, configured to construct first input content for guiding a target model based on the attribute information, the global descriptive information, and a preset knowledge base; and a generation module, configured to generate a global damage assessment report for the area to be assessed based on the first input content and the target model.
[0006] Thirdly, embodiments of this application provide a computer device, including a memory and a processor, wherein the memory stores a computer program that can run on the processor, and the processor executes the program to implement some or all of the steps in the above-described method.
[0007] Fourthly, embodiments of this application provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements some or all of the steps in the above-described method.
[0008] Fifthly, embodiments of this application provide a computer program product, including a computer program or instructions, which, when executed by a processor, implement some or all of the steps in the above-described method.
[0009] Technical Effects: By extracting attribute information of target objects and global description information of regions from remote sensing images, and constructing the first input content in conjunction with a preset knowledge base, and then using the target model to generate a structured damage assessment report, efficient and accurate analysis of damage is achieved. The damage assessment method provided in this application overcomes the shortcomings of traditional methods in terms of flexibility, generalization ability, and professionalism. It is applicable to various application scenarios such as natural disasters, emergencies, and battlefield situations, improving the adaptability and accuracy of damage assessment in various application scenarios.
[0010] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and are not intended to limit the technical solutions of this application. Attached Figure Description
[0011] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with this application and, together with the specification, serve to explain the technical solutions of this application.
[0012] Figure 1 A schematic diagram illustrating the implementation process of a damage assessment method provided in this application embodiment; Figure 2 A schematic diagram illustrating the implementation process of a damage assessment method provided in this application embodiment; Figure 3 A schematic diagram illustrating the implementation process of a damage assessment method provided in this application embodiment; Figure 4 A schematic diagram illustrating the implementation process of a damage assessment method provided in this application embodiment; Figure 5 A schematic diagram illustrating the implementation process of a damage assessment method provided in this application embodiment; Figure 6 A schematic diagram illustrating the implementation process of a damage assessment method provided in this application embodiment; Figure 7A schematic diagram illustrating the implementation process of a damage assessment method provided in this application embodiment; Figure 8 A schematic diagram illustrating the implementation process of a multimodal large-model scene target perception method for remote sensing damage assessment provided in this application embodiment; Figure 9 A schematic diagram of a damage assessment framework provided for an embodiment of this application; Figure 10 This is a schematic diagram of the composition structure of a damage assessment device provided in an embodiment of this application; Figure 11 This is a schematic diagram of the hardware entity of a computer device provided in an embodiment of this application. Detailed Implementation
[0013] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application are further described in detail below with reference to the accompanying drawings and embodiments. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0014] In the following description, references to "some embodiments" refer to a subset of all possible embodiments. It is understood that "some embodiments" may be the same or different subsets of all possible embodiments and may be combined with each other without conflict. The terms "first / second / third" are used merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first / second / third" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.
[0015] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains. The terminology used herein is for descriptive purposes only and is not intended to be limiting of this application.
[0016] This application provides a damage assessment method, which can be executed by a processor of a computer device. The computer device can refer to a server, laptop computer, tablet computer, desktop computer, smart TV, set-top box, mobile device (such as a mobile phone, portable video player, personal digital assistant, dedicated messaging device, portable gaming device), or other device with data processing capabilities.
[0017] Figure 1 This is a schematic diagram illustrating the implementation flow of a damage assessment method provided in an embodiment of this application. This method can be executed by the processor of a computer device. Figure 1As shown, the method includes the following steps S101 to S103, combining... Figure 1 The steps are explained below.
[0018] Step S101: Extract global description information of the region to be evaluated and attribute information of each target object in the region to be evaluated from the remote sensing image of the region to be evaluated.
[0019] The attribute information includes the target object's category, location, and preliminary damage rating.
[0020] The global description information includes at least the attribute characteristics of each type of land parcel in the area to be evaluated.
[0021] In some embodiments, the area to be assessed may be an area divided according to the assessment task. For example, if the assessment task is to assess the damage caused by an earthquake disaster, then the area involved in the earthquake is taken as the area to be assessed.
[0022] In some embodiments, remote sensing images refer to surface image data acquired by satellites, drones, or other remote sensing platforms, and include geographic information and target features of the area to be evaluated.
[0023] In some embodiments, remote sensing images can be image data of different resolutions, such as images with resolutions from 1024×1024 to 4096×4096.
[0024] In some embodiments, the target object can be an independent individual or a detection object in the area to be evaluated. For example, if the area to be evaluated is an airport area, the target object can be an aircraft, a warehouse, or other similar object in the airport.
[0025] In some embodiments, the attribute information of a target object refers to the structured information of the target object extracted from the remote sensing image. The attribute information includes the category of each target object (such as building, road, bridge, etc.), location coordinates (x1, y1, x2, y2), and preliminary damage rating (intact, slightly damaged, severely damaged, completely damaged), used to describe the state of a single target. For example, after a building is detected, its category will be automatically labeled as an industrial plant, its location coordinates will be labeled as (100, 200, 300, 400), and the damage level of the building will be determined to be slightly damaged based on the image content.
[0026] Among them, the attribute information can be extracted based on the fine-grained target perception big model. The fine-grained target perception big model can identify target objects of different scales and classify and judge the degree of damage of the target objects. The structure of the attribute information can be: "objects": ["object_1": {"class": "class_1", "Degree of damage": "mild", "location": (x1, y1, x2, y2)}.
[0027] In some embodiments, the global description information of the area to be assessed may be a structured text description of the overall environment of the area to be assessed, including macro information such as terrain features, functional zoning (such as industrial areas and residential areas), and building cluster distribution patterns, to provide contextual support for damage assessment.
[0028] In some embodiments, to achieve a systematic analysis of damage, it is first necessary to extract two levels of information from remote sensing images: one is a macroscopic description of the overall area, such as global descriptive information; the other is the detailed characteristics of each specific target object within the area to be evaluated, such as attribute information. Global descriptive information is used to provide scene context to understand the environmental factors that caused the damage, while attribute information is used to accurately identify and locate damaged target objects and provide a preliminary damage rating based on attribute information.
[0029] In some embodiments, the attribute characteristics of each type of land parcel refer to a comprehensive description of different types of land use in the assessment area, covering both natural geographical elements and artificial construction elements. For example, geographical environmental characteristics include topography (such as plains, mountains, and hills), water systems (such as rivers, lakes, and oceans), and vegetation cover type and proportion (such as forest coverage, grassland area, etc.); regional functional attributes refer to land use types dominated by human activities, including industrial zones, farmland, residential areas, commercial zones, transportation hubs, etc., and assessors record the distribution range and proportion of regional functional attributes; building cluster distribution patterns are used to describe the building agglomeration status in the built-up area, including building density (high / medium / low), arrangement (regular grid / irregular dispersion), and height level (low-rise residential buildings / high-rise buildings), etc.
[0030] In some embodiments, global descriptive information of the area to be evaluated can be generated by a global descriptive model. It is understood that user-input prompts and remote sensing images are input into the global descriptive model to generate a scene-level text description containing macroscopic information such as geographic environmental features, regional functional attributes, and building cluster distribution patterns.
[0031] In some embodiments, before generating global description information, it is also necessary to perform multimodal fusion of remote sensing images and user-input prompts to generate global description information from the fused content.
[0032] Step S102: Construct the first input content to guide the target model based on the attribute information, the global description information, and the preset knowledge base.
[0033] In some embodiments, the first input content is used to drive the target model to generate a damage assessment report.
[0034] In some embodiments, the first input content includes attribute information and global description information extracted from remote sensing images, as well as professional evaluation standards and historical cases matched from an external preset knowledge base.
[0035] In some embodiments, the preset knowledge base is a structured database that stores various damage assessment specifications, engineering standards, historical event records, and other content, which is indexed and retrieved in vector form.
[0036] In some embodiments, firstly, based on the attribute information and the global description information, relevant damage assessment specifications, engineering standards, historical event records, and other content are matched from a preset knowledge base. Then, based on the attribute information, the global description information, and the matched damage assessment specifications, engineering standards, historical event records, etc., the first input content is constructed.
[0037] For example, in a certain assessment task, the target model may need to refer to the damage assessment standards for industrial plants in the "Code for Safety Assessment of Civil Buildings" or cite historical assessment reports of similar events as a basis for comparison. This combines the damage assessment standards in the "Code for Safety Assessment of Civil Buildings" and the professional knowledge contained in historical assessment reports of similar events with the structured information of the current image to generate the first input content.
[0038] In some embodiments, the target model can be a pre-trained large language model, such as a large language model using Retrieval Augmented Generation (RAG) technology, which can receive structured input while calling relevant information from an external pre-set knowledge base to generate a damage assessment report that conforms to professional standards.
[0039] For example, if the attribute information and global description information relate to the damage status of a certain type of industrial plant, the target model retrieves relevant building assessment standards from the preset knowledge base and incorporates the relevant building assessment standards into the generated report, thereby improving the accuracy and authority of the report.
[0040] Step S103: Based on the first input content and the target model, generate a global damage assessment report for the area to be assessed.
[0041] In some embodiments, after receiving the first input, the target model can generate a damage assessment report that conforms to professional assessment standards by calling relevant information from a preset knowledge base. The target model can not only output the damage level of the target object, but also comprehensively analyze the damage situation of the entire area and generate a structured text report.
[0042] For example, for an urban area that has been attacked, the target model can determine which industrial plants have been severely damaged and which roads have been blocked based on the attribute information of the target objects. It can also assess the overall connectivity loss of the transportation system by combining regional functional attributes and building distribution patterns. Finally, the model will generate a formal assessment report containing damage level, impact range, and repair recommendations for decision-makers to reference.
[0043] In some embodiments, the target model, by introducing the RAG mechanism, achieves natural language-driven damage feature description and customized evaluation dimensions, enabling flexible evaluation across scenarios and tasks without retraining. The method of achieving natural language-driven damage feature description and customized evaluation dimensions by introducing the RAG mechanism significantly improves the efficiency and accuracy of damage assessment, while ensuring the professionalism and consistency of the evaluation results.
[0044] In this embodiment, by extracting the attribute information of the target object and the global description information of the region from remote sensing images, and constructing the first input content in conjunction with a preset knowledge base, and then using the target model to generate a structured damage assessment report, efficient and accurate analysis of the damage situation is achieved. The damage assessment method provided in this embodiment overcomes the shortcomings of traditional methods in terms of flexibility, generalization ability, and professionalism, and is applicable to various application scenarios such as natural disasters, emergencies, and battlefield situations.
[0045] Figure 2 This is a schematic diagram illustrating the implementation flow of a damage assessment method provided in an embodiment of this application. This method can be executed by the processor of a computer device. Based on Figure 1 The preset knowledge base includes multiple query vectors, and each query vector corresponds to at least one text content of a damage assessment standard; Figure 1 Step S102 can be updated to steps S201 to S203, combining... Figure 2 The steps shown are explained.
[0046] Step S201: Generate second input content based on the attribute information and the global description information.
[0047] In some embodiments, attribute information is local information of the area to be evaluated, and global description information is global information of the area to be evaluated. That is, it is necessary to integrate the micro-target objects and macro-plots of the area to be evaluated in order to avoid the separation of the two types of information.
[0048] In some embodiments, attribute information representing local information of the region to be evaluated and global description information representing global information of the region to be evaluated are mapped to the same spatial coordinate system to establish a correlation between attribute information and global description information, thereby fusing attribute information and global description information into the second input content based on the correlation.
[0049] In some embodiments, firstly, the spatial relationship between the attribute information of the target object in the area to be evaluated and the global description information of the area to be evaluated is determined, and the attribute information and the global description information are mapped to the same spatial coordinate system. The spatial relationship between the two types of information is determined based on the coordinates of the attribute information and the coordinates of the global description information. For example, if the location of the target object in the attribute information is the latitude and longitude of T001, and it falls within the range of plot D01 in the global description information, then T001 is determined to belong to D01. Secondly, the semantic relationship between the attribute information of the target object in the area to be evaluated and the global description information of the area to be evaluated is determined, such as the global cultivated land plot D02 corresponding to wheat field targets T005-T010 in the attribute information. Finally, the two types of information are fused into the second input content based on the spatial relationship and the semantic relationship.
[0050] For example, if the attribute information is that the roof of a 6-story residential building has a damaged area of 28% and the wall crack length is 1.5m, and the global description information is that an earthquake occurred in the residential area in the north of the city, then the second input content can be that an earthquake occurred in the residential area in the north of the city, wherein the roof of a 6-story residential building has a damaged area of 28% and the wall crack length is 1.5m.
[0051] Step S202: Determine the target query vector based on the similarity between the vector corresponding to the second input content and each of the query vectors.
[0052] In some embodiments, a query vector is a vector representation of each damage assessment criterion in a preset knowledge base, and each query vector corresponds to the text content of at least one damage assessment criterion.
[0053] In some embodiments, the second input content is first converted into a high-dimensional semantic vector. Then, the similarity between the vector corresponding to the second input content and all query vectors is calculated, typically using methods such as cosine similarity or Euclidean distance for matching. The query vector with the highest similarity is selected as the target query vector, and the damage assessment criterion associated with the target query vector is the most relevant assessment basis in the current scenario.
[0054] In some embodiments, the vector similarity matching mechanism can quickly locate the most relevant damage assessment criteria, enabling efficient retrieval of the knowledge base and improving the response speed and professional adaptability of damage assessment.
[0055] Step S203: Based on the text content of the damage assessment standard corresponding to the target query vector and the second input content, construct the first input content.
[0056] In some embodiments, the first input content serves as input prompts to guide the target model in generating the final evaluation report. The first input content combines the damage assessment standard text corresponding to the target query vector with information from the second input content to form a structured input format. For example, an evaluation report is generated based on the following target damage information and scene description, combined with the damage assessment standard. This structured input ensures that the model can correctly reference professional knowledge during the generation process and maintain the professionalism and accuracy of the output content.
[0057] In some embodiments, by fusing damage assessment criteria with specific scenario information, structured input content is constructed, enabling the target model to generate damage assessment reports that conform to domain specifications without relying on fine-tuning. The construction of this structured input content significantly improves the interpretability and task adaptability of the model.
[0058] In this embodiment, attribute information and global description information are extracted and fused to generate a second input content, which serves as a unified data carrier. Next, vector similarity matching technology is used to find the most relevant damage assessment criteria from a pre-defined knowledge base, forming a target query vector. Finally, the target query vector is combined with the second input content to construct structured input content for use by the target model. The entire processing from raw remote sensing data to a professional damage assessment report is automated, thereby improving the system's intelligence and assessment efficiency.
[0059] Figure 3 This is a schematic diagram illustrating the implementation flow of a damage assessment method provided in an embodiment of this application. This method can be executed by the processor of a computer device. Based on Figure 2 , Figure 2 Step S201 can be updated to steps S301 to S303, combining Figure 3 The steps shown are explained.
[0060] Step S301: Extract the attribute features of the target object based on the attribute information, and extract the global features of the region to be evaluated based on the global description information.
[0061] In some embodiments, attribute feature extraction is typically achieved using a fine-grained target perception big model (such as the Osmo series), which can identify the specific attributes of a target and output them in a structured manner.
[0062] In some embodiments, global feature extraction relies on a large scene global description model, which extracts textual descriptions of elements such as geographic environment, building density, and road network by performing multimodal encoding on remote sensing images.
[0063] Step S302: Based on the mapping relationship of the spatial coordinate system, establish a text matching relationship between the attribute features of the target object and the global features of the region to be evaluated.
[0064] In some embodiments, a spatial coordinate system provides a unified reference frame, enabling a logical connection between the location of a target object and its surrounding regional context. For example, in a target hangar located in an industrial area, the extent of damage to the target object may be affected by factors such as the layout of surrounding factories and traffic conditions.
[0065] In some embodiments, a pre-trained multimodal model is used to perform cross-modal alignment between the target object's attribute information (such as location coordinates and damage level) and global descriptive information (such as regional functional attributes and building cluster distribution patterns). Cross-modal alignment is typically achieved through vector similarity calculation within the embedding space, ultimately forming a semantic mapping relationship between the target object and the scene.
[0066] In some embodiments, the spatial coordinate system is not only used to locate the spatial position of the target object, but also serves as a bridge connecting attribute information and global description information. The spatial coordinate system ensures the alignment of attribute information and global description information in spatial dimensions, and enhances the accuracy and interpretability of semantic matching.
[0067] Step S303: Based on the text matching relationship, semantically concatenate the attribute information of the target object and the global description information to obtain the second input content.
[0068] In some embodiments, various methods can be used for concatenation, such as direct concatenation, weighted averaging, or attention-based fusion. This enables the language big model to better understand the complex relationship between the target object and its environment, thereby generating more accurate and professional damage assessment reports.
[0069] In some embodiments, semantic concatenation not only improves the information density of the input content, but also enhances the language model's ability to understand damage features, providing high-quality input support for the target object evaluation task.
[0070] In some embodiments, features of the target object and the region are extracted based on attribute information and global description information, respectively. A semantic matching relationship between the target object and the region is established using a spatial coordinate system, and then the second input content is generated through semantic concatenation. This approach allows for simultaneous consideration of the target object's state and the influence of the overall environment during the assessment process, thereby improving the comprehensiveness and accuracy of the damage assessment and ultimately generating an assessment report that conforms to professional standards.
[0071] In this embodiment, basic information about the target object and the region is obtained through feature extraction. Next, a spatial and semantic relationship is established between them using a spatial coordinate system. Finally, semantic concatenation integrates the attribute information of the target object and the global description information of the region to be evaluated into a coherent and context-rich input for further processing by a large language model. This ensures the integrity and consistency of information transmission and improves the accuracy of the damage assessment report.
[0072] Figure 4 This is a schematic diagram illustrating the implementation flow of a damage assessment method provided in an embodiment of this application. This method can be executed by the processor of a computer device. Based on Figure 2 , Figure 2 Step S202 can be updated to steps S401 to S403, combining Figure 4 The steps shown are explained.
[0073] Step S401: Extract features from the second input content to obtain the attribute feature vector of the target object and the global feature vector of the region to be evaluated.
[0074] In some embodiments, feature extraction can be performed on the second input content based on a trained deep neural network model to extract feature vectors corresponding to attribute information in the second input content and feature vectors corresponding to global description information of the region to be evaluated.
[0075] In some embodiments, the attribute feature vector is used to describe a local part of the target object in the area to be evaluated. Examples include building type, damage level, and spatial location.
[0076] In some embodiments, the global feature vector is used to describe the macro-environmental information of the entire area to be evaluated, such as topography, building distribution patterns, and vegetation cover, to provide contextual knowledge support required for the evaluation.
[0077] In some embodiments, the attribute feature vector and the global feature vector together constitute the semantic association between the target object and its environment, which helps to improve the accuracy and comprehensiveness of damage assessment. By fusing the attribute features of the target object with the overall regional features, the scope and severity of the damage event can be understood more comprehensively.
[0078] In this embodiment, by introducing attribute feature vectors and global feature vectors, the target object and its environmental state can be more accurately characterized. Introducing attribute feature vectors and global feature vectors enhances the context awareness of the damage assessment system, thereby improving the accuracy of damage assessment and enabling it to adapt to damage identification needs in various complex scenarios.
[0079] Step S402: Determine the similarity between the attribute feature vector and the global feature vector and each of the query vectors.
[0080] In some embodiments, similarity is an indicator that characterizes the degree of semantic similarity between two feature vectors, and can be quantified using methods such as cosine similarity, Euclidean distance, and Manhattan distance.
[0081] In some embodiments, similarity is used to assess the degree of matching between the characteristics of a target object and known damage standards or historical cases, thereby determining whether the current target meets a certain type of damage condition.
[0082] In some embodiments, the query vector represents a specific damage characteristic or assessment criterion, such as minor damage, road interruption, or building collapse. By calculating the similarity between the target object's attribute feature vector and global feature vector and the query vector, the situation where the damage criterion best matches the target object can be identified, thus providing a basis for subsequent damage classification and rating.
[0083] Step S403: Determine the query vectors whose similarity reaches the similarity threshold as the target query vector.
[0084] In some embodiments, a similarity threshold is used to distinguish between query vectors that match highly and those that are irrelevant.
[0085] In some embodiments, the similarity threshold can be dynamically adjusted based on task requirements, damage type, and historical data to balance the relationship between recall and precision. When the similarity between a query vector and a target feature vector is higher than the preset similarity threshold, the query vector is considered relevant and used as a reference for the final evaluation.
[0086] In some embodiments, by setting a similarity threshold, only query vectors that highly match the target object are retained, thereby filtering out noise and irrelevant information and improving the accuracy of the evaluation results. Furthermore, the similarity threshold can be set with differentiated criteria based on different damage levels; for example, higher matching requirements can be applied to severe damage, while lower similarity is allowed for minor damage, thus achieving more flexible and refined damage classification.
[0087] In some embodiments, by setting a reasonable similarity threshold, the most relevant damage criteria or assessment cases can be selected. By setting a reasonable similarity threshold, the false positive rate can be significantly reduced, thereby improving the accuracy and stability of damage assessment and meeting the high standards required by professional fields for assessment results.
[0088] In this embodiment, firstly, attribute feature vectors of the target object and global feature vectors of the region are obtained through feature extraction to construct a complete semantic representation. Secondly, similarity calculations are performed between the attribute feature vectors and global feature vectors and multiple predefined query vectors to identify the most matching damage criteria. Finally, based on a set similarity threshold, target query vectors that meet the conditions are selected, thereby completing the damage assessment of the target object. This achieves a closed-loop processing from data input to feature modeling and then to matching decision-making, improving the accuracy of subsequent assessment results.
[0089] Figure 5 This is a schematic diagram illustrating the implementation flow of a damage assessment method provided in an embodiment of this application. This method can be executed by the processor of a computer device. Based on Figure 1 , Figure 1 Step S101, which involves extracting the global description information of the region to be evaluated, can be updated to steps S501 and S502, combining... Figure 5 The steps shown are explained.
[0090] Step S501: In response to user input, construct prompt words.
[0091] The prompt word carries a generation standard for the global description information; the generation standard is determined at least based on the attribute characteristics of each type of land parcel in the area to be evaluated.
[0092] In some embodiments, user input can be questions or instructions in natural language, such as requesting a text description of the topography and building distribution of the area to be evaluated. The model automatically extracts key semantic information from the input and converts it into structured prompts.
[0093] In some embodiments, prompts are instructional texts used to guide the model in generating output with a specific format and content. They typically contain information such as task objectives, key elements, and output format. The process of constructing prompts includes keyword extraction, semantic parsing, and contextual modeling, ensuring that the generated prompts accurately reflect the user's intent.
[0094] In some embodiments, the generation standard refers to the specific rules and constraints followed when generating the global description information. The generation standard determines the granularity, coverage, and terminology used in the output content. The generation standard can be flexibly adjusted according to different application scenarios; for example, in the military field, the focus may be more on the damage level and strategic value of targets, while in urban planning, the emphasis may be more on land use and infrastructure distribution.
[0095] In some embodiments, the types of land parcels in the area to be evaluated may include building areas, water areas, farmland areas, etc.
[0096] In some embodiments, the attribute characteristics of different types of land parcels refer to the physical, functional, and spatial distribution characteristics of different land cover types identified in remote sensing images. For example, built-up areas may have characteristics such as high density, regular arrangement, and multi-layered structure; farmland areas may exhibit characteristics such as large-scale continuous planting and periodic changes; and water areas may present characteristics such as smooth boundaries and low reflectivity. By analyzing the attribute characteristics of land parcels, data support can be provided for generating standards, thereby ensuring that the global descriptive information is both comprehensive and meets the needs of actual scenarios.
[0097] In some embodiments, the prompt includes a user intent and a generation criterion. The user intent is determined based on user input information, and the generation criterion is obtained from a pre-set criterion library based on matching the user intent.
[0098] For example, if the user inputs "Please generate the topography and building distribution of the area to be evaluated", the system will identify the user's intent to generate the topography and building distribution of the area to be evaluated, and then match macro-information such as the geographical environment characteristics, regional functional attributes, and building cluster distribution patterns of the area to be evaluated from the standard library based on the user's intent.
[0099] Step S502: Based on the prompt words and the remote sensing image, generate global description information for the area to be evaluated using a global description model.
[0100] In some embodiments, the global description model is a multimodal large model capable of processing visual and linguistic information. Specifically, the global description model is pre-trained based on remote sensing images and corresponding text descriptions. The trained global description model can generate structured or unstructured text descriptions based on the input image and text content.
[0101] In some embodiments, based on the prompt words and the remote sensing image, the global description model performs encoding operations on the remote sensing image to extract its visual features. Then, the task instructions and generation criteria from the prompt words are jointly processed with the remote sensing image features, achieving the fusion of visual and linguistic information through a cross-modal attention mechanism. Finally, the global description model outputs one or more descriptive texts detailing the geographical environmental features, regional functional attributes, building cluster distribution patterns, and other information in the remote sensing image.
[0102] In some embodiments, the generated global description information includes a macro-level scene overview of the area to be assessed, as well as attribute characteristics down to specific plots. For example, for an urban area, the global description information includes the topography, water system distribution, vegetation coverage, and spatial layout of industrial and residential areas. This multi-layered description helps users quickly understand the overall situation of the area to be assessed and provides important reference for subsequent damage assessment. For example, if the input image is a remote sensing image of a coastal city, the global description information is as follows: the geographical environment of the coastal city is mainly plains, bordering the ocean to the east and having rivers flowing through it to the west; functional zones include a northern industrial area (dense factories), a central commercial area (high-rise building complexes), and a southern residential area (low-rise townhouses); the building complexes exhibit a density decreasing pattern from the center to the periphery.
[0103] In this embodiment, prompt words are constructed in response to user input, and global descriptive information is generated based on these prompt words and remote sensing images. This ensures that the generated descriptions meet the user's actual needs, thereby improving the professionalism and relevance of damage assessment, and ultimately providing more scientific and efficient decision support for scenarios such as disaster emergency response and battlefield situation assessment.
[0104] In this embodiment, the user first triggers the construction of prompt words by inputting information. These prompt words further define generation criteria, which are then integrated with the remote sensing image input. The integrated result is used as the input parameter of the global description model. During processing, the global description model fuses visual and linguistic information, ultimately outputting global description information that meets the user's needs.
[0105] Figure 6 This is a schematic diagram illustrating the implementation flow of a damage assessment method provided in an embodiment of this application. This method can be executed by the processor of a computer device. Based on Figure 5 , Figure 5 Step S502 can be updated to steps S601 to S603, combining Figure 6 The steps shown are explained.
[0106] Step S601: Determine the text feature vector of the prompt word and the semantic feature vector of the remote sensing data.
[0107] In some embodiments, a text feature vector refers to a high-dimensional vector representation obtained by converting natural language prompts into discrete ID sequences through a tokenizer within a multimodal large model, and then mapping them to an embedding space. Text feature vectors can capture semantic information in prompts, such as damage type, assessment dimension, and geographical features. For example, when the prompt is to identify building damage within an industrial area, the text feature vector will encode semantic representations of keywords such as industrial area, building, and damage.
[0108] In some embodiments, semantic feature vectors refer to high-dimensional semantic feature vectors extracted by a visual encoder (such as SigLIP-SO400M) after the remote sensing image is input. These semantic feature vectors not only contain pixel information from the remote sensing image but also incorporate the structured semantics of the remote sensing image content, such as topography, building distribution, and vegetation cover. In remote sensing images of coastal cities, semantic feature vectors can represent key geographical features such as the eastern ocean, western rivers, and densely populated industrial areas in the central region.
[0109] In some embodiments, cross-modal information alignment can be achieved by extracting textual feature vectors and semantic feature vectors separately. This alignment enhances the model's understanding of the damage scene. Textual feature vectors provide semantic guidance of user intent, while semantic feature vectors reflect objective environmental information in remotely sensed images. Combining textual and semantic feature vectors can improve the accuracy of subsequent processing.
[0110] Step S602: Spatially align the text feature vector and semantic feature vector to obtain the target feature vector.
[0111] In some embodiments, spatial alignment refers to aligning feature vectors from different modalities (text and images) in a unified embedding space so that these feature vectors are comparable and consistent.
[0112] In some embodiments, spatial alignment can be achieved through a multilayer perceptron (MLP) projection layer to ensure that the text feature vector and the semantic feature vector of the remote sensing image are in the same semantic representation space. For example, the dimension of the result obtained after projecting the text feature vector through the MLP matches the semantic feature vector of the remote sensing image, thus allowing for comparison and fusion in the same space.
[0113] In some embodiments, the target feature vector is a spatially aligned joint feature representation that integrates semantic information from the cue words and visual information from the remote sensing image. The target feature vector can serve as the foundational input for generating subsequent global description information, enabling the global description model to understand the damage scene in a more accurate context. Spatial alignment allows for a better capture of the relationship between the damaged target and its environment.
[0114] Step S603: Based on the target feature vector and the global description model, generate global description information for the region to be evaluated.
[0115] In some embodiments, the global description model can be a multimodal large model based on the Transformer architecture, capable of generating structured text descriptions based on the input target feature vector.
[0116] In some embodiments, the global description model outputs a text description containing geographic environmental features, regional functional attributes, and building distribution patterns through a token-by-token prediction method using a decoder. When the input is a remote sensing image of a coastal city, the description generated by the global description model might include: the geographic environment is predominantly plains, bordering the ocean to the east, with a river flowing through the west; functional zones include a northern industrial area (dense factories), a central commercial area (high-rise buildings), and a southern residential area (low-rise townhouses).
[0117] In some embodiments, the generated global description information not only provides users with an intuitive overview of the scene but also offers contextual support for subsequent professional evaluation reports. By combining target feature vectors, more targeted and accurate descriptions can be generated, avoiding generalized or vague outputs. The generated global description information can assist the Retrieval Enhanced Generation (RAG) module, which uses external knowledge bases to access relevant professional knowledge, thereby improving the professionalism and reliability of the final evaluation results.
[0118] In this embodiment, a target feature vector is generated by spatially aligning the text feature vector of the prompt words with the semantic feature vector of the remote sensing image. This target feature vector is then input into a global description model to generate a scene-level text description. This effectively integrates the user's query intent with the objective content of the remote sensing image, thereby improving the contextual understanding and semantic expression capabilities of damage assessment. Ultimately, this results in generating more accurate, professional, and domain-compliant damage assessment reports.
[0119] Figure 7 This is a schematic flowchart illustrating the implementation of a damage assessment method provided in an embodiment of this application. The method can be executed by a processor of a computer device. The method includes steps S701 to S703, combining... Figure 7 The steps shown are explained.
[0120] Step S701: Obtain the text content corresponding to at least one type of ground feature damage standard.
[0121] In some embodiments, ground feature damage standards refer to professional specifications and indicators used to assess the degree of damage to ground features, typically including building damage classification, infrastructure damage assessment criteria, and vegetation cover change thresholds. Ground feature damage standards can be derived from damage assessment specifications formulated by the state or industry, such as the "Standard for Grading Damage to Civil Building Structures" and the "Guideline for Identifying Surface Damage in Remote Sensing Images."
[0122] In some embodiments, by obtaining descriptions of textual content corresponding to damage standards for multiple types of land features, a knowledge base covering different land feature types and damage scenarios can be constructed to provide accurate reference for subsequent damage assessment.
[0123] Step S702: Identify each text content based on the intent recognition model to obtain the query vector corresponding to each text content.
[0124] In some embodiments, the intent recognition model is a deep learning model that extracts the task objective from the text to understand the semantics of the text content corresponding to each land feature damage standard, and maps the text content corresponding to the land feature damage standard into a machine-processable mathematical representation, such as a query vector.
[0125] In some embodiments, by processing the text content corresponding to the damage standards of ground features through an intent recognition model, complex natural language descriptions can be transformed into structured numerical representations, thereby enabling efficient semantic retrieval and matching and improving the accuracy of damage assessment.
[0126] Step S703: Generate the preset knowledge base based on the entries formed by the query vectors corresponding to each text content.
[0127] In some embodiments, the preset knowledge base is a database structure consisting of multiple query vectors and the text content of the corresponding land cover damage standards for each query vector, used to store and manage information on various land cover damage standards. Each entry uses a query vector as an index and the text content of the corresponding land cover damage standard as the entry content, thereby supporting fast retrieval and matching.
[0128] In some embodiments, by constructing a knowledge base based on query vectors, when user input or remote sensing image data is received, the text content corresponding to the ground feature damage standards related to the query vector knowledge base can be quickly invoked for matching and reasoning, thereby outputting professional assessment results that conform to domain specifications.
[0129] In this embodiment, a pre-defined knowledge base is constructed by acquiring the text content corresponding to the damage standards for ground features and generating query vectors. This method enables efficient management and retrieval of various damage standards, thereby improving the accuracy of damage assessment and meeting the need for rapid response in multiple scenarios.
[0130] The following describes an exemplary application of a damage assessment method provided in this application in a real-world scenario.
[0131] In key areas such as security defense, maritime traffic management, environmental monitoring, disaster emergency response, and counter-terrorism, there is an urgent need for rapid and accurate damage assessment after sudden events such as natural disasters, regional conflicts, and man-made sabotage. With the rapid development of remote sensing technology, high-resolution satellite and UAV images provide rich surface information, which is crucial for rapid disaster response, disaster relief planning, assessment of damage to land features, and environmental remediation. However, traditional remote sensing damage assessment methods often rely on manual visual interpretation or automated processing using a single model. These methods are inadequate for handling large-scale, diverse damage data. While accurate, manual visual interpretation is time-consuming and labor-intensive, and cannot meet the demands of rapid response. Automated processing using a single model is limited by its generalization capabilities and struggles to handle complex and varied damage scenarios. In practical applications, the damage scenarios depicted by remote sensing images are often complex and varied, requiring simultaneous assessment of multiple damage types and degrees. For example, in monitoring natural disasters or man-made damage, it is not only necessary to accurately identify and assess the damage to individual buildings, but also to comprehensively analyze the various types of damage across the entire affected area. The system should be highly flexible, capable of handling diverse damage queries, and possess a deep understanding of complex damage characteristics to ensure seamless switching between different damage scenarios and meet the needs of real-time assessment and efficient analysis.
[0132] With the rapid development of multimodal large-scale models, new opportunities have emerged for remote sensing damage assessment. Multimodal large-scale models achieve efficient alignment of image and text descriptions by fusing visual and linguistic information. However, in practical applications, multimodal large-scale models also face challenges such as large parameter scale, high computational resource consumption during retraining, and insufficient fine-grained information extraction capabilities. Because damage assessment is a domain-specific and specialized application, traditional data-driven large-scale models often suffer from issues such as de-specialization of output content and AI illusions when performing tasks. Especially in the precise assessment of damage location and extent, existing models often only provide image-level text alignment, lacking the ability to perform fine-grained analysis on specific damaged targets, thus limiting assessment accuracy.
[0133] Related technologies are typically optimized for single tasks, lacking flexibility and struggling to adapt to diverse damage assessment needs. Generative multimodal large models, when dealing with damage assessment tasks, lack specialized domain knowledge and cannot provide domain-based professional answers. Furthermore, simply using generative multimodal large models has limitations in positioning accuracy, lacking fine-grained information extraction capabilities, leading to inaccurate damage target location detection. Deep learning-based damage assessment frameworks require retraining or fine-tuning processes when facing new damage types or assessment tasks, significantly increasing computational and deployment costs.
[0134] To address the aforementioned issues, this application proposes a multimodal large-scale scene target perception method for remote sensing damage assessment. (Reference) Figure 8 , Figure 8 A schematic diagram illustrating the implementation process of a multimodal large-model scene target perception method for remote sensing damage assessment, provided in this application embodiment, includes the following steps: Step S801: Obtain the high-resolution remote sensing image to be evaluated, extract the precise location information and preliminary damage rating of all targets in the scene through the fine-grained target perception large model, and generate a structured description containing target type, physical attributes and spatial distribution relationship.
[0135] In some embodiments, the initial damage rating may be intact, slightly damaged, severely damaged, etc.
[0136] In some embodiments, the structure is described as follows: "objects": [ "object_1":{"class":"class_1", Damage level: "Minor" "Position": (x1, y1, x2, y2)}.
[0137] Step S802: Input the original remote sensing image and prompt words into the scene global description model, and generate a scene-level text description containing macro information such as geographical environmental features, regional functional attributes, and building cluster distribution patterns through multimodal joint encoding, providing global contextual knowledge.
[0138] In some embodiments, the prompts include user intent and generation requirements; wherein, the user intent may be to generate scene-level text descriptions based on input remote sensing images; the generation requirements may be geographical environmental features: topography (plains / mountains / hills), water systems (rivers / lakes / oceans), vegetation cover type and proportion; regional functional attributes: distribution range of industrial areas, farmland, residential areas, commercial areas, and transportation hubs; building cluster distribution patterns: density (high / medium / low), arrangement (regular grid / irregular dispersion), and height hierarchy (low-rise residential buildings / high-rise buildings).
[0139] In some embodiments, the remote sensing image and the prompt words first need to be fused in a multimodal manner, including: converting the text into a discrete ID sequence (such as BPE word segmentation) using the tokenizer built into the large language model and mapping it to the embedding space; converting the image into high-dimensional semantic features using the SigLIP-SO400M encoder; aligning it to the embedding space through the MLP projection layer; and concatenating the text embedding with the visual feature sequence to form a multimodal input sequence (such as: [text_tokens] [image_features]).
[0140] In some embodiments, the scene-level text description (corresponding to the global description information in the above embodiments) may be: the input image is a remote sensing image of a coastal city, and the text description is: the geographical environment is mainly plains, bordering the ocean to the east and a river running through the west; the functional zones include the northern industrial zone (dense factories), the central commercial zone (high-rise building complex), and the southern residential zone (low-rise townhouses); the building complex shows a distribution pattern of decreasing density from the center to the periphery.
[0141] Step S803: Construct a language big model processing framework based on RAG technology, semantically align the preliminary damage rating, target type, and scene-level text description, call the professional assessment knowledge base through the retrieval enhancement generation mechanism, and finally output a professional assessment report containing elements such as target individual damage degree analysis, key facility impact assessment, and overall scene damage index.
[0142] In some embodiments, it is first necessary to construct input information (corresponding to the second input content in the above embodiments) for the professional knowledge base (corresponding to the preset knowledge base in the above embodiments); the input information of the professional knowledge base can be a semantically aligned string: "overall scene situation: "+global_answer+'damage status of each target in the scene: '+object_answer; where, an example of global_answer: the geographical environment is mainly plains, the eastern region is close to the ocean, and the western region has rivers running through it; the functional zoning of the city includes the northern region as an industrial area (dense factories), the central region as a commercial area (high-rise building complexes) and the southern region as a residential area (low-rise townhouses); the building complexes of the city show a distribution pattern of gradually decreasing density from the central area to the outer area; an example of object_answer: the {}th target: type {}, damage level {}, location {}.
[0143] In some embodiments, damage assessment criteria are matched from a professional knowledge base based on input information. This includes converting the input information into a retrieval query, which can be achieved by extracting keywords or by embedding (encoding) the question text of the input information. Specifically, in a pre-built external knowledge base (such as documents, databases, or APIs), the most relevant text information chunks to the current retrieval query are found through vector similarity search (such as FAISS, Chroma) or keyword matching (such as Elasticsearch). The retrieval results are then sorted, deduplicated, or filtered, retaining highly relevant content as the basis for generation.
[0144] In some embodiments, the retrieved damage assessment criteria and input information for the knowledge base are combined to form a complete question, which is then input into the language model to obtain a professional assessment report.
[0145] In some embodiments, the final knowledge base structure can be briefly described as follows: {Text block 1 query vector: Text block 1,} Text Block 2 Query Vector: Text Block 2, Text block 3 query vector: text block 3, ...}.
[0146] In some embodiments, the knowledge base construction process includes: first, cleaning and converting various format texts (pdf, txt, doc) into plain text; dividing the plain text into text blocks (natural segmentation, equal-length segmentation) according to different standards; and constructing query vectors for each block and storing them in the database.
[0147] In some embodiments, the output of the professional assessment report is as follows: Question: Based on the above situation, strictly follow the format requirements to generate a damage report. Do not generate content that does not meet the format requirements. An example is shown below: Based on the images you provided, and in conjunction with the target damage rule knowledge base, the criteria and conclusions for determining the scene damage level are as follows (assessment report): Damage Situation 1: Target: Hangar Damage Level: Severe Quantity: 3 locations; Criteria for judgment: The hangar is unusable and difficult to repair, and most of its supplies (equipment and ammunition) have been destroyed. The text assessment report concludes: On August 26, 2024, based on satellite reconnaissance imagery, an international airport was attacked, with three hangars in a localized area severely damaged and unlikely to be repaired in the short term.
[0148] In some embodiments, the aforementioned fine-grained target perception large model (Osmo series) is built on a multi-task pre-training framework, supporting parallel tasks such as target detection, instance segmentation, and attribute recognition. It employs a cascaded feature pyramid network to process targets of different scales, and dynamically adapts to image input resolutions from 1024×1024 to 4096×4096 using dynamic convolutional kernels. For ultra-large images, a 2048×2048 sliding window is used for block processing, and a non-maximum suppression fusion strategy is employed for overlapping window regions.
[0149] In some embodiments, the large-scale scene description model is pre-trained based on multimodal contrastive learning, fusing a remote sensing image encoder and a text decoder. The encoder adopts a ResNet-Transformer hybrid architecture, supporting panoramic segmentation and dense description generation; the decoder generates structured scene description text containing elements such as terrain features, building density, and road networks through an attention routing mechanism.
[0150] In some embodiments, a professional assessment engine based on a RAG-based language large model integration retrieval enhancement generation mechanism includes a professional assessment knowledge base covering fields such as civil engineering, civil architecture, and infrastructure. A dual-path attention architecture is employed, where the retrieval path invokes assessment standard clauses through semantic similarity matching, and the generation path integrates target data, scene features, and retrieval results to generate graded and categorized damage assessment conclusions according to preset assessment procedures.
[0151] In some embodiments, this application achieves professional-grade remote sensing damage assessment without fine-tuning model parameters through multi-model collaboration and knowledge enhancement technology. While maintaining the general capabilities of large models, it effectively addresses the shortcomings of traditional methods in damage feature understanding, environmental context association, and professional standard adaptation, meeting the rapid response needs of scenarios such as battlefield assessment, disaster relief, and facility maintenance.
[0152] In some embodiments, this application addresses scenarios such as disaster emergency response, battlefield situation assessment, and infrastructure maintenance. It utilizes remote sensing images acquired by satellite or airborne platforms, containing targets such as buildings, roads, and vegetation, to achieve professional-grade damage assessment without parameter fine-tuning through multimodal large-scale model collaboration. The specific implementation process is as follows: S1. Data Acquisition and Preprocessing: S11. Obtain the dataset of remote sensing images to be evaluated: high-resolution remote sensing images with different resolutions (0.3m to 2m) and multispectral bands, covering typical scenes such as urban building complexes, transportation hubs, and natural landforms, with image sizes ranging from 2048×2048 to 8192×8192.
[0153] S12. Construct a professional assessment knowledge base: Integrate professional knowledge such as standards for damage to land features, safety specifications for civil buildings, and damage classification of infrastructure to form a structured retrieval database, which includes core elements such as damage level definition, impact weight parameters, and mapping relationships of related facilities.
[0154] S2, Target-level fine-grained feature extraction: S21. Input the original remote sensing image into the fine-grained target perception large model (Osmo 1.0 / 2.0 / 3.0). Through pre-trained target detection and attribute analysis capabilities (Osmo 1.0 is based on a cascaded feature pyramid network, Osmo 2.0 integrates a dynamic sparse attention mechanism, and Osmo 3.0 supports ultra-large-scale image processing of 8192×8192), output the precise location coordinates (x1, y1, x2, y2) of all targets in the scene and the preliminary damage rating (intact, slightly damaged, severely damaged, completely damaged), and generate a structured description of the target type, physical attributes and spatial distribution relationship.
[0155] S22. For ultra-large images, a sliding window block processing strategy is adopted (Osmo 3.0 supports a 2048×2048 sliding window). The boundary effect is eliminated by the overlapping area feature fusion technology to ensure the consistency of target positioning and rating.
[0156] S3, Scene Global Feature Description: S31. Input the same remote sensing image into the scene global description large model (multimodal large models such as Janus-7B / 13B), and generate scene-level text descriptions containing terrain features, building density distribution, road network connectivity, and vegetation coverage based on the vision-language joint pre-training capability (Janus-7B supports real-time processing of 4096×4096 images, and Janus-13B integrates panoramic segmentation and dense semantic reasoning) as evaluation context knowledge.
[0157] S32. Extract multi-scale visual context information (Janus-13B supports multi-scale attention feature encoding) to form scene semantic prior knowledge that associates geospatial relationships with functional attributes.
[0158] S4. Multi-model collaborative evaluation process: S41. Target-Scene Feature Alignment: Through spatial coordinate system mapping, target-level damage information is semantically associated with the global scene description, establishing a text matching relationship between target attributes and scene elements.
[0159] S42, Enhanced evaluation of large language models based on RAG (GPT-3.5 / 4, DeepSeek-1.5B / 7B): Search phase: Based on target type, damage level, and scene characteristics, a search vector is generated to retrieve matching assessment clauses, standards and specifications, and historical cases from the professional assessment knowledge base.
[0160] Generation phase: Input the target damage data, scene context features and retrieval results into the model (GPT-3.5 supports rapid generation of structured reports, GPT-4 enhances complex reasoning capabilities, DeepSeek-1.5B adapts to edge computing, and DeepSeek-7B supports battlefield-level simulation), and fuse information through a multimodal attention mechanism to generate an evaluation report containing the following: a) individual target damage level; b) a comprehensive scene damage report.
[0161] Figure 9 This is a schematic diagram of a damage assessment framework provided in an embodiment of this application, wherein 901 is a fine-grained target perception large model, 902 is a large model capable of global description, 903 is a large language model, 904 is a word segmentation embedding block model, and 905 is a vector knowledge base.
[0162] In some embodiments, the fine-grained target perception large model 901 generates damage information for each target in the area to be evaluated based on high-resolution remote sensing images, including the target type, location, and initial damage level.
[0163] In some embodiments, the global description large model 902 generates global scene damage information based on high-resolution remote sensing images and cue words.
[0164] In some embodiments, the word segmentation embedding block model converts each damage assessment criterion in the damage assessment knowledge base into a vector, generating a vector knowledge base 905.
[0165] In some embodiments, the large language model 903 matches relevant damage assessment criteria from the vector knowledge base 905 based on the damage information of each target and the global damage information of the scene, and generates a textual result of the damage assessment based on the damage assessment criteria, the damage information of each target and the global damage information of the scene.
[0166] In some embodiments, this application proposes a zero-shot remote sensing damage assessment framework based on a pre-trained multimodal large model. By dynamically fusing damage descriptions at different levels, an alignment mapping of the visual-linguistic feature space is constructed, enabling cross-modal retrieval and quantitative assessment of damaged targets without fine-tuning.
[0167] In some embodiments, this application constructs a multi-scale spatial damage information fusion mechanism, adopting a strategy that combines target-level damage information with regional-level damage information to achieve hierarchical evaluation from image-level global perception to target-level precise positioning without the need for parameter optimization.
[0168] In some embodiments, this application establishes a dynamic knowledge base retrieval system, integrates multiple damage assessment standards according to task requirements, and achieves self-correction of assessment results through multi-source information fusion and logical consistency verification, ensuring that the output conforms to the knowledge specifications of the professional field.
[0169] The damage assessment method provided in this application has at least the following technical effects: 1. Adaptive Enhancement of Assessment Task: Traditional damage assessment models rely on predefined task frameworks, making it difficult to cope with dynamic assessment requirements. This solution, through deep coupling of a dynamic knowledge base and the RAG strategy, constructs an open semantic interaction interface, supporting natural language-driven damage feature description, customized assessment dimensions, and cross-scenario transfer, achieving dynamic expansion of domain knowledge and improved task generalization capabilities without retraining; 2. Fine-grained semantic association text input: To address the problem of decoupling local features and global context of damaged targets, a knowledge-guided visual-language alignment architecture is designed. Through real-time retrieval of professional terms and damage mechanism features from a dynamic knowledge base, combined with multi-scale damage description semantic text, multi-scale semantic association analysis of the damaged area is achieved. 3. Lightweight Professional Capability Injection: Breaking through the traditional large model parameter fine-tuning paradigm, a domain knowledge enhancement framework based on RAG is proposed. Through real-time semantic retrieval and feature mapping of dynamic knowledge base, professional knowledge such as geospatial rules and damage assessment standards are seamlessly integrated into the visual feature encoding and assessment report generation process, significantly reducing the computational overhead of domain adaptation.
[0170] Based on the foregoing embodiments, this application provides a damage assessment device, which includes various units and modules included in each unit. It can be implemented by a processor in a computer device; of course, it can also be implemented by specific logic circuits. In the implementation process, the processor can be a central processing unit (CPU), a microprocessor unit (MPU), a digital signal processor (DSP), or a field programmable gate array (FPGA), etc.
[0171] Figure 10 This is a schematic diagram of the composition structure of a damage assessment device provided in an embodiment of this application, as shown below. Figure 10As shown, the damage assessment device 1000 includes: an extraction module 1001, a construction module 1002, and a generation module 1003, wherein: the extraction module 1001 is used to extract global descriptive information of the area to be assessed and attribute information of each target object in the area to be assessed from a remote sensing image of the area to be assessed; the attribute information includes the category, location, and preliminary damage rating of the target object; the global descriptive information includes at least the attribute characteristics of each type of land parcel in the area to be assessed; the construction module 1002 is used to construct first input content for guiding the target model based on the attribute information, the global descriptive information, and a preset knowledge base; the generation module 1003 is used to generate a global damage assessment report for the area to be assessed based on the first input content and the target model.
[0172] In some embodiments, the preset knowledge base includes multiple query vectors, each query vector corresponding to at least one damage assessment standard text content; the construction module 1002 is further configured to generate second input content based on the attribute information and the global description information; determine a target query vector based on the similarity between the vector corresponding to the second input content and each query vector; and construct the first input content based on the text content of the damage assessment standard corresponding to the target query vector and the second input content.
[0173] In some embodiments, the construction module 1002 is further configured to extract attribute features of the target object based on the attribute information, and extract global features of the region to be evaluated based on the global description information; establish a text matching relationship between the attribute features of the target object and the global features of the region to be evaluated based on the mapping relationship of the spatial coordinate system; and semantically concatenate the attribute information of the target object and the global description information based on the text matching relationship to obtain the second input content.
[0174] In some embodiments, the construction module 1002 is further configured to perform feature extraction on the second input content to obtain the attribute feature vector of the target object and the global feature vector of the region to be evaluated; determine the similarity between the attribute feature vector and the global feature vector and each of the query vectors; and determine the query vectors whose similarity reaches a similarity threshold as the target query vectors.
[0175] In some embodiments, the extraction module 1001 is further configured to perform feature extraction on the second input content to obtain the attribute feature vector of the target object and the global feature vector of the region to be evaluated; determine the similarity between the attribute feature vector and the global feature vector and each of the query vectors; and determine the query vectors whose similarity reaches a similarity threshold as the target query vectors.
[0176] In some embodiments, the extraction module 1001 is further configured to determine the text feature vector of the prompt word and the semantic feature vector of the remote sensing data; spatially align the text feature vector and the semantic feature vector to obtain a target feature vector; and generate global description information for the region to be evaluated based on the target feature vector and the global description model.
[0177] In some embodiments, the generation module 1003 is further configured to acquire text content corresponding to at least one type of ground feature damage standard; identify each text content based on an intent recognition model to obtain a query vector corresponding to each text content; and generate the preset knowledge base based on the entries formed by the query vectors corresponding to each text content.
[0178] The descriptions of the apparatus embodiments above are similar to those of the method embodiments above, and have similar beneficial effects. In some embodiments, the functions or modules included in the apparatus provided in this application can be used to perform the methods described in the method embodiments above. For technical details not disclosed in the apparatus embodiments of this application, please refer to the descriptions of the method embodiments of this application for understanding.
[0179] It should be noted that, in the embodiments of this application, if the above-described methods are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, or the parts that contribute to related technologies, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), magnetic disks, or optical disks. Thus, the embodiments of this application are not limited to any specific hardware, software, or firmware, or any combination of hardware, software, and firmware.
[0180] This application provides a computer device including a memory and a processor. The memory stores a computer program that can run on the processor. When the processor executes the program, it implements some or all of the steps in the above-described method.
[0181] This application provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements some or all of the steps in the above-described method. The computer-readable storage medium can be transient or non-transient.
[0182] This application provides a computer program including computer-readable code, wherein when the computer-readable code is executed in a computer device, a processor in the computer device performs some or all of the steps in the above-described method.
[0183] This application provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program. When the computer program is read and executed by a computer, it implements some or all of the steps in the above-described method. This computer program product can be implemented specifically through hardware, software, or a combination thereof. In some embodiments, the computer program product is specifically embodied as a computer storage medium; in other embodiments, the computer program product is specifically embodied as a software product, such as a software development kit (SDK), etc.
[0184] It should be noted that the descriptions of the various embodiments above tend to emphasize the differences between them, while their similarities or commonalities can be referred to interchangeably. The descriptions of the above embodiments of the device, storage medium, computer program, and computer program product are similar to the descriptions of the above method embodiments and have similar beneficial effects. For technical details not disclosed in the embodiments of the device, storage medium, computer program, and computer program product of this application, please refer to the descriptions of the method embodiments of this application for understanding.
[0185] Figure 11 This application provides a hardware entity diagram of a computer device as an embodiment of the present application, such as... Figure 11 As shown, the hardware entity of the computer device 1100 includes a processor 1101 and a memory 1102, wherein the memory 1102 stores a computer program that can run on the processor 1101, and the processor 1101 executes the program to implement the steps in the method of any of the above embodiments.
[0186] The memory 1102 stores computer programs that can run on the processor. The memory 1102 is configured to store instructions and applications that can be executed by the processor 1101. It can also cache data to be processed or already processed (e.g., image data, audio data, voice communication data, and video communication data) in the processor 1101 and various modules in the computer device 1100. It can be implemented by flash memory or random access memory (RAM).
[0187] The processor 1101 executes the steps of any of the above methods when executing a program. The processor 1101 typically controls the overall operation of the computer device 1100.
[0188] This application provides a computer storage medium that stores one or more programs, which can be executed by one or more processors to implement the steps of the methods described in any of the above embodiments.
[0189] It should be noted that the descriptions of the storage medium and device embodiments above are similar to the descriptions of the method embodiments above, and have similar beneficial effects. For technical details not disclosed in the storage medium and device embodiments of this application, please refer to the descriptions of the method embodiments of this application for understanding.
[0190] The aforementioned processor can be at least one of the following: Application Specific Integrated Circuit (ASIC), Digital Signal Processor (DSP), Digital Signal Processing Device (DSPD), Programmable Logic Device (PLD), Field Programmable Gate Array (FPGA), Central Processing Unit (CPU), Controller, Microcontroller, and Microprocessor. It is understood that other electronic devices can also implement the functions of the aforementioned processor, and this application does not specifically limit the specific implementation.
[0191] The aforementioned computer storage media / memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), magnetic random access memory (FRAM), flash memory, magnetic surface memory, optical disc, or compact disc read-only memory (CD-ROM), etc.; or it can be various terminals that include one or any combination of the above-mentioned memories, such as mobile phones, computers, tablet devices, personal digital assistants, etc.
[0192] It should be understood that the phrase "one embodiment" or "an embodiment" throughout the specification means that a specific feature, structure, or characteristic related to the embodiment is included in at least one embodiment of this application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. It should be understood that in the various embodiments of this application, the sequence numbers of the above steps / processes do not imply a sequential order of execution; the execution order of each step / process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application. The sequence numbers of the above embodiments of this application are merely descriptive and do not represent the superiority or inferiority of the embodiments.
[0193] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0194] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple units or components can be combined, or integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed can be through some interfaces, and the indirect coupling or communication connection between devices or units can be electrical, mechanical, or other forms.
[0195] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units. They may be located in one place or distributed across multiple network units. Some or all of the units may be selected to achieve the purpose of this embodiment according to actual needs.
[0196] Furthermore, in the various embodiments of this application, all functional units can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in a combination of hardware and software functional units. Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, read-only memory (ROM), magnetic disks, or optical disks.
[0197] Alternatively, if the integrated units described above are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, or the part that contributes to related technologies, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, ROM, magnetic disks, or optical disks.
[0198] The above description is merely an embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.
Claims
1. A damage assessment method, characterized in that, The method includes: From the remote sensing image of the area to be evaluated, extract the global description information of the area to be evaluated and the attribute information of each target object in the area to be evaluated. The attribute information includes the category, location and preliminary damage rating of the target object. The global description information includes at least the attribute characteristics of each type of land parcel in the area to be evaluated. Based on the attribute information, the global description information, and the preset knowledge base, a first input content is constructed to guide the target model; Based on the first input content and the target model, a global damage assessment report for the area to be assessed is generated.
2. The method according to claim 1, characterized in that, The preset knowledge base includes multiple query vectors, and each query vector corresponds to at least one text content of a damage assessment standard. The construction of the first input content for guiding the target model based on the attribute information, the global description information, and the preset knowledge base includes: Based on the attribute information and the global description information, the second input content is generated; The target query vector is determined based on the similarity between the vector corresponding to the second input content and each of the query vectors. The first input content is constructed based on the text content of the damage assessment standard corresponding to the target query vector and the second input content.
3. The method according to claim 2, characterized in that, The step of generating the second input content based on the attribute information and the global description information includes: Based on the attribute information, the attribute features of the target object are extracted, and based on the global description information, the global features of the region to be evaluated are extracted. Based on the mapping relationship of the spatial coordinate system, a text matching relationship is established between the attribute features of the target object and the global features of the region to be evaluated; Based on the text matching relationship, the attribute information of the target object and the global description information are semantically concatenated to obtain the second input content.
4. The method according to claim 2, characterized in that, The step of determining the target query vector based on the similarity between the vector corresponding to the second input content and each of the query vectors includes: Feature extraction is performed on the second input content to obtain the attribute feature vector of the target object and the global feature vector of the region to be evaluated; Determine the similarity between the attribute feature vector and the global feature vector and each of the query vectors; The query vectors whose similarity reaches the similarity threshold are determined as the target query vectors.
5. The method according to any one of claims 1 to 4, characterized in that, The step of extracting global descriptive information of the region to be evaluated from the remote sensing image of the region to be evaluated includes: In response to user input, prompt words are constructed; the prompt words carry generation criteria for the global description information; the generation criteria are determined at least based on the attribute characteristics of each type of land parcel in the area to be evaluated; Based on the prompt words and the remote sensing image, a global description information for the area to be evaluated is generated through a global description model.
6. The method according to claim 5, characterized in that, The step of generating global descriptive information for the region to be evaluated based on the prompt words and the remote sensing image through a global description model includes: Determine the text feature vector of the prompt word and the semantic feature vector of the remote sensing data; The text feature vector and semantic feature vector are spatially aligned to obtain the target feature vector; Based on the target feature vector and the global description model, global description information for the region to be evaluated is generated.
7. The method according to any one of claims 1 to 4, characterized in that, The method further includes: Obtain the text content corresponding to at least one type of ground feature damage standard; Based on the intent recognition model, the text content is identified to obtain the query vector corresponding to each text content; The preset knowledge base is generated based on the entries formed by the query vectors corresponding to each text content.
8. A damage assessment device, characterized in that, The device includes: The extraction module is used to extract global descriptive information of the area to be evaluated and attribute information of each target object in the area to be evaluated from the remote sensing image of the area to be evaluated. The attribute information includes the category, location and preliminary damage rating of the target object; the global descriptive information includes at least the attribute characteristics of each type of land parcel in the area to be evaluated. The construction module is used to construct the first input content for guiding the target model based on the attribute information, the global description information, and the preset knowledge base; The generation module is used to generate a global damage assessment report for the area to be assessed based on the first input content and the target model.
9. A computer device comprising a memory and a processor, the memory storing a computer program executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.