Geological disaster intelligent analysis method and system based on large-scale language model

By combining large-scale language models with anomaly detection and visual-language alignment techniques, geological disaster analysis texts are automatically generated, solving the problems of manual reliance and model deployment in existing technologies, and realizing automated and standardized identification and analysis of geological disasters.

CN121456331AActive Publication Date: 2026-02-03安徽明生恒卓科技有限公司 +1
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202511313032.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-15
Publication Date
2026-02-03
Estimated Expiration
2045-09-15

AI Technical Summary

Technical Problem

Existing geological disaster detection methods rely on human expert judgment, which is inefficient, lacks standardization, makes it difficult to generate structured text, and high-precision models are not suitable for real-time deployment under emergency conditions.

Method used

By employing a large-scale language model combined with the anomaly detection model PatchCore, the visual-language alignment model Q-Former, and a fully connected network, analytical text that conforms to the standards of the geological disaster field is automatically generated, and disaster detection and analysis are carried out using UAV or satellite imagery.

Benefits of technology

It enables automated and standardized identification and analysis of geological hazards, generates structured text, and is applicable to various types of geological hazard scenarios, thereby improving the efficiency and accuracy of emergency response.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121456331A_ABST
    Figure CN121456331A_ABST
Patent Text Reader

Abstract

The invention relates to a geological disaster intelligent analysis method and system based on a large-scale language model in the field of artificial intelligence, remote sensing monitoring and geological disaster prevention and control, and the method is used for inputting a to-be-processed image in a target area into a trained geological disaster analysis model. Outputting a geological disaster analysis text and a visual disaster detection image which conform to geological disaster field specifications; the geological disaster analysis model comprises a data preprocessing module; an anomaly detection model PatchCore is established; a visual language alignment model Q-Former is established; a fully connected network; according to the method, the problems that geological disaster samples are scarce and marking cost is high are effectively relieved by introducing the anomaly detection model, the automation and generalization ability of disaster recognition is improved, and the technical problems that an existing disaster judgment system is low in judgment efficiency, prone to being affected by subjectivity and lack of automatic text analysis are solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence, remote sensing monitoring and geological disaster prevention and control, and particularly relates to a geological disaster intelligent analysis method based on a large-scale language model and a geological disaster intelligent analysis system based on a large-scale language model. BACKGROUND

[0002] Geological disasters (such as landslides, debris flows, and rock avalanches) have the characteristics of strong suddenness, great destructive power, and wide influence, and pose a serious threat to people's life and property safety and the operation of infrastructure. In order to improve the ability to prevent and reduce disasters, government departments and research institutions widely use remote sensing technology for disaster monitoring and assessment. Remote sensing images have become an important data source for geological disaster monitoring due to their large coverage, fast update speed, and multi-source availability (such as optical, radar, and infrared).

[0003] Traditional geological disaster detection methods rely on machine vision and deep learning technology [1] , and the typical process is as follows: first, a suspected disaster area is extracted using an image segmentation model, and then an artificial expert analyzes the disaster type, range, and damage degree based on the image and segmentation results. This method has improved the automation level of disaster identification to some extent, but still has obvious shortcomings in practical application: first, it is highly dependent on artificial labor and low in efficiency, requiring experts to analyze the segmentation results one by one, resulting in a large workload and slow response speed; second, the results are not standardized enough, and different experts have different standards for identifying the same disaster, affecting the consistency of emergency command; third, there is a lack of automated expression, making it difficult to directly generate structured and machine-readable disaster analysis texts, thereby limiting seamless integration with emergency dispatch systems; fourth, deployment is limited, and high-precision deep learning models have high computational overhead, making them unsuitable for real-time deployment and reasoning on edge devices under emergency conditions.

[0004] In view of the above problems, in recent years, academia and industry have carried out a lot of research [2] in the field of remote sensing image intelligent analysis and visual language fusion. For example, large-scale language models (LLM) have made significant breakthroughs in natural language understanding and generation, and multi-modal models (Qwen-VL, Qwen2-VL) can handle both image and text inputs, have cross-modal information fusion and reasoning capabilities, and perform well in knowledge question answering and image description tasks [3][4]However, the application of these technologies in geological disaster scenarios is still limited to general visual tasks, lacking deep fusion of remote sensing images and disaster segmentation results and professional adaptation. Existing researches mostly focus on disaster area identification or simple text description, failing to achieve disaster judgment and automatic generation of standardized analysis text, let alone solving the problem of lightweight deployment of high-precision models under emergency conditions. The present application proposes an intelligent question and answer system for geological disasters based on large-scale language models. The system first processes remote sensing images and unmanned aerial vehicle inspection images using an anomaly detection model to generate disaster detection results and pixel-level feature information. Then, the pixel-level feature information is input as prompt information into a public visual language model (MiniGPT-4) [5] to automatically generate text descriptions that conform to the specifications of the geological disaster field, thereby constructing a data set of images and texts. On this basis, the image features and generated text semantic information are jointly input into a visual language alignment module to achieve accurate fusion and alignment of cross-modal features. Further, a fully connected network and a large-scale language model are combined for fine-tuning, enabling the model to accurately understand the fine-grained semantic features of geological disasters. Ultimately, the system can not only achieve automatic identification and classification of multiple types of geological disasters, but also generate structured disaster analysis text and support intelligent question and answer for disaster scenarios. Through lightweight network design and effective fusion of domain knowledge, the system can efficiently run on edge devices while ensuring accuracy and professionalism, thereby significantly improving the automation, standardization, and intelligence level of geological disaster detection and analysis.

[0005] References:

[0006] [1]. Roth K, Pemula L, Zepeda J, et al. Towards total recall in industrial anomaly detection [C] / / Proceedings of the IEEE / CVF conference on computer vision and pattern recognition. 2022: 14318-14328.

[0007] [2]. Li J, Li D, Xiong C, et al. Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation [C] / / International conference on machine learning. PMLR, 2022: 12888-12900.

[0008] [3]. Li C, Li Z, Jing C, et al. Searchlvlms: A plug-and-play framework for augmenting large vision-language models by searching up-to-date internet knowledge [J]. Advances in Neural Information Processing Systems, 2024, 37:64582-64603.

[0009] [4]. Yu Y, Shi C, Tang J, et al. Qwen-VL2 Model with NEFTune technique For Medical Report Generation [C] / / 2025 4th International Symposium on Computer Applications and Information Technology (ISCAIT). IEEE, 2025:165-168.

[0010] [5]. Zhu D, Chen J, Shen X, et al. Minigpt-4: Enhancing vision-language understanding with advanced large language models [J]. arXiv preprint arXiv:2304.10592, 2023. SUMMARY

[0011] (1) Technical problems to be solved

[0012] In order to solve the technical problems of low discrimination efficiency, easy to be affected by subjectivity and lack of automatic text analysis of the existing disaster discrimination system, the present application provides a geological disaster intelligent analysis method and system based on a large-scale language model.

[0013] (2) Technical solutions

[0014] In a first aspect, the present application discloses a geological disaster intelligent analysis method based on a large-scale language model, which inputs the to-be-processed image in a target area into a trained geological disaster analysis model, and outputs a geological disaster analysis text conforming to the specification of the geological disaster field and a visual disaster detection image; the geological disaster analysis model comprises:

[0015] a data preprocessing module configured to process the image to be processed into a format image for calibration;

[0016] an anomaly detection model PatchCore configured to output pixel-level features and a visualized disaster detection image according to the format image;

[0017] a visual language alignment model Q-Former configured to output visual features aligned with disaster information according to the pixel-level features;

[0018] a fully connected network configured to output a soft prompt projection vector according to the visual features;

[0019] a large-scale language model Qwen2.0 configured to output the geological disaster analysis text according to the soft prompt projection vector and a calibrated prompt template.

[0020] As an improvement of the above scheme, the data preprocessing method of the data preprocessing module comprises:

[0021] scaling the image to be processed in size, so that the short side is scaled to a calibrated length in proportion;

[0022] performing center cropping on the scaled image to obtain an image region with a fixed resolution;

[0023] performing color space unification and data type conversion on the cropped image to ensure that the image is in a three-channel floating-point format;

[0024] performing pixel value normalization on the converted image to scale the pixel value to the interval [0, 1, 1, 1];

[0025] on the basis of the normalization result, normalizing the image according to the calibrated channel mean and standard deviation to make it conform to the model input standard;

[0026] converting the normalized image into a tensor form and adjusting the channel order to generate a calibrated format image conforming to the input specification of the anomaly detection model;

[0027] As an improvement of the above scheme, the anomaly detection method of the anomaly detection model PatchCore comprises:

[0028] inputting the calibrated format image into a feature extraction network to extract multi-level deep feature representations;

[0029] performing dimension reduction and feature embedding processing on the extracted deep feature representations to reduce redundancy and highlight key features;

[0030] comparing the processed key features with a core feature library constructed from normal samples to generate an anomaly score corresponding to each pixel position, thereby obtaining pixel-level features of the anomaly detection image;

[0031] Thresholding and spatial smoothing are performed on pixel-level features to obtain the detection results of abnormal regions;

[0032] The detection results are overlaid on the original image to generate a visualized disaster detection image.

[0033] As an improvement to the above scheme, the visual language alignment method of the Q-Former visual language alignment model includes:

[0034] Pixel-level features are input into the visual-language alignment model Q-Former, and contextual features of each pixel-level feature are extracted through a self-attention mechanism.

[0035] The extracted contextual features are interactively calculated with the trained disaster domain cue vector to achieve the association between pixel-level features and disaster domain knowledge;

[0036] The features obtained from interactive computation are weighted, fused, and linearly mapped to generate a low-dimensional and compact visual representation.

[0037] The visual representation is normalized to obtain visual features aligned with disaster information.

[0038] As an improvement to the above scheme, the data processing method for fully connected networks includes:

[0039] Visual features aligned with disaster information are input into the input layer of a fully connected network and linearly mapped to adjust the feature dimensions.

[0040] The mapped features are subjected to nonlinear activation processing to enhance their expressive power.

[0041] The activated features are subjected to layer-by-layer linear transformation and normalization to form a compact feature vector;

[0042] After iterative processing through multiple fully connected layers, the feature vectors are finally linearly projected.

[0043] Based on the projection results, soft cue projection vectors are generated for large-scale language models, which can be used to guide the generation of geological disaster analysis text.

[0044] The intelligent analysis method for geological hazards, and the data processing method of the large-scale language model Qwen2.0, include:

[0045] Input the soft cue projection vector and the calibrated cue template into the large-scale language model Qwen2.0;

[0046] In the model's encoding layer, the soft cue projection vector is fused with the cue template as contextual information to form an enhanced initial semantic representation;

[0047] Under the action of self-attention and cross-layer information interaction, the enhanced initial semantic representation is gradually contextually expanded and feature transferred;

[0048] Through the decoder layer, the contextually expanded semantic representation is converted into serialized text features;

[0049] During the generation process, the structured information in the prompt template and the domain knowledge constraints are combined to predict and select each output token;

[0050] Finally, the geological disaster analysis text conforming to the geological disaster field specification is output.

[0051] As an improvement of the above scheme, the training method of the geological disaster analysis model comprises:

[0052] Collect sample images of disaster areas, including normal images before the disaster area occurs as a training set and abnormal images after the disaster occurs as a test set;

[0053] Abnormal detection model PatchCore training: for normal image data, the abnormal detection model PatchCore is trained to realize automatic detection of abnormal areas under the condition of no manual annotation. Through this process, pixel-level features corresponding to abnormal images and their visual detection results can be obtained;

[0054] Further, the obtained pixel-level features of the abnormal images are input into the pre-trained MiniGPT-4 model to generate corresponding text descriptions, thereby constructing a "picture-text pair" data set of abnormal images to avoid relying on large-scale manual text annotation.

[0055] Visual language alignment model Q-Former training: the "picture-text pair" data set is used as a training sample to train the visual language alignment model Q-Former; specifically, the pixel-level features of the abnormal images and the initialized prompt vector are fused through cross-attention mechanism to obtain a visual prompt vector containing disaster features; then, the visual prompt vector and the corresponding text features are aligned, and are optimized through a contrast loss, thereby obtaining a disaster field prompt vector with strong semantic correlation in the disaster field;

[0056] The joint training of the visual language alignment model Q-Former and the large-scale language model Qwen2.0: after obtaining the disaster field prompt vector, the disaster field prompt vector is jointly trained with the large-scale language model Qwen2.0; specifically, the pixel-level features are fused with the disaster field prompt vector optimized by the first training again through the cross-attention mechanism to generate low-dimensional and compact visual representations; after the visual representations are normalized, the visual features aligned with the disaster semantics are obtained, and the soft prompt projection vector is generated by inputting the visual features into the full connection network; then, the soft prompt projection vector and the calibrated prompt template are input into the large-scale language model Qwen2.0 to generate the geological disaster analysis text, and based on the similarity constraint between the generated text and the reference text in the “image-text pair” data set, the Q-Former and the full connection network are further optimized to improve the generation accuracy of the model in the disaster analysis task.

[0057] As an improvement of the above scheme, the image to be processed can be collected through satellite images or unmanned aerial vehicle inspection images or aerial photography images, and the geological disaster analysis text includes disaster types, spatial distribution, disaster-affected range, potential impact and danger level.

[0058] In a second aspect, the present application discloses a geological disaster intelligent analysis system based on a large-scale language model, which uses the geological disaster intelligent analysis method disclosed in the first aspect.

[0059] (3) Advantageous effects

[0060] 1. The present application can automatically identify potential abnormal areas and generate pixel-level feature information by introducing an anomaly detection model for unsupervised feature modeling of disaster images, effectively alleviating the problems of scarce geological disaster samples and high labeling cost, improving the automation and generalization ability of disaster identification, and solving the problems of low discrimination efficiency, easy subjective influence and lack of automatic text analysis technology of the existing disaster discrimination system.

[0061] 2. The present application automatically generates text description of the disaster area by inputting the pixel-level features of the disaster image into the visual language model MiniGPT-4, thereby constructing the “image-text pair” data set of the geological disaster image. This method does not require manual labeling and can efficiently generate standardized training corpus under low-cost conditions, effectively alleviating the problem of scarce geological disaster samples and improving the consistency and reliability of cross-disaster and multi-scene analysis.

[0062] 3. The present application realizes the alignment learning of disaster image features and disaster text features through the visual language alignment module Q-Former, and accesses the frozen large-scale language model Qwen2.0 under the support of a learnable full connection network, so as to fine-tune the identification and analysis process of the geological disaster image, thereby establishing a high-precision mapping relationship between multi-modal features and semantics. This method not only improves the identification accuracy of the model under complex terrain conditions, but also realizes fine-grained feature representation and deep semantic understanding of the disaster area.

[0063] 4. The geological disaster intelligent analysis system constructed by the present application can automatically generate structured analysis text conforming to the field specification for the input geological disaster image, and can be applied to multiple types of geological disaster scenes such as landslides, debris flows, and collapses, realizing the systematization and automation of disaster identification and analysis, and having wide engineering application value. BRIEF DESCRIPTION OF DRAWINGS

[0064] Figure 1 is a module diagram of the geological disaster intelligent analysis model based on a large-scale language model.

[0065] Figure 2 is an image preprocessing method flowchart.

[0066] Figure 3 is an anomaly detection method flowchart of the anomaly detection model PatchCore.

[0067] Figure 4 is a visual language alignment method flowchart of the visual language alignment model Q-Former.

[0068] Figure 5 is a data processing method flowchart of the full connection network.

[0069] Figure 6 is a data processing method flowchart of the large-scale language model Qwen2.0.

[0070] Figure 7 is a geological disaster analysis model training flowchart. DETAILED DESCRIPTION

[0071] The technical solutions in the embodiments of the present application will be described clearly and completely below. Obviously, the described embodiments are only part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.

[0072] It should be understood that when an element, or components, is referred to as being "on" another element, or components, it can be directly on another element or intervening elements can also be present. In contrast, when an element is referred to as being "directly on" another element, there are no intervening elements present. It will be understood that when an element is referred to as being "coupled" or "connected" to another element, it can be directly coupled or connected to the other element or intervening elements can be present. In contrast, when an element is referred to as being "directly coupled" or "directly connected" to another element, there are no intervening elements present.

[0073] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used in the description herein is for describing particular embodiments only and is not intended to be limiting of the application. As used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises" and / or "comprising", or "includes" and / or "including" when used herein, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0074] The present application provides a geological disaster intelligent analysis method based on a large-scale language model. The operation steps of the geological disaster intelligent analysis method are as follows: inputting the to-be-processed image in a target area into a trained geological disaster analysis model, and outputting a geological disaster analysis text conforming to the specification of the geological disaster field and a visualized disaster detection image. The geological disaster analysis method can be applied to real-time monitoring, disaster assessment and emergency decision support of geological disasters. Please refer to Figure 1 , Figure 1 The module diagram of the geological disaster intelligent analysis model is shown in the figure. The model includes:

[0075] I. An image acquisition module for acquiring a to-be-processed image in a target area;

[0076] II. A data preprocessing module for processing the to-be-processed image into a calibrated format image;

[0077] III. An anomaly detection model PatchCore for outputting pixel-level features and visualized disaster detection images according to the format image;

[0078] IV. A visual language alignment model Q-Former for outputting visual features aligned with disaster information according to the pixel-level features;

[0079] V. A fully connected network for outputting a soft prompt projection vector according to the visual features;

[0080] VI. A large-scale language model Qwen2.0 for outputting the geological disaster analysis text according to the soft prompt projection vector and the calibrated prompt template.

[0081] The application significantly improves the recognition accuracy of the model in complex terrain and various disaster scenes, reduces the misjudgment and omission caused by background interference and sample imbalance, and solves the technical problems of the existing disaster discrimination system, such as relying on artificial experience, lacking cross-modal semantic understanding, and insufficient robustness of new disaster recognition.

[0082] The modules will be introduced one by one as follows.

[0083] I. Image acquisition module

[0084] The target area of the image to be processed is collected, and the data sources can include satellite remote sensing platforms, unmanned aerial vehicle aerial photography systems and aerial photography equipment. In this embodiment, the system obtains the image to be processed from satellite, unmanned aerial vehicle, aerial photography and other multi-source sensing platforms.

[0085] II. Data preprocessing module

[0086] The image to be processed is first subjected to preprocessing operation, including but not limited to geometric correction, radiation correction, denoising and color balance, to eliminate distortion and brightness difference caused by different sensors and imaging conditions, and to ensure the accuracy of subsequent analysis. In this embodiment, due to the difference in imaging conditions of different sensors, the image to be processed may contain geometric distortion, radiation deviation and noise. Through geometric correction, radiation correction, denoising and color balance, the system standardizes the image to be processed, so that the image to be processed is converted into a calibrated format image, i.e. the image to be processed is converted into an image that can be received by the anomaly detection model. Thus, the stability and reliability of subsequent feature extraction are ensured. The principle of this stage is to eliminate the interference of non-disaster factors by using image processing algorithm, so that the image received by the model remains consistent in statistical features. Please refer to Figure 2 , Figure 2 The flowchart of the image preprocessing method, the data preprocessing method comprises:

[0087] The image to be processed is scaled in size, and the short side is scaled to a preset length in proportion;

[0088] The scaled image is center cropped to obtain an image region with a fixed resolution;

[0089] The cropped image is subjected to color space unification and data type conversion to ensure that the image is in three-channel floating point format;

[0090] The converted image is subjected to pixel value normalization to scale the pixel value to the interval of 0, 10, 10, 1;

[0091] On the basis of the normalization result, the standardization processing is performed according to the preset channel mean and standard deviation, so as to conform to the model input distribution;

[0092] The standardized image is converted into a tensor form and the channel order is adjusted, so as to generate a calibrated format image conforming to the input specification of the anomaly detection model.

[0093] III. Anomaly detection model.

[0094] The calibrated format image is input into the anomaly detection model, and the anomaly detection model realizes automatic identification of potential abnormal geological image regions by constructing a regional feature distribution of normal geological images and detecting deviation patterns. The output result includes pixel-level feature information and visual disaster images, and labels the location, boundary and morphological features of potential disasters. This step effectively solves the problems of high cost of high-quality labeling and strong dependence on manual experience, and lays a foundation for full-process automatic analysis. The anomaly detection model in this embodiment adopts PatchCore. Based on the principle of high-dimensional feature representation and nearest neighbor retrieval, PatchCore can automatically identify potential abnormal regions in calibration. Specifically, PatchCore first extracts features from the input normal geological image, constructs a local feature library, and then performs nearest neighbor matching on the feature library samples to calculate the anomaly score of each pixel or region, thereby generating pixel-level feature information and visual disaster detection images. Please refer to Figure 3 , Figure 3 The anomaly detection method of the anomaly detection model PatchCore includes:

[0095] The format image is input into a pre-trained feature extraction network to extract multi-level deep feature representations;

[0096] The extracted features are dimensionally reduced and embedded to reduce redundancy and highlight key features;

[0097] In the feature space after dimension reduction, the processed features of the format image are compared with the core feature library constructed based on existing normal samples for similarity, so as to generate an anomaly score corresponding to each pixel position, and obtain pixel-level features of the anomaly detection image;

[0098] The pixel-level features are threshold segmented and spatially smoothed to obtain the detection result of the abnormal region;

[0099] The detection result is superimposed and rendered with the original image to generate a visual disaster detection image.

[0100] IV. Visual language alignment model Q-Former.

[0101] The visual language alignment model can convert pixel-level feature information into visual features aligned with disaster information. In this embodiment, the visual language alignment module adopted is Q-Former. Please refer to Figure 4 ,Figure 4 A visual language alignment method flowchart for a visual language alignment model Q-Former, the visual language alignment method of the visual language alignment model Q-Former comprising:

[0102] Inputting the pixel-level features into the visual language alignment model Q-Former, and extracting the context information of each pixel through a self-attention mechanism;

[0103] Interactively calculating the extracted context features and the trained disaster field prompt vector, to realize the association of the pixel-level features and the disaster field knowledge;

[0104] Performing weighted fusion and linear mapping on the features obtained through the interactive calculation, to generate a low-dimensional and compact visual representation;

[0105] Performing normalization processing on the visual representation, to obtain visual features aligned with disaster information.

[0106] Five, a fully connected network.

[0107] In this embodiment, a learnable fully connected network is used to convert the visual features aligned with the disaster information into a soft prompt vector, and the fully connected network is connected with the Q-Former and a large-scale language model respectively. The fully connected network converts the visual disaster information generated by the Q-Former into a soft prompt vector, and transmits the soft prompt vector to the large-scale language model. Please refer to Figure 5 , Figure 5 A data processing method flowchart for the fully connected network, the data processing method of the fully connected network comprising:

[0108] Inputting the visual features aligned with the disaster information into the input layer of the fully connected network, and performing linear mapping on it to adjust the feature dimension;

[0109] Performing nonlinear activation processing on the mapped features, to enhance the expression ability of the features;

[0110] Performing layer-by-layer linear transformation and normalization processing on the activated features, to form a compact feature vector;

[0111] After the iterative processing of the multiple fully connected layers, performing final linear projection on the feature vector;

[0112] Generating a soft prompt projection vector for the large-scale language model according to the projection result, which can be used to guide the generation of the geological disaster analysis text.

[0113] Six, a large-scale language model Qwen2.0.

[0114] Large-scale language models can generate structured analysis text that meets the specifications of the geological disaster field based on the soft prompt projection of the input, under the control of the calibrated prompt template. The text content includes disaster type, spatial distribution, disaster area, potential impact, and danger level, etc., and can generate a brief report or detailed analysis report. The final generated analysis text can be directly applied to emergency command systems, risk assessment platforms, and disaster warning systems, realizing intelligent identification, professional interpretation, and rapid response of geological disasters.

[0115] In this embodiment, the prompt template is an input text that can be input in a manual interactive manner or fixedly written, for example, a value of 0 represents a non-flood area, which can be regarded as a background; a value of 1 represents a flood area, and output disaster cause elements and disaster impact analysis text.

[0116] This embodiment uses Qwen2.0 as a large-scale language model, and maps visual features to language understanding space through a multi-modal interface to realize comprehensive analysis of geological disasters. The model can perform semantic reasoning based on image pattern information and geological disaster field knowledge to identify disaster types, assess disaster areas, judge potential impact areas, and analyze danger levels. At the same time, the spatial constraint information provided by the pixel-level feature information makes the discrimination result more accurate, significantly improving the robustness and reliability of the model in complex terrain and diverse disaster scenarios. Please refer to Figure 6 , Figure 6 The data processing method of the large-scale language model Qwen2.0 includes:

[0117] The soft prompt projection vector and the calibrated prompt template are input into the large-scale language model Qwen2.0;

[0118] In the encoding layer of the model, the soft prompt projection vector is fused with the prompt template as context information to form an enhanced initial semantic representation;

[0119] Under the action of self-attention and cross-layer information interaction, the fused semantic representation is gradually contextually expanded and feature transferred;

[0120] Through the decoder layer, the contextually expanded semantic representation is converted into a serialized text feature;

[0121] During the generation process, the structured information in the prompt template and the domain knowledge constraints are combined to predict and select each output token;

[0122] Finally, the geological disaster analysis text that meets the specifications of the geological disaster field is output.

[0123] The original Qwen2.0 is a large-scale language model framework open-sourced by Ali Cloud, which belongs to the public prior art. Its basic structure, training method and model interface have been disclosed in public channels (such as GitHub and technical documents). The present application optimizes the structure and function on the basis of the original Qwen2.0. Specifically, at the input end, a learnable fully connected network is added to map the fine-grained disaster feature vector output by Q-Former to a unified semantic space, realizing the fusion of visual features and language semantics; at the output end, a disaster analysis task head is added for disaster type classification, disaster area assessment, potential impact area analysis and danger level determination. Through these optimizations, the model can efficiently convert multi-modal features into structured and professional geological disaster analysis text, realizing accurate identification and comprehensive evaluation of geological disasters.

[0124] The framework of Qwen2.0 as a whole can be regarded as a highly modular large-scale language model system. It takes the basic language model as the core, covering different scales from hundreds of millions of parameters to hundreds of billions of parameters (such as 7B, 13B, 72B, etc.), which can handle multi-language text understanding and generation tasks. Through pre-training, these basic models have mastered rich language knowledge and semantic relationships, and in the instruction fine-tuning stage, the model further learns how to generate high-quality responses according to specific task requirements, such as question answering, dialogue, abstract generation, etc. This hierarchical training method makes the model have both broad general ability and excellent performance on specific tasks.

[0125] In addition to basic language capabilities, Qwen2.0 also introduces multi-modal processing capabilities (Qwen2-VL) that can accept visual inputs such as images and videos, and combine them with text information for understanding and generation. For example, it can analyze picture content to answer questions, or generate corresponding image information based on text descriptions, which has significant advantages in visual question answering, document understanding, image content generation, etc. At the same time, Qwen2.0 adopts the Mixture-of-Experts (MoE) architecture, which dynamically selects some experts to participate in calculation during inference, thereby expanding the model capacity while ensuring computational efficiency, so that large-scale models can also run efficiently on limited hardware resources.

[0126] The entire usage process starts from the input end, whether it is pure text or multi-modal information, which is first converted into Token or feature sequence that the model can process through a unified segmentation and coding interface (such as AutoTokenizer), and then the prediction or generation result is obtained through the model calculation. Model configuration (AutoConfig) and loading (AutoModel) provide flexible parameter management and device adaptation capabilities, so that the same model can run efficiently on CPU, GPU or multi-node cluster. Finally, through high-performance deployment framework (such as vLLM) or application integration tool (such as LangChain), these model capabilities are applied to actual scenarios to realize question and answer systems, dialogue robots, document analysis, multi-modal content understanding and other functions.

[0127] Through the organic cooperation of the above modules, the geological disaster analysis text generation process can realize the deep fusion and professional expression of multi-modal information. The geological disaster analysis model used in this embodiment is a trained geological disaster analysis model. Please refer to Figure 7 , Figure 7 The training flowchart of the geological disaster analysis model, the training method of the geological disaster analysis model comprises:

[0128] Collect sample images of disaster areas, including normal images before the occurrence of disaster areas as a training set and abnormal images after the occurrence of disasters as a test set;

[0129] Abnormal detection model PatchCore training: for normal image data, the abnormal detection model PatchCore is trained to realize automatic detection of abnormal areas without manual annotation. Through this process, the pixel-level features corresponding to the abnormal images and their visual detection results can be obtained.

[0130] Further, the obtained pixel-level features of the abnormal images are input into the pre-trained MiniGPT-4 model to generate the corresponding text description, thereby constructing the “image-text pair” data set of the abnormal images to avoid relying on large-scale manual text annotation.

[0131] Visual language alignment model Q-Former training: after obtaining the “image-text pair” data set, it is used as a training sample to train the visual language alignment model Q-Former. Specifically, the pixel-level features of the image and the initialized prompt vector are fused through the cross-attention mechanism to obtain a visual prompt vector containing disaster features; then, the visual prompt vector and the corresponding text features are aligned, and the contrast loss is optimized to obtain a disaster field prompt vector with strong semantic correlation in the disaster field.

[0132] Joint training of visual language alignment model Q-Former and large-scale language model Qwen2.0: After obtaining the disaster field prompt vector, it is jointly trained with the large-scale language model Qwen2.0. Specifically, the pixel-level features and the disaster field prompt vector optimized by the first training are fused again through the cross-attention mechanism to generate low-dimensional and compact visual representations; after normalization processing of the visual representations, the visual features aligned with the disaster semantics are obtained, and the soft prompt projection vector is generated by inputting the full connection network. Then, the soft prompt projection vector and the calibrated prompt template are input into the large-scale language model Qwen2.0 to generate the geological disaster analysis text, and based on the similarity constraint between the generated text and the reference text in the "image-text pair" data set, the Q-Former and the full connection network are further optimized to improve the generation accuracy of the model in the disaster analysis task.

[0133] Based on the geological disaster intelligent analysis method based on the large-scale language model proposed in the above embodiment, the present application further proposes a geological disaster intelligent analysis system based on a large-scale language model, which uses the geological disaster intelligent analysis method based on a large-scale language model.

[0134] In summary, the key innovations of the present application are as follows:

[0135] 1. Abnormality detection and image-text pair generation, the potential disaster area in the target area is identified by the abnormality detection model PatchCore, and the pixel-level features are input into the visual language model MiniGPT-4 to automatically generate corresponding text description and build an image-text pair data set to provide high-quality corpus for multi-modal training.

[0136] 2. Multi-modal feature fusion, the visual language alignment module Q-Former is used to train the image-text pair data to realize the alignment of fine-grained visual features and text semantics, generate semantic consistent disaster feature vectors, and improve the accuracy and robustness of cross-modal information fusion.

[0137] 3. Fine-tuning of large-scale language model, the trained Q-Former output is connected to the frozen large-scale language model Qwen2.0 through a learnable full connection network, and fine-tuning is performed to realize unified mapping of visual features and language semantics and optimize disaster identification and analysis performance.

[0138] 4. Structured disaster text generation, based on the fine-tuned large-scale language model, the multi-modal fusion disaster feature vector is converted into a structured analysis text conforming to the geological disaster field specification, including disaster type, spatial distribution, disaster range, potential impact and danger level, to realize automatic and standardized professional analysis report generation.

[0139] Any combination of the technical features in the above-described embodiments can be made, and for the sake of brevity, not all possible combinations are described, however, as long as there is no conflict, any combination of the technical features should be considered within the scope of the present disclosure.

[0140] The above-described embodiments only express several implementation manners of the present application, and the description is relatively specific and detailed, but it should not be understood as a limitation on the scope of the patent. It should be pointed out that for ordinary skilled persons in the art, some modifications and improvements can be made without departing from the concept of the present application, and these are within the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.

Claims

1. A method for intelligent analysis of geological hazards based on a large-scale language model, characterized in that, The image to be processed within the target area is input into the trained geological hazard analysis model, which outputs geological hazard analysis text that conforms to the standards of the geological hazard field and visualized hazard detection images. Geological hazard analysis models include: The data preprocessing module is used to process the image to be processed into a calibrated image format; The PatchCore anomaly detection model is used to output pixel-level features and visualized disaster detection images based on formatted images. The visual language alignment model Q-Former is used to output visual features aligned with disaster information based on pixel-level features; A fully connected network is used to output soft cue projection vectors based on visual features; The large-scale language model Qwen2.0 is used to output the geological disaster analysis text based on soft cue projection vectors and calibrated cue templates.

2. The intelligent geological disaster analysis method based on a large-scale language model according to claim 1, characterized in that, The data preprocessing methods in the data preprocessing module include: The image to be processed is resized so that its shorter side is proportionally scaled to the specified length. Perform a center crop on the scaled image to obtain a fixed-resolution image region; Perform color space unification and data type conversion on the cropped image to ensure that the image is in three-channel floating-point format; The converted image is normalized by scaling the pixel values ​​to the range of 0, 10, 10, 1. Based on the normalization results, the image is standardized according to the calibrated channel mean and standard deviation to make it conform to the model input standard; The standardized image is converted into tensor form, and the channel order is adjusted to generate a calibrated image that conforms to the input specifications of the anomaly detection model.

3. The intelligent geological disaster analysis method based on a large-scale language model according to claim 1, characterized in that, The anomaly detection methods of the PatchCore anomaly detection model include: The calibrated image is input into the feature extraction network to extract multi-level deep feature representations; The extracted deep feature representations are subjected to dimensionality reduction and feature embedding to reduce redundancy and highlight key features; The processed key features are compared with the core feature library constructed from normal samples to generate an anomaly score corresponding to each pixel position, thus obtaining the pixel-level features of the anomaly detection image. Thresholding and spatial smoothing are performed on pixel-level features to obtain the detection results of abnormal regions; The detection results are overlaid on the original image to generate a visualized disaster detection image.

4. The intelligent geological disaster analysis method based on a large-scale language model according to claim 1, characterized in that, The visual-language alignment method of the Q-Former visual-language alignment model includes: Pixel-level features are input into the visual-language alignment model Q-Former, and contextual features of each pixel-level feature are extracted through a self-attention mechanism. The extracted contextual features are interactively calculated with the trained disaster domain cue vector to achieve the association between pixel-level features and disaster domain knowledge; The features obtained from interactive computation are weighted, fused, and linearly mapped to generate a low-dimensional and compact visual representation. The visual representation is normalized to obtain visual features aligned with disaster information.

5. The intelligent geological disaster analysis method based on a large-scale language model according to claim 1, characterized in that, Data processing methods for fully connected networks include: Visual features aligned with disaster information are input into the input layer of a fully connected network and linearly mapped to adjust the feature dimensions. The mapped features are subjected to nonlinear activation processing to enhance their expressive power. The activated features are subjected to layer-by-layer linear transformation and normalization to form a compact feature vector; After iterative processing through multiple fully connected layers, the feature vectors are finally linearly projected. Based on the projection results, soft cue projection vectors are generated for use in large-scale language models to guide the generation of geological disaster analysis text.

6. The intelligent geological disaster analysis method based on a large-scale language model according to claim 1, characterized in that, The data processing methods for the large-scale language model Qwen2.0 include: Input the soft cue projection vector and the calibrated cue template into the large-scale language model Qwen2.0; In the model's encoding layer, the soft cue projection vector is fused with the cue template as contextual information to form an enhanced initial semantic representation; With the help of self-attention and cross-layer information interaction, the enhanced initial semantic representation is progressively extended into context and features are passed on. The decoder layer converts the context-expanded semantic representation into serialized text features. During the generation process, the structured information in the prompt template and the domain knowledge constraints are combined to predict and select each output tag; Finally, the output is a geological hazard analysis text that conforms to the standards in the field of geological hazards.

7. The intelligent geological disaster analysis method based on a large-scale language model according to claim 1, characterized in that, Training methods for geological hazard analysis models include: Collect sample images of the disaster area, including normal images of the area before the disaster as the training set and abnormal images of the area after the disaster as the test set; PatchCore anomaly detection model training: For normal image data, the PatchCore anomaly detection model is trained to achieve automatic detection of abnormal regions without manual annotation, and can obtain pixel-level features and visualization detection results corresponding to abnormal images. Furthermore, the pixel-level features of the obtained anomalous images are input into the pre-trained MiniGPT-4 model to generate corresponding text descriptions, thereby constructing a dataset of "image-text pairs" of anomalous images to avoid relying on large-scale manual text annotation. Training the visual-language alignment model Q-Former: The visual-language alignment model Q-Former is trained using the "image-text pair" dataset as training samples. Specifically, the pixel-level features of the abnormal image are fused with the initialized cue vector through a cross-attention mechanism to obtain a visual cue vector containing disaster features. Subsequently, the visual cue vector is aligned with the corresponding text features and optimized through contrastive loss to obtain a disaster domain cue vector with strong semantic relevance in the disaster domain. Joint training of the visual-language alignment model Q-Former and the large-scale language model Qwen2.0: After obtaining the disaster domain cue vector, it is jointly trained with the large-scale language model Qwen2.

0. Specifically, pixel-level features and the disaster domain cue vector optimized by the first training are fused again through a cross-attention mechanism to generate a low-dimensional and compact visual representation. After normalization of this visual representation, visual features aligned with disaster semantics are obtained and input into a fully connected network to generate soft cue projection vectors. Subsequently, the soft cue projection vectors and the calibrated cue templates are input into the large-scale language model Qwen2.0 to generate geological disaster analysis text. Based on the similarity constraints between the generated text and the reference text in the "image-text pair" dataset, Q-Former and the fully connected network are further optimized, thereby improving the model's generation accuracy in disaster analysis tasks.

8. The intelligent geological disaster analysis method based on a large-scale language model according to claim 1, characterized in that, The images to be processed are acquired through satellite imagery, drone inspection images, or aerial photography.

9. The intelligent geological disaster analysis method based on a large-scale language model according to claim 1, characterized in that, Geological hazard analysis texts include hazard type, spatial distribution, affected area, potential impact, and hazard level.

10. A geological disaster intelligent analysis system based on a large-scale language model, characterized in that, It employs the intelligent geological disaster analysis method based on a large-scale language model as described in any one of claims 1-9.

Citation Information

Patent Citations

  • Industrial image anomaly detection method based on multi-modal large model

    CN119762891A

  • Remote sensing image surface anomaly detection method and device based on image-text cooperative processing

    CN120217263A

  • Steel bridge disease detection and identification method based on large language model

    CN120496073A

  • Data analysis problem generation method based on image input and large model combination

    CN120632138A

  • Disaster information fusion and semantic reasoning algorithm based on cross-modal graph attention mechanism

    CN120633837A