Geoscience big data element extraction model training and application method, device and medium

By calculating cross-modal entropy to optimize the pruning strategy of the geoscientific big data feature extraction model, the problems of semantic fragmentation and false positives and false negatives in geoscientific feature extraction are solved, and higher feature extraction completeness and accuracy are achieved.

CN122196559BActive Publication Date: 2026-08-04INSTITUTE OF GEOLOGY AND GEOPHYSICS CHINESE ACADEMY OF SCIENCES
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
INSTITUTE OF GEOLOGY AND GEOPHYSICS CHINESE ACADEMY OF SCIENCES
Filing Date
2026-05-15
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

Existing geoscientific big data feature extraction models are prone to semantic fragmentation when processing continuous geoscientific targets, resulting in breaks or missed detections. They are also sensitive to background interference, leading to a high false detection rate.

Method used

By acquiring geoscientific modal sample data and text task prompts, a visual language model is used to extract visual token sequences and text token sequences to calculate cross-modal entropy, determine the optimal clipping, and introduce cross-modal entropy as a loss function during training to optimize model parameters and improve the completeness and accuracy of feature extraction.

Benefits of technology

It reduces semantic fragmentation, improves the completeness and accuracy of feature extraction, reduces false detections and false negatives, and enhances the clarity of segmentation boundaries and the continuity of linear features.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122196559B_ABST
    Figure CN122196559B_ABST
Patent Text Reader

Abstract

This application discloses a method, device, and medium for training and applying a geoscientific big data feature extraction model, relating to the fields of geoscientific big data processing and artificial intelligence technology. The method includes: cropping geoscientific modal sample data to obtain several candidate cropping samples corresponding to each geoscientific modal sample data; extracting visual token sequences and text token sequences from the candidate cropping samples and text task prompts; calculating cross-modal entropy based on the visual token sequences and text token sequences, and determining the optimal cropping; inputting the geoscientific modal sample data, the optimal cropping, and the text task prompts into a large language model to obtain sample feature extraction results; and training a trained geoscientific big data feature extraction model using a total loss function, where the total loss function includes the cross-modal entropy based on the candidate cropping samples and the sample feature extraction results. This method avoids the occurrence of feature extraction fragmentation or missed detection problems, and also reduces background interference, thus reducing false positives or false negatives.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of geoscience big data processing and artificial intelligence technology, and in particular to a method, device and medium for training and applying a geoscience big data element extraction model. Background Technology

[0002] Geoscientific big data typically features "large format, high resolution, multi-source heterogeneity, and spatiotemporal multi-scale", such as satellite / aerial remote sensing imagery, DEM, geophysical raster, geological maps, and text reports. In practical applications, geoscientific elements (faults, folds, river networks, landslides, lithological contact zones, etc.) often exhibit a continuous distribution in long strips or patches, with a large scale span.

[0003] Current geoscientific analysis systems based on deep learning or large multimodal models generally employ fixed-patch segmentation mechanisms of visual encoders (such as ViT segmenting images into fixed-size patches). For lightweight models or models with limited input windows, fixed resolution and fixed pruning strategies are often used to ensure computational efficiency, resulting in: fragmented continuous geoscientific targets: for example, fault lines are segmented into multiple segments, geological boundaries are discontinuous, and text on the map is cut off; and the extraction results are sensitive to background interference: when early layer-to-modal alignment is insufficient, the model's attention is easily diverted to irrelevant areas, leading to false positives or false negatives. Summary of the Invention

[0004] The purpose of this application is to provide a method, device and medium for training and applying geoscientific big data feature extraction models, which can reduce semantic fragmentation, improve feature integrity from the source, and avoid the occurrence of feature extraction breaks or missed detections; it can also reduce background interference and reduce false detections or missed detections.

[0005] To achieve the above objectives, this application provides the following solution.

[0006] Firstly, this application provides a method for training a geoscientific big data feature extraction model, comprising: acquiring several geoscientific modality sample data and text task prompts corresponding to each geoscientific modality sample data; cropping each geoscientific modality sample data to obtain a sample candidate cropping set corresponding to each geoscientific modality sample data; each sample candidate cropping set including several sample candidate croppings; inputting the sample candidate croppings and text task prompts into the visual encoder and text encoder in a visual language model, respectively, to extract the visual token sequence of the sample candidate croppings and the text token sequence of the text task prompts; and calculating the sample... The cross-modal entropy corresponding to this candidate clipping is used to determine the optimal clipping for the geoscientific modality sample data based on the cross-modal entropy corresponding to all sample candidate clippings. The cross-modal entropy is used to measure the semantic integrity of the relevant region of the text task prompt in the sample candidate clipping. The geoscientific modality sample data, the optimal clipping corresponding to the geoscientific modality sample data, and the text task prompt are input into the large language model in the visual language model to extract geoscientific features and obtain the sample feature extraction results. The model parameters of the visual language model are updated using the total loss function to obtain the trained geoscientific big data feature extraction model. The total loss function includes the cross-modal entropy based on the sample candidate clipping and the sample feature extraction results.

[0007] Secondly, this application provides a method for applying a geoscientific big data feature extraction model, comprising the following steps: acquiring geoscientific modal target data and text task prompts for a target area; cropping the geoscientific modal target data to obtain a target candidate cropping set; each target candidate cropping set includes several target candidate croppings; inputting the target candidate croppings and text task prompts into a trained geoscientific big data feature extraction model to obtain the target feature extraction result; the visual encoder in the visual language model extracts the visual token sequence of the target candidate croppings, and the text encoder extracts the text task prompts. The text token sequence is used; the cross-modal entropy corresponding to the target candidate cropping is calculated based on the visual token sequence of the target candidate cropping and the text token sequence of the text task prompt; the optimal cropping corresponding to the geoscience modal target data is determined based on the cross-modal entropy corresponding to all target candidate croppings; the geoscience modal target data, the optimal cropping corresponding to the geoscience modal target data and the text task prompt are input into the large language model in the visual language model to obtain the target element extraction result; the trained geoscience big data element extraction model is the model trained using the above-mentioned geoscience big data element extraction model training method.

[0008] Thirdly, this application provides a computer device, including: a memory, a processor, and a computer program stored in the memory and capable of running on the processor, wherein the processor executes the computer program to implement the steps of the above-described geoscientific big data feature extraction model training method or the above-described geoscientific big data feature extraction model application method.

[0009] Fourthly, this application provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the above-described geoscientific big data feature extraction model training method or the above-described geoscientific big data feature extraction model application method.

[0010] According to the specific embodiments provided in this application, this application has the following technical effects: This application provides a method, device, and medium for training and applying a geoscientific big data feature extraction model. After inputting the sample candidate cropping and text task prompts into the visual encoder and text encoder in the visual language model, respectively, and extracting the visual token sequence of the sample candidate cropping and the text token sequence of the text task prompts, the cross-modal entropy corresponding to the sample candidate cropping is calculated based on the visual token sequence and the text token sequence. The optimal cropping corresponding to the geoscientific modality sample data is determined based on the cross-modal entropy corresponding to all sample candidate croppings, making the prompt-related areas more spatially continuous, reducing semantic fragmentation, and improving feature integrity from the source. By introducing a loss function through cross-modal entropy for training, the model's attention is more focused on key visual tokens, reducing background interference, reducing false detections or missed detections, and improving the clarity of segmentation boundaries, the continuity of linear features, and the accuracy of OCR. Attached Figure Description

[0011] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 This is an application environment diagram of a geoscientific big data feature extraction model training method or a geoscientific big data feature extraction model application method in the embodiments of this application.

[0013] Figure 2 This is a flowchart illustrating a geoscientific big data feature extraction model training method provided in Embodiment 1 of this application.

[0014] Figure 3 This is a schematic diagram illustrating the specific process of a geoscientific big data feature extraction model training method provided in Embodiment 1 of this application.

[0015] Figure 4 This is a schematic diagram of the structure of a computer device provided in Embodiment 3 of this application. Detailed Implementation

[0016] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0017] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0018] Example 1: The geoscientific big data feature extraction model training method provided in this application can be applied to, for example... Figure 1 In the application environment shown, terminal 102 communicates with server 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be set up independently, integrated into server 104, or placed in the cloud or on another server. Terminal 102 can send geoscientific modality sample data and text task prompts to server 104. After receiving the geoscientific modality sample data and text task prompts, server 104 performs cropping on the geoscientific modality sample data to obtain several candidate cropping samples corresponding to each geoscientific modality sample data. It extracts the visual token sequence of the candidate cropping samples and the text token sequence of the text task prompts, calculates the cross-modal entropy based on the visual token sequence and the text token sequence, determines the optimal cropping, and inputs the geoscientific modality sample data, the optimal cropping, and the text task prompts into a large language model to obtain the sample element extraction results. The model is then trained using the total loss function to obtain a trained geoscientific big data element extraction model. Server 104 can feed back the trained geoscientific big data element extraction model to terminal 102. In addition, in some embodiments, the geoscientific big data feature extraction model training method can also be implemented by the server 104 or the terminal 102 separately. For example, the terminal 102 can directly train the model based on the geoscientific modal sample data and text task prompts, or the server 104 can obtain the geoscientific modal sample data and text task prompts from the data storage system and train the model based on the geoscientific modal sample data and text task prompts.

[0019] The terminal 102 can be, but is not limited to, various desktop computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, and smart in-vehicle devices. Portable wearable devices can include smartwatches, smart bracelets, and head-mounted devices. The server 104 can be implemented using a standalone server or a server cluster composed of multiple servers, or it can be a cloud server.

[0020] In one exemplary embodiment, such as Figure 2 As shown, a method for training a geoscientific big data feature extraction model is provided. This method is executed by a computer device, specifically by a terminal or server alone, or by both a terminal and a server. In this embodiment, the method is applied to... Figure 1 Taking server 104 as an example, the explanation includes the following steps 201 to 206.

[0021] Step 201: Obtain several geoscientific modal sample data and text task prompts corresponding to each geoscientific modal sample data.

[0022] Step 202: Prune the sample data of each geoscience modality to obtain the sample candidate pruning set corresponding to each geoscience modality sample data; each sample candidate pruning set includes several sample candidate prunings.

[0023] Step 203: Input the sample candidate cropping and text task prompts into the visual encoder and text encoder in the visual language model, respectively, and extract the visual token sequence of the sample candidate cropping and the text token sequence of the text task prompts.

[0024] Step 204: Calculate the cross-modal entropy corresponding to the candidate cropping of the samples based on the visual token sequence and the text token sequence, and determine the optimal cropping corresponding to the geoscience modality sample data based on the cross-modal entropy corresponding to all candidate cropping of the samples; the cross-modal entropy is used to measure the semantic integrity of the relevant regions of the text task prompt in the candidate cropping of the samples.

[0025] Step 205: Input the geoscientific modality sample data, the optimal cropping corresponding to the geoscientific modality sample data, and the text task prompts into the large language model in the visual language model to extract geoscientific features and obtain the sample feature extraction results.

[0026] Step 206: Update the model parameters of the visual language model using the total loss function to obtain the trained geoscientific big data feature extraction model; the total loss function includes the cross-modal entropy based on the sample candidate pruning and the sample feature extraction results.

[0027] Steps 201 to 206 are performed as described above. The cross-modal entropy corresponding to candidate clippings is calculated based on the visual and text token sequences. The optimal clipping for the geoscientific modality sample data is then determined based on the cross-modal entropy corresponding to all candidate clippings. This makes the cue-related region more spatially continuous, reduces semantic fragmentation, and improves element integrity from the source. By introducing a loss function through cross-modal entropy during training, the model's attention is more focused on key visual tokens, reducing background interference, false positives or false negatives, and improving the clarity of segmentation boundaries, the continuity of linear elements, and OCR accuracy. Cross-modal entropy is used to measure the degree of semantic fragmentation of the "cue-related region"; the smaller the cross-modal entropy, the more spatially continuous and complete the cue-related region is.

[0028] like Figure 3 As shown, the training method can be divided into the following steps S101~S106.

[0029] The corresponding module can be divided into the following sub-modules: Geoscience Multi-Source Data Access and Preprocessing Module: Completes coordinate unification, resolution matching, cloud removal, noise reduction, normalization, and tile conversion for images / DEMs / vectors / texts; Candidate Crop Generation Module: Generates candidate cropping windows according to preset cropping ratio and scale sets, while maintaining alignment across modal spaces; Cross-Modal Entropy Calculation Module: Extracts visual and text tokens from early layers of the model, calculates the normalized attention probability distribution, and obtains cross-modal entropy; Adaptive Crop Selection Module: Sorts the cross-modal entropy of candidate croppings and selects the crop with the smallest (or nth smallest) cross-modal entropy as the optimal cropping (which can be merged with the original image); Feature Extraction and Enhancement Module: Performs segmentation / detection / OCR / vectorization based on the selected cropping, and performs boundary enhancement using attention heatmaps; Entropy Regularization Training Module: Uses a total loss function during fine-tuning to improve the robustness of key token identification and feature extraction; Stitching and Georegistration Output Module: Maps the cropping domain results back to the original coordinate system and stitches them together, outputting raster / vector features and reports.

[0030] Step S101 (Geoscience Data Access and Preprocessing): Geoscience modal sample data includes at least one or more spatial raster data, such as remote sensing image I (which may be multispectral / hyperspectral / SAR), digital elevation model H (DEM / DSM) and / or vector layer V (geological boundaries, faults, roads, water systems, etc.) and other geoscience sample data.

[0031] Preprocessing of geoscientific modal sample data includes, but is not limited to, coordinate unification, resolution matching, cloud removal and noise reduction, normalization and / or tile conversion.

[0032] Coordinate unification: Projection and affine transformations are performed on I, H, and V to obtain a unified spatial reference; Resolution / scale matching: Data at different resolutions is resampled to align multimodal data on the cell grid; Quality control: Cloud / shadow detection and masking, noise suppression, and radiometric normalization; Large-format processing: Pyramid construction or tile indexing is performed on ultra-large images to support rapid reading of subsequent candidate cropping.

[0033] Step S102, Generate candidate cropping (scale / proportion): Based on the preset cropping scale set R={r1, r2, ..., rQ} and scale set S={s1, s2, ..., sP}, generate cropping at the same center point or on the candidate ROI set. There are 10 candidate clipping windows. Each window corresponds to a clipped multimodal sub-block X_{q,p} (denoted as the sample candidate clipping), and maintains the spatial consistency (i.e., aligned clipping) between the remote sensing image I and the raster data such as the digital elevation model H. q=1, 2, ..., Q; p=1, 2, ..., P.

[0034] The cropping ratio can cover the range of [1, 4] (e.g., 1:1, 4:3, 16:9, 2:1, 1:2, etc.), and the scale can be adaptively set according to the pixel resolution and the size of the task target (e.g., landslide scale, river network width, fault length).

[0035] Step S103, Cross-modal entropy calculation (early layer): For each candidate pruning X_{q,p}, the system inputs it along with text task prompts into the geoscience multimodal large model (i.e., the Vision-Language Model, VLM)). The system extracts visual and text tokens in the early layer k (early layer refers to the geoscience multimodal large model or the visual-language model): the visual encoder maps spatial data into a sequence of visual tokens. (N is the number of visual tokens, D is the dimension); The text encoder maps text task cues to a sequence of text tokens. (M represents the number of prompt tokens). When the geoscientific modal sample data includes geoscientific sample data from multiple modalities, the remote sensing image I is processed by a visual encoder to obtain the first set of visual token sequences; the digital elevation model H is processed by a visual encoder (or a dedicated encoder) to obtain the second set of visual token sequences; and the vector layer V is processed by an encoder to obtain the third set of visual token sequences.

[0036] The architecture of the visual language model includes: a first layer (encoding layer): a visual encoder, responsible for mapping geoscientific modality sample data (remote sensing image I, DEM H, vector layer V, etc.) into visual token sequences; a second layer (encoding layer): a text encoder, responsible for mapping text task prompts T into text token sequences; and a third layer (inference layer): a large language model, which uses visual tokens and text tokens as joint inputs to perform geoscientific feature extraction inference (including segmentation, detection, OCR, question answering, and other tasks). In this embodiment, the inputs of the large language model are geoscientific modality sample data, the optimal cropping corresponding to the geoscientific modality sample data, and text task prompts.

[0037] Furthermore, the visual language model introduces a projection layer between the visual encoder and the large language model, and inserts low-rank adaptation (LoRA) modules into each attention layer of both the visual encoder and the large language model. During visual language model training, the original model parameters of the text encoder, visual encoder, and large language model are frozen, and only the model parameters of the projection layer and the LoRA module are updated. The projection layer is responsible for modal feature alignment, mapping visual tokens from the visual space dimension to the feature space dimension of the large language model, enabling joint attention computation between the two. Cross-modal entropy computation is performed after extracting visual and text tokens in the early layer k (usually k=1 or 2), without adding additional inference burden.

[0038] The cue conditional distribution is constructed based on the visual token sequence and the text token sequence, as shown in formula (1).

[0039] (1).

[0040] in, For the first A visual token Normalized attention probability distribution under text prompt conditions; The number of visual tokens in the visual token sequence; For text tokens in a text token sequence Quantity; The base is the natural number; Indicates inner product similarity; For the first A text token; For the first A visual token; Temperature coefficient ( The smaller the size, the sharper the distribution.

[0041] The formula for calculating cross-modal entropy is shown in equation (2) below.

[0042] (2).

[0043] in, The cross-modal entropy corresponding to the candidate sample pruning; To avoid The constant is . When candidate cropping can more completely cover the relevant geoscientific target, its visual token probability distribution is more concentrated and the cross-modal entropy is smaller; conversely, if the target is fragmented by fixed blocks or submerged by irrelevant background, the distribution is more dispersed and the cross-modal entropy is larger.

[0044] Geoscience modal sample data can include geoscience sample data of multiple modalities; at this time, the sample candidate clipping set corresponding to each geoscience modal sample data includes the sample candidate clipping set of all geoscience sample data in the geoscience modal sample data; each sample candidate clipping includes geoscience sample data of multiple modalities at the same spatial location; the visual token sequence of the sample candidate clipping includes the visual token sequence of the sample candidate clipping of all geoscience sample data in the geoscience modal sample data.

[0045] In step 204 above, the cross-modal entropy corresponding to the candidate clipping of the samples is calculated based on the visual token sequence and the text token sequence, and the optimal clipping corresponding to the geoscience modality sample data is determined based on the cross-modal entropy corresponding to all candidate clipping of the samples. Specifically, the optimal clipping corresponding to the geoscience modality sample data can be obtained using either method 1 or method 2.

[0046] Method 1: Align the visual token sequences of candidate cropped samples from each modality in the geoscience modality sample data according to their spatial positions and concatenate them into a geoscience visual token sequence. Calculate the cross-modal entropy corresponding to the candidate cropped samples based on the geoscience visual token sequence and the text token sequence, and determine the optimal cropping for the geoscience modality sample data based on the cross-modal entropy corresponding to all candidate cropped samples. In other words, align the visual tokens obtained from different modalities through their respective encoders according to their spatial positions and concatenate them into a geoscience visual token. Then substitute the geoscientific visual token into the visual tokens in equations (1) and (2). , and calculate cross-modal entropy with text token.

[0047] Method 2: The text token sequence is combined with the visual token sequence of the candidate cropping data for each modality to calculate the corresponding cross-modal entropy of the geoscientific sample data. The cross-modal entropies corresponding to all modalities of geoscientific sample data are weighted and summed to obtain the cross-modal entropy corresponding to the candidate cropping data. The optimal cropping for the geoscientific modality sample data is then determined based on the cross-modal entropies corresponding to all candidate cropping data. When the geoscientific modality sample data includes data from three modalities: remote sensing image I, SAR image I_sar, and DEM H, the cross-modal entropies HS_I, HS_H, and HS_sar for each modality are calculated separately, and then the cross-modal entropies of each modality are fused according to their weights. The formula is as follows: Where w_m can be adaptively determined by data quality (such as cloud cover, noise) or task type. Represents mode, .

[0048] Step S104, Select the minimum entropy crop input: Calculate and sort the cross-modal entropy of all candidate croppings under the same image / same ROI, and select the candidate cropping with the minimum cross-modal entropy as the optimal cropping. This can be used as the final input; or, in engineering practice, the candidate clipping corresponding to the nth smallest cross-modal entropy, sorted from smallest to largest, can be selected as the optimal clipping to balance coverage and stability.

[0049] To avoid over-pruning leading to missing context, the system can perform optimal pruning. The input is fused with the original view or thumbnail of the geoscientific sample data. For example, the large language model adopts a dual-path input structure of "primary clipping (optimal clipping) + global thumbnail" to maintain both local fine detail and global positional relationship.

[0050] Step S105, Feature Extraction / Inference: The system is based on Perform geoscientific feature extraction tasks, which include at least one of the following: a) Area features: segmentation and vectorization of landslide bodies / debris flow deposition areas / floodplains / mining subsidence areas, etc.; b) Linear features: detection and connectivity restoration of faults / fissures / river networks / roads, etc.; c) Point features: location of mining sites / hazard points / well locations, etc.; d) Map text and symbols: geological map annotation / OCR and symbol recognition; e) Question and answer and explanation: perform multi-step reasoning and output explanations based on prompts such as "possible landform types in this area" and "whether there are signs of landslides in this area".

[0051] In at least one embodiment, the training method for the geoscientific big data feature extraction model further includes: mapping the normalized attention probability distribution of each visual token under text prompt conditions onto the visual token to form an attention heatmap; the attention heatmap is used as attention guidance for segmentation boundary enhancement and feature detection in geoscientific feature extraction.

[0052] The system maps the normalized attention probability distribution of visual tokens under text prompts back to a spatial patch grid to form an attention heatmap A. Thresholds, morphological operations, and topological constraints (such as river network connectivity and fault continuity) are used to enhance the segmentation / detection results, thereby improving boundary integrity and geometric rationality.

[0053] The spatial patch grid here refers to the two-dimensional grid structure formed by the visual encoder after dividing a candidate crop into fixed-size patches (e.g., the visual encoder divides a 224 patch into a 2D grid). A 224-bit image is cut into 14 bits. 14 = 196 patches, forming 14 lines (14-column spatial grid).

[0054] Each visual token corresponds to a patch segmented by the visual encoder, and the calculated... Each visual token Normalized attention probability distribution under text prompt conditions This indicates the semantic relevance of the visual token (corresponding to a patch at a certain spatial location in the original image) under the current text prompt conditions. Fill the original patch grid with the spatial coordinates corresponding to the visual token (e.g., 14). A matrix of 14 is used to obtain the attention heatmap A. Regions with higher values ​​in the attention heatmap A represent regions with stronger semantic relevance to the text task cues T, and can be used for attention guidance in subsequent segmentation boundary enhancement and feature detection.

[0055] Step S106, stitching, georegistration, and output: When the input is a large-format image or multi-tile data, the system maps the cropping domain results back to the original coordinate system through the recorded affine transformation matrix, and performs fusion on overlapping areas (such as IoU-based merging and confidence-based non-maximum suppression (NMS)). The final sample feature extraction results include: 1) Raster results: probability map, segmentation mask, change detection map; 2) Vector results: point / line / polygon features (Shapefile / GeoJSON, etc.) and attributes; 3) Text results: explanation report, metadata, and quality assessment indicators. The above outputs can also be provided externally through Web services / APIs.

[0056] To enhance the robustness and generalization ability of geoscientific element extraction, this application further introduces entropy regularization training (ERT) in the training / fine-tuning stage.

[0057] Step T1 (Training Data Construction): The system acquires labeled geoscientific training samples, including: raster labels: landform / surface cover categories, hazard segmentation masks, rasterized results of fault lines / river networks; vector labels: fault lines, geological boundaries, landslide boundary polygons; text labels: feature names, attribute descriptions, question-answer pairs, etc. The system registers the labels with the spatial data, generating (X, T, Y) triples, where Y is the task label. The task label is the supervised learning objective corresponding to the input sample (geological modality sample data X, text task prompt T), including three categories: raster labels (such as semantic segmentation masks), vector labels (such as feature boundary coordinates), and text labels (such as geological description text).

[0058] Step T2 (Freezing the backbone and making minor tweaks): In one embodiment, to reduce the training cost of geoscientific big data, the original model parameters of the text encoder, visual encoder, and large language model are frozen, and only a small number of adaptation parameters (such as the model parameters of the LoRA module, projection layers, or task heads) are updated during training to achieve rapid transfer under limited computing power. This strategy also helps to maintain the original general inference capabilities of the large model.

[0059] During training, only the model parameters of the projection layer and the LoRA module are updated, while the remaining parameters of the visual encoder and language model are frozen to reduce training costs. The model parameters of the LoRA module consist of low-rank matrix pairs A and B, i.e., incremental weight delta_W = A dot B, where dot represents matrix multiplication.

[0060] Step T3 (Calculate cross-modal entropy and add it to training loss): For each training sample, the system calculates the cross-modal entropy in the early layer k and adds it as a regularization term to the task loss to obtain the total loss function, as shown in Equation (3).

[0061] (3).

[0062] in, This is the total loss function; The task loss function; This is geoscientific modal sample data; Provide text task prompts; Tag for task; This is the entropy regularization term; For regularization weights, ; For consistency loss function; The weights are for the consistency loss function.

[0063] For generative tasks (such as multimodal question answering / report generation), the task loss function can be used to predict the loss for the next token, as shown in equation (4) below.

[0064] (4).

[0065] in, For the first Each output token (referring to a text token, i.e., a token in the output text sequence generated by the autoregression of a large language model, such as geological descriptive words, labeled category words, etc.); For the front -1 generated token (also referring to text token); For optimal cropping; Provide text task prompts; The total length of the output sequence is the standard autoregressive next-token prediction loss.

[0066] For segmentation / detection tasks, the task loss function can be composed of cross-entropy loss, Dice loss, boundary loss, etc., and the expression is shown in equation (5) below.

[0067] (5).

[0068] in, The loss is calculated as pixel-level cross-entropy. The Dice loss measures the overlap between the prediction mask and the ground truth mask. For boundary loss, the prediction accuracy of the feature boundary region is specifically constrained; , , These are the weight hyperparameters for each item, which can be set to 1.0, 1.0, or 0.5.

[0069] The mechanism of entropy regularization is that when the visual token is strongly correlated with the text task cue T, its normalized attention probability distribution... Typically larger; minimizing cross-modal entropy makes the probability distribution more "sharp," thereby enhancing key tokens, suppressing irrelevant tokens, and improving feature extraction, localization, and boundary quality.

[0070] In the training phase, a consistency loss is added to make the key token space distribution of different modalities consistent under cue conditions, thereby improving robustness in missing test / noise scenarios. The consistency loss function expression is shown in Equation (6).

[0071] (6).

[0072] in, The consistency loss function is the average of the KL divergence among all mode pairs. This refers to the number of modes in the geoscientific modal sample data (e.g., C=3 corresponds to optical image I, DEM H, and SAR image I_sar); when the geoscientific modal sample data is single-mode data, i.e. When =1, =0; This is the KL divergence, used to measure the difference between two modal attention distributions; For the first Normalized attention probability distribution of visual tokens of each modality under textual prompts; For the first Normalized attention probability distribution of visual tokens of each modality under textual prompts.

[0073] The smaller the consistency loss function value, the more consistent the spatial regions of interest of each modality under the cue conditions. This helps to improve the overall robustness by utilizing the consistency constraints of other modalities when a certain modality is missing or noisy (such as when cloud cover obscures optical images).

[0074] Step T4 (Post-training Deployment): After training is complete, the system will update the adaptation parameters and inference configuration (candidate pruning sets R, S, and early layer parameters). , (etc.) are solidified into a model version, resulting in a trained geoscientific big data feature extraction model. During deployment, the inference phase executes the same steps as steps S101-S106 to achieve adaptive pruning and enhanced extraction.

[0075] The following is a detailed explanation of the above steps, using landslide area extraction as an example: 1) Input: High-resolution remote sensing image I and DEM H of a mountainous area. The text task prompt T is "Please identify suspected landslides and output boundary vectors, while explaining the criteria." 2) Candidate cropping: The system generates multiple cropping ratios (e.g., 1:1, 4:3, 16:9, 1:2) around the suspected area and sets multi-scale windows to cover the possible long axis directions of the landslide. 3) Cross-modal entropy: The system calculates the cross-modal entropy of each candidate cropping in the early layer. It finds that the cropping containing the tension cracks at the rear edge of the landslide, the deposited tongue at the front edge, and the valley background has lower cross-modal entropy because its visual token probability is more concentrated. 4) Extraction and enhancement: The system performs segmentation and boundary refinement on the optimal cropping, combines attention heatmaps to repair connectivity at boundary fractures, and outputs GeoJSON boundaries. 5) Explanatory output: The system provides explanatory text based on the prompt-related tokens, such as "tone / texture abrupt changes, topographic slope breaks, valley depositional morphology," etc.

[0076] This application also provides an application scenario in which the above-described geoscientific big data feature extraction model training method is applied. Specifically, the geoscientific big data feature extraction model training method provided in this embodiment can be applied in a feature extraction scenario. The feature extraction scenario includes a content production stage, a model training chain, and a feature extraction stage; geoscientific modal sample data and text task prompts enter the model training chain from the content production stage, and through human-computer collaboration, a trained geoscientific big data feature extraction model is obtained, which then enters the downstream feature extraction stage. The geoscientific big data feature extraction model training method provided in this embodiment belongs to the model training chain. Specifically, in the model training process for geoscientific modality sample data and text task prompts, the geoscientific modality sample data can be cropped to obtain several candidate cropping samples for each geoscientific modality sample data. The visual token sequence of the candidate cropping samples and the text token sequence of the text task prompts are extracted. The cross-modal entropy is calculated based on the visual token sequence and the text token sequence, and the optimal cropping is determined. The geoscientific modality sample data, the optimal cropping and the text task prompts are input into the large language model to obtain the sample element extraction results. The trained geoscientific big data element extraction model is obtained by using the total loss function.

[0077] Example 2: The geoscientific big data feature extraction model application method provided in this application embodiment can be applied to, for example... Figure 1 In the application environment shown, terminal 102 communicates with server 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be set up independently, integrated into server 104, or placed in the cloud or on another server. Terminal 102 can send geoscientific modality target data and text task prompts to server 104. After receiving the geoscientific modality target data and text task prompts, server 104 performs cropping on the geoscientific modality target data to obtain a target candidate cropping set. The target candidate cropping set and the text task prompts are then input into a trained geoscientific big data element extraction model to obtain the target element extraction results. Server 104 can then feed back the obtained target element extraction results based on the geoscientific modality target data and text task prompts to terminal 102. In addition, in some embodiments, the application method of the geoscience big data feature extraction model can also be implemented by the server 104 or the terminal 102 separately. For example, the terminal 102 can directly extract features from the geoscience modal target data and text task prompts, or the server 104 can obtain the geoscience modal target data and text task prompts from the data storage system and extract features from the geoscience modal target data and text task prompts.

[0078] The terminal 102 can be, but is not limited to, various desktop computers and laptops. The server 104 can be implemented using a standalone server or a server cluster consisting of multiple servers, or it can be a cloud server.

[0079] In an exemplary embodiment, a method for applying a geoscientific big data feature extraction model is provided, comprising: acquiring geoscientific modal target data and text task prompts for a target region; cropping the geoscientific modal target data to obtain a target candidate cropping set; each target candidate cropping set includes several target candidate crops; inputting the target candidate crops and text task prompts into a trained geoscientific big data feature extraction model to obtain target feature extraction results; the visual encoder in the visual language model extracts the visual token sequence of the target candidate crops, and the text encoder extracts the text token sequence of the text task prompts; calculating the cross-modal entropy corresponding to the target candidate crops based on the visual token sequence of the target candidate crops and the text token sequence of the text task prompts, and determining the optimal cropping corresponding to the geoscientific modal target data based on the cross-modal entropy corresponding to all target candidate crops; inputting the geoscientific modal target data, the optimal cropping corresponding to the geoscientific modal target data, and the text task prompts into the large language model in the visual language model to obtain target feature extraction results; the trained geoscientific big data feature extraction model is a model trained using the geoscientific big data feature extraction model training method described in Embodiment 1.

[0080] This application also provides an application scenario in which the above-mentioned geoscientific big data element extraction model application method is applied. Specifically, the geoscientific big data element extraction model application method provided in this embodiment can be applied in a content distribution scenario. The content distribution scenario includes a content production stage, an element extraction chain, and a content distribution stage; geoscientific modality target data and text task prompts enter the element extraction chain from the content production stage, and obtain the corresponding target element extraction results through human-machine collaboration, and then enter the downstream content distribution stage. The geoscientific big data element extraction model application method provided in this embodiment belongs to the machine labeling stage in the element extraction chain. Specifically, in the element extraction chain process for geoscientific modality target data and text task prompts, the geoscientific modality target data can be cropped to obtain a target candidate cropping set of geoscientific modality target data. The target candidate cropping and text task prompts are input into the trained geoscientific big data element extraction model to obtain the target element extraction results.

[0081] This application utilizes the cross-modal attention distribution of early layers in a large model / visual language model to calculate the cross-modal entropy between "image (or geoscientific raster) - text task," which serves as a semantic integrity indicator. During the inference phase, the minimum cross-modal entropy is used to select the optimal clipping ratio and scale, preserving as much continuity as possible the geoscientific target areas (e.g., faults, landslides, river networks, lithological boundaries, map text annotations, etc.) related to the task prompts. During the training / fine-tuning phase, the cross-modal entropy is added as a regularization term to the task loss, prompting the model to focus its attention more on key visual tokens and suppress irrelevant tokens, thereby enhancing the quality of feature extraction.

[0082] Traditional feature extraction methods suffer from cropping mismatches with tasks: the region of interest for the same image varies under different tasks, but fixed cropping cannot adapt to text instructions; one related technique uses saliency detection, edge density, or target proposal networks to generate cropping windows, but its cropping decisions do not explicitly depend on text task prompts, easily resulting in "saliency but irrelevant" cropping, especially when the same image carries multiple types of geoscientific targets; another related technique obtains a saliency map by calculating the gradient of the target category or output logit and guides the cropping, which can be associated with the task to some extent, but requires at least one complete forward-backward computation, resulting in higher computational costs and sensitivity to noise; yet another related technique enhances continuity by modifying the visual encoder, but usually requires changes to the model structure, has high training costs, and is complex to implement in engineering on multi-source geoscientific data.

[0083] This application draws on the idea that "the entropy of cross-modal attention distribution in early layers can characterize the integrity of cue-related regions," and uses cross-modal entropy as a computable and optimizable indicator: the minimum cross-modal entropy is used for pruning during the inference phase; and the minimum cross-modal entropy is used during the training phase. Regularized attention addresses the semantic fragmentation and unstable feature extraction issues of large-format geoscientific imagery with fixed tile divisions. Stable metrics can be obtained in early layers through cross-modal entropy, resulting in low inference overhead. Meanwhile, the cropping strategy is explicitly constrained by textual cues, making it suitable for multi-task scenarios involving the same image.

[0084] This application has the following advantages: 1) Improved integrity of relevant areas: Traditional high-resolution remote sensing processing often uses fixed tiles or fixed-ratio cropping, which easily fragments continuous structures such as faults, river networks, lithological boundaries, and geological map annotations, leading to fragmented or missed element extraction. This application selects the optimal cropping based on cross-modal entropy in steps S103-S104, making the relevant areas of the prompt more continuous in space, reducing semantic fragmentation, and improving the integrity of elements from the source. 2) Cross-modal adaptation to avoid the cropping bias of "only looking at the picture and not the question": Existing cropping / scaling strategies mostly rely on visual saliency or pure geometric rules, ignoring the constraints of text task prompts, and easily cropping to "salient but irrelevant" areas. This application incorporates text task prompts into the normalized attention probability distribution and cross-modal entropy calculation under the text prompt conditions (Equation (1)-Equation (2)), realizing true cross-modal adaptive cropping, which can automatically focus on different spatial areas for different tasks (landslides, faults, river networks, mineralization alteration, etc.). 3) Improved feature extraction quality and controllable overhead: This application only uses early layers of the model to calculate cross-modal entropy, without needing to perform full inference on all candidate clippings; at the same time, the number of candidate clippings is controllable (e.g., or The overall computational cost is low. By introducing entropy regularization training in step T3, the model's attention is more focused on key visual tokens, reducing background interference and improving the clarity of segmentation boundaries, the continuity of linear features, and the accuracy of OCR. (4) Adapting to multi-source geoscientific data and engineering deployment: This application provides a multimodal extension method that can process optical, SAR, DEM and other data simultaneously, and supports batch processing and online services.

[0085] Example 3: In an exemplary embodiment, a computer device is provided, which may be a server or a terminal, and its internal structure diagram may be as follows. Figure 4 As shown, the computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operating system and computer programs in the non-volatile storage media to run. The database stores geoscientific feature extraction data. The I / O interfaces are used for information exchange between the processor and external devices. The communication interface is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements a geoscientific big data feature extraction model training method or a geoscientific big data feature extraction model application method.

[0086] Those skilled in the art will understand that Figure 4 The structures shown are merely block diagrams of some structures related to the present application and do not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than shown in the figures, or combine certain components, or have different component arrangements. In an exemplary embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.

[0087] Example 4: In an exemplary embodiment, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.

[0088] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Moreover, the collection, use and processing of the relevant data are carried out in compliance with the relevant data protection laws and policies of the country where the location is located, and with the authorization granted by the owner of the corresponding device.

[0089] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM).

[0090] The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0091] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0092] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A method for training a geoscientific big data feature extraction model, characterized in that, The training method for the geoscientific big data feature extraction model includes: Acquire several geoscientific modal sample data and corresponding text task prompts for each geoscientific modal sample data; Each geoscientific modality sample data is cropped to obtain a sample candidate cropping set corresponding to each geoscientific modality sample data; each sample candidate cropping set includes several sample candidate croppings. The candidate cropping samples and text task prompts are input into the visual encoder and text encoder of the visual language model, respectively, to extract the visual token sequence of the candidate cropping samples and the text token sequence of the text task prompts. The geoscience modality sample data includes geoscience sample data of multiple modalities. The candidate cropping set corresponding to each geoscience modality sample data includes the candidate cropping set of all geoscience sample data in the geoscience modality sample data. Each candidate cropping sample includes geoscience sample data of multiple modalities at the same spatial location. The visual token sequence of the candidate cropping samples includes the visual token sequence of the candidate cropping samples of all geoscience sample data in the geoscience modality sample data. The cross-modal entropy corresponding to the candidate clipping of samples is calculated based on the visual token sequence and the text token sequence. Then, the optimal clipping for the geoscientific modality sample data is determined based on the cross-modal entropy corresponding to all candidate clippings. Specifically, this includes: The visual token sequences of candidate cropping samples from each modality in the geoscience modality sample data are aligned according to their spatial positions and concatenated into a geoscience visual token sequence. The cross-modal entropy corresponding to the candidate cropping samples is calculated based on the geoscience visual token sequence and the text token sequence, and the optimal cropping corresponding to the geoscience modality sample data is determined based on the cross-modal entropy corresponding to all candidate cropping samples. Alternatively, the text token sequence can be combined with the visual token sequence of the candidate cropping of the geoscientific sample data for each modality to calculate the corresponding cross-modal entropy of the geoscientific sample data for each modality; the cross-modal entropy of the corresponding geoscientific sample data for all modalities can be weighted and summed to obtain the cross-modal entropy corresponding to the candidate cropping of the sample; and the optimal cropping corresponding to the geoscientific sample data can be determined based on the cross-modal entropy corresponding to all candidate cropping of the sample; the cross-modal entropy is used to measure the semantic integrity of the relevant region of the text task prompt in the candidate cropping of the sample; Geomorphic modal sample data, the optimal cropping of the corresponding geomorphic modal sample data, and text task prompts are input into a large language model in the visual language model to extract geomorphic elements, resulting in sample element extraction results. The geomorphic element extraction task includes at least one of the following: a) Area elements: segmentation and vectorization of landslide bodies / debris flow deposition areas / floodplains / mining subsidence areas; b) Linear elements: detection and connectivity restoration of faults / fissures / river networks / roads; c) Point elements: location of mining sites / hazard points / well locations; d) Map text and symbols: geological map annotation / OCR and symbol recognition; e) Question answering and explanation: performing multi-step reasoning and outputting explanations for prompts such as "possible landform types in this area" and "whether there are signs of landslides in this area". The model parameters of the visual language model are updated using the total loss function to obtain a trained geoscientific big data feature extraction model. The total loss function includes the cross-modal entropy based on the sample candidate cropping and the sample feature extraction results. The total loss function is expressed as follows: ; ; in, This is the total loss function; The task loss function; This is geoscientific modal sample data; Provide text task prompts; Tag for task; This is the entropy regularization term; Regularization weights; For consistency loss function; This refers to the number of modes in the geoscientific modal sample data; when the geoscientific modal sample data is single-modal data, that is... When =1, =0; These are the weights of the consistency loss function; Let KL divergence be a metric. For the first Normalized attention probability distribution of visual tokens of each modality under textual prompts; For the first Normalized attention probability distribution of visual tokens of each modality under textual prompts.

2. The method for training a geoscientific big data feature extraction model according to claim 1, characterized in that, The formula for calculating cross-modal entropy is as follows: ; ; in, The cross-modal entropy corresponding to the candidate sample pruning; For the first A visual token Normalized attention probability distribution under text prompt conditions; It is a constant; The number of visual tokens in the visual token sequence; For text tokens in a text token sequence Quantity; Indicates inner product similarity; For the first A text token; For the first A visual token; This is the temperature coefficient.

3. The method for training a geoscientific big data feature extraction model according to claim 2, characterized in that, The training method for the geoscience big data feature extraction model also includes: The normalized attention probability distribution of each visual token under textual prompts is mapped onto the visual token to form an attention heatmap; the attention heatmap is used as attention guidance for segmentation boundary enhancement and feature detection in geoscientific feature extraction.

4. The method for training a geoscientific big data feature extraction model according to claim 1, characterized in that, The method for training the geoscientific big data feature extraction model also includes: preprocessing geoscientific modal sample data; the preprocessing includes coordinate unification, resolution matching, cloud removal and noise reduction, normalization and / or tile conversion.

5. The method for training a geoscientific big data feature extraction model according to claim 1, characterized in that, The visual language model introduces a projection layer between the visual encoder and the large language model, and inserts low-rank adaptation modules into each attention layer of the visual encoder and the large language model. During the training of the visual language model, the original model parameters of the text encoder, the visual encoder and the large language model are frozen, and only the model parameters of the projection layer and the low-rank adaptation modules are updated.

6. A method for applying a geoscientific big data feature extraction model, characterized in that, The application methods of the geoscience big data feature extraction model include: Acquire geoscientific modal target data and text task prompts for the target area; The target data of geoscience modalities are cropped to obtain a target candidate cropping set; each target candidate cropping set includes several target candidate croppings. The target candidate cropping and text task prompts are input into the trained geoscientific big data feature extraction model to obtain the target feature extraction results. The visual encoder in the visual language model extracts the visual token sequence of the target candidate cropping, and the text encoder extracts the text token sequence of the text task prompts. The cross-modal entropy corresponding to the target candidate cropping is calculated based on the visual token sequence of the target candidate cropping and the text token sequence of the text task prompts, and the optimal cropping corresponding to the geoscientific modality target data is determined based on the cross-modal entropy corresponding to all target candidate croppings. The geoscientific modality target data, the optimal cropping corresponding to the geoscientific modality target data, and the text task prompts are input into the large language model in the visual language model to obtain the target feature extraction results. The trained geoscientific big data feature extraction model is a model trained using the geoscientific big data feature extraction model training method described in any one of claims 1-5.

7. A computer device, comprising: A memory, a processor, and a computer program stored in the memory and capable of running on the processor, characterized in that the processor executes the computer program to implement the geoscientific big data feature extraction model training method according to any one of claims 1-5 or the geoscientific big data feature extraction model application method according to claim 6.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the geoscientific big data feature extraction model training method according to any one of claims 1-5 or the geoscientific big data feature extraction model application method according to claim 6.