Geological exploration data cloud management and analysis method, device and equipment

By constructing a cloud-based data lake and using tensor analysis technology, the problems of managing and allocating multi-source geological exploration data have been solved, enabling efficient data retrieval and intelligent allocation of exploration resources, reducing exploration costs and improving exploration efficiency.

CN120875482AInactive Publication Date: 2025-10-31ZHUHAI EAGLER SPECIALTY DRILLING EQUIP CO LTD +3
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511392024.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-26
Publication Date
2025-10-31
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing technologies lack effective means of fusion of multi-source geological exploration data, resulting in difficulties in data retrieval, serious duplication of storage, inconvenience in cross-departmental sharing, high exploration costs, unreasonable resource allocation, and difficulty in achieving collaborative management of multiple projects.

Method used

By constructing a cloud-based data lake for unified management, using tensor analysis technology to deeply mine data relationships, and combining field theory principles and evolutionary analysis, multi-project collaborative management solutions are generated, and intelligent optimization technology is used for resource allocation.

Benefits of technology

It has enabled unified management and efficient retrieval of exploration data, revealed the spatiotemporal correlations and causal relationships between different data sources, optimized the allocation of exploration resources, reduced exploration costs and improved exploration efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120875482A_ABST
    Figure CN120875482A_ABST
Patent Text Reader

Abstract

The invention discloses a geological exploration data cloud management and analysis method, device and equipment, and the method comprises the steps: achieving the preprocessing of multi-source exploration data through space-time marking and quality grading, and constructing a cloud data lake for structural storage; a geologic feature group is extracted through tensor analysis, data field intensity distribution is established, and a related project group is tracked and positioned through an equipotential line; deploying a data pulsation monitoring network, capturing update frequency fluctuation to form a respiratory rhythm diagram, identifying an active data source and extracting a collaborative triggering mechanism; carrying out evolution analysis after synchronously recombining the project group, obtaining a geological evolution track and extracting a trend vector; performing cross validation on the trend vector and geologic features to generate a confidence gradient field, and tracking a high-confidence path to deduce an exploration hot spot region; a resource allocation matrix is constructed based on hotspot spatial distribution, a project priority sequence is determined, a multi-project collaborative management scheme is output, and a complete technical system of geological exploration from data management, intelligent analysis to decision optimization is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of geological exploration data processing technology, and in particular to a method, apparatus and equipment for cloud-based management and analysis of geological exploration data. Background Technology

[0002] With the rapid development of geophysical exploration technology, the amount of data generated by geological exploration has increased dramatically, including various types such as seismic wave data, electromagnetic field data, and gravity field data. This data is typically stored on different local servers or workstations, lacking a unified management platform, leading to difficulties in data retrieval, significant duplication of storage, and inconvenience in cross-departmental sharing. Furthermore, traditional data processing methods rely heavily on human experience for interpretation, making it difficult to discover the inherent relationships between different types of data and to fully utilize the geological information contained within the massive datasets.

[0003] In practical exploration projects, it is often necessary to comprehensively analyze multiple geophysical data to accurately determine the underground geological structure and resource distribution. However, current technologies lack effective multi-source data fusion methods, making it difficult to establish spatiotemporal correspondences between different data sources. Inconsistent data quality also affects the reliability of comprehensive interpretation. Furthermore, when multiple exploration projects are carried out in parallel, resource allocation mainly relies on experience-based judgments, lacking scientific optimization methods, resulting in low equipment utilization, unreasonable personnel allocation, and high exploration costs.

[0004] Therefore, a method is urgently needed to solve at least one of the above problems. Summary of the Invention

[0005] This invention provides a cloud-based management and analysis method, apparatus, and equipment for geological exploration data. It aims to achieve unified management of multi-source exploration data by constructing a cloud-based data lake, deeply mine the correlation between data using tensor analysis technology, reveal the spatiotemporal laws of geological processes using field theory principles and evolutionary analysis, and generate multi-project collaborative management schemes through intelligent optimization, thus providing a data-driven solution for geological exploration.

[0006] The first aspect of this invention proposes a cloud-based management and analysis method for geological exploration data, comprising the following steps: Collect monitoring data from each exploration site, and perform quality grading on the monitoring data to generate a graded data queue; The hierarchical data queue is preprocessed at the edge to form structured data blocks, and the structured data blocks are dynamically stored and optimized to obtain a cloud data lake. Deep scanning is performed on the cloud data lake to extract geological feature groups. The structured data blocks are matched with the geological feature groups in multiple dimensions to obtain the correlation tensor. A data field strength distribution map is constructed based on the correlation tensor. Isopotential line tracing is performed on the data field strength distribution map to locate related project groups. A data pulsation monitoring network is established within the relevant project group. The data pulsation monitoring network is used to capture update frequency fluctuations to form a respiratory rhythm map. The respiratory rhythm map and the geological feature group are subjected to time series correlation analysis to obtain active data sources. Interactive analysis is performed on the active data sources to extract a collaborative triggering mechanism. The collaborative triggering mechanism is used to perform synchronous reorganization on the relevant project group to obtain the evolution sequence, curvature analysis is performed on the evolution sequence to obtain the geological evolution trajectory, and inflection point extraction is performed on the geological evolution trajectory to generate a trend vector; The trend vector and the geological feature group are cross-validated layer by layer to form a confidence gradient field. The confidence gradient field is used to trace the path to obtain the predicted path bundle. The predicted path bundle is used to perform convergence analysis to deduce exploration hotspot areas. Spatial clustering analysis is performed on the exploration hotspot areas to construct a resource allocation matrix. Based on the resource allocation matrix, a project priority sequence is generated, and a multi-project collaborative management scheme is output according to the project priority sequence.

[0007] A second aspect of this invention provides a cloud-based management and analysis device for geological exploration data, comprising: The data acquisition module is used to collect monitoring data from various exploration sites and to perform quality grading on the monitoring data to generate a graded data queue. The cloud storage module is used to perform edge preprocessing on the hierarchical data queue to form structured data blocks, and to perform dynamic storage optimization on the structured data blocks to obtain a cloud data lake; The project association module is used to perform deep scanning on the cloud data lake to extract geological feature groups, perform multi-dimensional matching between the structured data blocks and the geological feature groups to obtain association tensors, construct a data field strength distribution map based on the association tensors, and perform equipotential line tracing on the data field strength distribution map to locate related project groups. The data monitoring module is used to establish a data pulsation monitoring network within the relevant project group, capture update frequency fluctuations through the data pulsation monitoring network to form a respiratory rhythm map, perform time-series correlation analysis on the respiratory rhythm map and the geological feature group to obtain active data sources, and perform interactive analysis on the active data sources to refine the collaborative triggering mechanism. The trend analysis module is used to perform synchronous reorganization on the relevant project group using the collaborative triggering mechanism to obtain the evolution sequence, perform curvature analysis on the evolution sequence to obtain the geological evolution trajectory, and extract inflection points from the geological evolution trajectory to generate a trend vector. The hotspot prediction module is used to perform layer-by-layer cross-validation between the trend vector and the geological feature group to form a confidence gradient field, perform path tracing on the confidence gradient field to obtain a predicted path bundle, and perform convergence analysis on the predicted path bundle to deduce exploration hotspot areas. The solution output module is used to perform spatial cluster analysis on the exploration hotspot areas to construct a resource allocation matrix, generate a project priority sequence based on the resource allocation matrix, and output a multi-project collaborative management solution according to the project priority sequence.

[0008] A third aspect of the present invention provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of a cloud-based geological exploration data management and analysis method disclosed in the first aspect.

[0009] The beneficial effects of this invention are reflected in the following points: First, through the structured encapsulation and multi-level indexing mechanism of the cloud data lake, exploration data that was originally scattered and varied in format is now under unified management. The quality grading system automatically filters low-quality data, ensuring that the analysis process uses high-reliability data sources. Simultaneously, the establishment of the index table optimizes data retrieval from a full-database scan to targeted queries, improving retrieval efficiency. Second, by employing tensor analysis and data field strength modeling techniques, discrete exploration data points are transformed into continuous field distributions. Isopotential line tracing automatically identifies data convergence areas. Combined with the collaborative triggering mechanism discovered through data pulsation monitoring, the spatiotemporal correlations and causal relationships between different data sources can be revealed, upgrading the identification of geological anomalies from single-data judgment to multi-source data collaborative analysis. Third, the geological evolution trajectory and trend vector extracted through curvature analysis of evolutionary sequences, verified by a confidence gradient field, generate exploration hotspots. Combined with the resource allocation matrix and project priority sequence formed by spatial clustering, intelligent allocation of exploration resources and multi-project collaborative operations are achieved, avoiding resource conflicts and redundant investment in traditional experience-based management.

[0010] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description

[0011] The accompanying drawings illustrate specific examples of the technical solutions described in this invention and, together with the detailed embodiments, form part of the specification, serving to explain the technical solutions, principles, and effects of this invention.

[0012] Unless otherwise specified, the same reference numerals in different figures represent the same or similar technical features, and different reference numerals may be used to represent the same or similar technical features.

[0013] Figure 1 This is a flowchart illustrating a cloud-based method for managing and analyzing geological exploration data according to the present invention.

[0014] Figure 2This is a structural block diagram of a cloud-based geological exploration data management and analysis device according to the present invention.

[0015] Figure 3 This is a schematic diagram of the structure of a computer device according to the present invention. Detailed Implementation

[0016] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0017] It should be noted that all directional indicators (such as up, down, left, right, front, back, etc.) in the embodiments of this application are only used to explain the relative positional relationship and movement of each component in a certain specific posture (as shown in the figure). If the specific posture changes, the directional indicator will also change accordingly.

[0018] It should also be noted that when a component is described as "fixed to" or "set on" another component, it can be directly on the other component or there may be an intervening component present. When a component is described as "connected to" another component, it can be directly connected to the other component or there may be an intervening component present.

[0019] Furthermore, the use of terms such as "first" and "second" in this application is for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. Additionally, the technical solutions of the various embodiments can be combined with each other, but only on the basis of being achievable by those skilled in the art. When the combination of technical solutions is contradictory or impossible to implement, such a combination of technical solutions should be considered non-existent and not within the scope of protection claimed in this application.

[0020] The technical solutions of the embodiments of this application will be described below.

[0021] like Figure 1 As shown, this embodiment of the invention provides a cloud-based management and analysis method for geological exploration data, including the following steps S110-S170: Step S110: Collect monitoring data from each exploration site, and perform quality grading on the monitoring data to generate a graded data queue.

[0022] Monitoring data is collected from various exploration sites. A distributed acquisition architecture is used to acquire multi-source heterogeneous monitoring data from sensor networks at each exploration site, including geophysical parameters such as seismic waveform data, electromagnetic field strength, gravity gradient, and geothermal distribution. Sampling frequencies are set differently based on the measurement type: 1000Hz for seismic data, 100Hz for electromagnetic data, 10Hz for gravity gradient, and 1Hz for geothermal data. An edge server is deployed at each site, connecting to the sensor array via TCP / IP protocol, supporting 128 concurrent data streams. The equipment operation log is in JSON format, containing five core fields: timestamp, device number, status code, sensor reading, and error flag. The sensor reading field nests and records four real-time parameters: voltage value, temperature value, GPS satellite count, and signal strength. The error flag uses a 32-bit mask to record various error types, such as hardware failure and communication anomalies. A breakpoint resume mechanism is introduced for data acquisition; data is automatically cached when the network is interrupted, and transmission resumes from the breakpoint upon recovery, forming a raw data stream containing four types of geophysical parameters and a complete equipment operation log.

[0023] The monitoring data is graded to generate a tiered data queue. First, spatiotemporal annotations are applied to the raw data stream to construct labeled data packets. A GPS timing system with microsecond-level accuracy is used, with each data point appended with a 64-bit timestamp. Spatial annotation employs RTK differential positioning technology with centimeter-level planar accuracy and within 5 cm elevation accuracy, using the WGS84 coordinate system. The annotation strategy is designed based on sampling frequency differences. A spatiotemporal fusion algorithm fuses GPS time with the real-time clock time extracted from sensor readings using Kalman filtering, with the formula T_fusion = 0.8 × T_GPS + 0.2 × T_RTC + K × (T_GPS - T_RTC_pred), where T_fusion is the fused timestamp, T_GPS is the GPS timing system time, T_RTC is the sensor real-time clock time, T_RTC_pred is the RTC prediction value based on historical data, and K is the Kalman filter gain coefficient. The labeled data packets use a three-layer nested structure to store spatiotemporal metadata, data type parameters, and the raw measurement value array. Next, the equipment operation logs are parsed and anomaly filtered to extract key status nodes. Structured content is extracted using a JSON parser, and anomaly detection employs multi-dimensional analysis: a voltage deviation exceeding 220V standard value by more than 10% is recorded as an anomaly; a temperature exceeding the operating range of -20°C to 60°C triggers an alarm; fewer than 4 GPS satellites indicate insufficient positioning; and a signal strength weaker than -85dBm indicates a communication anomaly. A predictive model is built for time-series sensor readings, and a deviation between measured and predicted values ​​exceeding 2 times the standard deviation is considered an anomaly. Time points with a severity score exceeding 0.6 are marked as critical state nodes. These critical state nodes are then used to assess and classify the quality of marked data packets, generating a graded data queue. Spatial quality is graded based on GPS accuracy and satellite count: accuracy <5cm and ≥8 satellites indicate excellent positioning; 5-20cm and 6-7 satellites indicate good positioning; 20-50cm and 4-5 satellites indicate usable positioning; and other values ​​indicate degraded positioning. The quality is divided into five levels: Level A is high-quality seismic or electromagnetic high-frequency data with no anomalies; Level B allows good positioning and minor anomalies; Level C accepts usable positioning, moderate anomalies, or low-frequency gravity and geothermal data; Level D includes degraded positioning or severely anomaly data; and Level E is all data during equipment failure. The quality scoring formula is Q=0.4×L_spatial+0.3×(1-S_device)+0.2×T_precision+0.1×D_integrity, where L_spatial is the spatial positioning quality (high-quality positioning = 1.0, good = 0.7, usable = 0.4, degraded = 0.1), S_device is the equipment anomaly severity score (range 0-1, no anomaly = 0), T_precision is the normalized value of time precision (microsecond level = 1.0, millisecond level = 0.5), and D_integrity is the data integrity index (complete = 1.0, minor missing = 0.8, severe missing = 0.5).Five priority queues store data packets of grades A to E respectively, with queue capacity ratios of 20%, 40%, 25%, 10%, and 5%. Each data packet carries quality grade and rating information.

[0024] Step S120: Perform edge preprocessing on the hierarchical data queue to form structured data blocks, and perform dynamic storage optimization on the structured data blocks to obtain a cloud data lake.

[0025] Specifically, edge preprocessing is performed on the hierarchical data queues to form structured data blocks. Each data packet in each queue retains complete quality level labels, spatiotemporal annotations, and arrays of raw measurements. Preprocessing strategies are designed differently based on quality level. High-quality data packets from the A and B queues undergo full preprocessing: first, raw measurements are extracted from the inner layer data; seismic data (1000Hz) is processed using a bandpass filter (1-250Hz) for 1024 data frames, and electromagnetic data (100Hz) is processed using a median filter for 256 data frames; then, characteristic parameters such as amplitude spectrum, phase spectrum, and instantaneous frequency are calculated; finally, normalization is performed. Data from the C queue is simplified to basic filtering and feature extraction, while data from the D and E queues only undergo outlier removal. The structured data block is designed into three parts: the metadata header (128 bytes) directly inherits the spatiotemporal annotation information (timestamp, WGS84 coordinates, GPS accuracy indicators) and quality level of the tagged data packet, and adds data type identifiers and processing parameters; the feature data area stores the preprocessed standardized parameter array; the quality tag area not only retains the original quality score Q value, but also queries the key status node information of the corresponding time period through the association of device ID and timestamp, and records the anomaly type and severity score.

[0026] In some embodiments, the step of dynamically optimizing the storage of the structured data blocks to obtain a cloud data lake includes: applying data fingerprint encoding to the structured data blocks to generate a unique identifier; constructing a dynamic storage topology graph based on the unique identifier; performing node density analysis on the dynamic storage topology graph to determine the optimal storage path; and performing distributed storage via the optimal storage path to form a cloud data lake.

[0027] Data fingerprinting is applied to structured data blocks to generate unique identifiers. Fingerprinting extracts key information from the data block: device number, timestamp, spatial coordinates, quality level, and data type are read from the metadata header; a content summary is calculated from the feature data area; and anomaly markers are obtained from the quality marker area. This information is concatenated and input into the SHA-256 algorithm to generate a 256-bit identifier. The first 64 bits serve as the partition key, the middle 128 bits as the block identifier, and the last 64 bits as the checksum. Crucially, an identifier index table is established. Each record contains: identifier, data block size (MB), quality level (AE), data type (earthquake / electromagnetic / gravity / geothermal), expected access frequency (preset based on quality level: 10 times / day for level A, 5 times / day for level B, 2 times / day for level C, and 0.5 times / day for levels D and E), device ID, and time range. The index table is stored in a distributed database, supporting rapid retrieval of the complete attributes of a data block via the identifier. A bidirectional mapping is established between the data block entity and the identifier, ensuring that subsequent processing can obtain all the necessary information.

[0028] A dynamic storage topology graph is constructed based on unique identifiers. Nodes in the topology graph represent storage locations. Data blocks are batch-read from the index table, providing attributes such as quality level, size, and type for initial allocation: A and B level data blocks are allocated to high-performance SSD nodes (by querying the `quality_level` field in the index table), C level blocks to standard storage, and D and E level blocks to cold storage. Edge weights are calculated using real data from the index table: W_edge = 0.4 / B + 0.3 × C + 0.3 × (1 / F), where B is the measured bandwidth, C is the current transmission cost, and F is obtained from the `expected_frequency` field in the index table. During construction, consecutive time-series data blocks with the same device ID (identified by the `device_id` and `time_range` fields in the index table) are adjacent, and data blocks with similar content (determined by Hamming distance from the identifier) ​​are clustered. Each node in the topology graph records information such as the list of currently stored data block identifiers, total capacity usage (accumulated from `block_size` in the index table), and quality level distribution (counting the number of each level). The attributes of the edges include dynamic parameters such as real-time bandwidth, latency, and cost, forming a complete topology that reflects data distribution and network status.

[0029] Node density analysis is performed on the dynamic storage topology to determine the optimal storage path. Node density calculation obtains accurate data from the index table: D_node = Σ(S_i × F_i) / C_max, where S_i is read from the block_size field of the index table, F_i is obtained from the expected_frequency field, and C_max is the maximum capacity of the storage node; these are actual values, not estimates. Statistics from the data_type field of the index table show that seismic and electromagnetic data occupy 80% of the capacity of hotspot nodes, and a traffic distribution strategy is formulated accordingly. When selecting the optimal path, the algorithm considers not only the topology but also the specific attributes of each data block: high-quality seismic data (type='seismic' AND qualityIN('A', 'B')) prioritizes low-latency paths, while low-quality data (qualityIN('D', 'E')) selects low-cost paths. The path planning results include: source node, target node sequence, estimated transmission time (calculated based on block_size and bandwidth), and storage cost (calculated based on the storage strategy according to the quality level). The path scheme for each data block is recorded in the storage_path field of the index table, which supports subsequent execution and adjustment.

[0030] Distributed storage is executed via the optimal path to reliably transmit data block entities and metadata to the cloud. During storage execution, complete information about the data block is read from the index table: the data block entity is located by its identifier, the storage tier is determined by the quality level, the block size is read for transmission fragmentation, and the storage path is used for routing. Level A data blocks are stored in the high-performance tier and maintained with three replicas; Levels B and C use standard storage with two replicas; and Levels D and E use cold storage with a single replica and erasure coding. During transmission, data blocks of 1-10MB are fragmented based on their actual size, while blocks larger than 5MB are divided into 1MB fragments for parallel transmission. In the object storage system of the cloud data lake, each data block is named and stored using an identifier. The metadata tag contains all fields of the index table, supporting multi-dimensional queries. The data lake directory service maintains the mapping from identifiers to storage locations, providing rich retrieval capabilities in conjunction with the index table: queries can be performed by time range, spatial region, quality level, data type, device origin, and other conditions. Through index-driven storage management, a cloud data lake is formed where data block entities and metadata are closely linked, achieving a complete data flow from hierarchical queues to cloud storage.

[0031] Step S130: Perform a deep scan on the cloud data lake to extract geological feature groups, perform multi-dimensional matching between structured data blocks and geological feature groups to obtain correlation tensors, construct a data field strength distribution map based on the correlation tensors, and perform equipotential line tracing on the data field strength distribution map to locate related project groups.

[0032] Specifically, a deep scan is performed on the cloud-based data lake to extract geological feature clusters. Through the query interface of the identifier index table, data is retrieved according to data type: seismic data blocks (data_type='seismic') are scanned to extract features such as reflection interfaces and wave impedance anomalies; electromagnetic data blocks (data_type='electromagnetic') identify features such as resistivity stratification and anomalies; gravity data blocks extract density interfaces and gravity gradient zones; and geothermal data blocks analyze geothermal gradient changes. The scanning algorithm performs pattern recognition on the feature data areas of each data type, extracting geologically significant feature vectors from preprocessed parameters such as amplitude spectrum, phase spectrum, and instantaneous frequency. Feature extraction utilizes quality levels to optimize the processing order, prioritizing A and B-level high-quality data blocks to ensure feature reliability, using C-level data as a supplement, and D and E-level data only for trend verification. Spatial clustering, based on the WGS84 coordinates in the data block metadata header, aggregates similar features in adjacent areas (within a 5km radius) to form feature clusters. Temporal correlation uses timestamp information to identify the evolution sequence and abrupt change points of features. The geological feature groups are organized in a hierarchical structure: the top layer is classified according to feature type (structural features, lithological features, fluid features), the middle layer is clustered according to spatial distribution, and the bottom layer records specific feature parameters and source data block identifiers. Each feature group includes attributes such as feature type, spatial range, time span, confidence level (calculated based on the source data quality level), and a list of associated data blocks, forming a structured geological feature group extracted from the cloud data lake.

[0033] Structured data blocks are matched with extracted geological feature groups in multiple dimensions to construct a tensor reflecting the correlation between data and geological features. First, a mapping from data blocks to features is established: complete information for each data block is queried using an identifier, including spatial coordinates, time range, data type, and quality level. Then, the intersection of this information with the spatial range and time span of the feature group is calculated. Spatial matching is accelerated using an R-tree index; an association is established when the data block coordinates fall within the spatial range of the feature group. Temporal matching checks whether the data acquisition time is within the feature's validity period. Type matching ensures the physical relevance between data type and feature type, such as matching seismic data with structural features and electromagnetic data with lithological features. The matching strength is calculated using the formula M = w1 × S_overlap + w2 × T_corr + w3 × Q_factor, where S_overlap is the spatial overlap (intersection area / feature area), T_corr is the temporal correlation coefficient, and Q_factor is the data quality factor (converted from quality_level), with weights w1 = 0.5, w2 = 0.2, and w3 = 0.3. The association tensor is designed as a three-dimensional structure [N_blocks×N_features×N_attributes], where N_blocks is the number of data blocks (approximately 100,000 based on the index table), N_features is the number of feature families (typically 1000), and N_attributes represents the association attribute dimensions (including eight attributes such as matching strength, spatial distance, temporal difference, and quality weight). Each element in the tensor records a pair of data block-feature association vectors, with non-zero elements indicating the existence of an association.

[0034] In some embodiments, constructing a data field intensity distribution map based on the correlation tensor includes: decomposing the correlation tensor into lithology-related components and tectonic-related components; applying stratigraphic compaction correction to the lithology-related components to generate a true lithology field; performing stress field inversion on the tectonic-related components to generate a paleotectonic stress field; and superimposing the true lithology field and the paleotectonic stress field to generate a data field intensity distribution map.

[0035] Tensor decomposition is performed based on the correlation tensor to separate complex data-feature associations into lithology-related and tectonic-related components. The CP decomposition (CANDECOMP / PARAFAC) algorithm is used to decompose the original tensor R into the sum of two low-rank tensors: R = R_lithology + R_structure + E, where E is the residual. During the decomposition process, the type labels of feature groups guide the decomposition direction: tensor elements related to lithology features (resistivity stratification, density interfaces) are aggregated into R_lithology, and elements related to tectonic features (reflection interfaces, faults) are aggregated into R_structure. The decomposition optimization objective is min||R - R_lithology - R_structure||_F, solved iteratively using alternating least squares. The lithology-related component R_lithology retains information such as the correlation strength, spatial distribution, and vertical variation between data blocks and lithology features. The tectonic-related component R_structure reflects the matching degree between data blocks and tectonic features, structural strike, dip angle, and other parameters. The decomposition convergence criterion is set to a relative error of less than 0.01, with typical convergence within 50 iterations. The physical meaning of the decomposition results is explained through the correlation attributes: high-value regions indicate a strong correlation between the data and the features, and tensor slices show correlation patterns at different depths or over time.

[0036] Formation compaction correction is applied to lithology-related components to eliminate the influence of burial depth on lithology parameters and restore the true lithology field. First, depth information is obtained: by tracing back to the cloud data lake using the data block identifier recorded in the correlation tensor, the WGS84 coordinates of the metadata header are queried, where the z-coordinate is the elevation. For seismic data blocks, two-way travel time and layer velocity parameters are extracted from the characteristic data area, and a time-depth conversion is performed to obtain the structural depth. The conversion formula is Depth=Σ(V_i×ΔT_i) / 2, where V_i is the layer velocity of the i-th layer, and ΔT_i is the two-way travel time of that layer. When establishing the normal compaction trend, density-depth data pairs are extracted from the correlated gravity data blocks, and an exponential model is fitted: ρ(z)=ρ0+(ρ_max-ρ0)(1-e^(-z / λ)), where ρ0 is the surface density (statistically obtained as 2.0 g / cm³ from shallow data). ³ ), ρ_max is the maximum density (statistically obtained from deep data: 2.7 g / cm³). ³λ is the compaction coefficient (obtained by least-squares fitting, typical value 800m). Correction is performed on the spatial point corresponding to each non-zero element in the lithological correlation component: first, the data block and feature location are found using tensor coordinates, and the observed lithological parameters and burial depth of that point are obtained; then, the decompaction value ρ_corrected = ρ_observed - ρ_compaction(z) is calculated. Resistivity correction uses a depth variant of the Archie formula, and velocity correction uses the Gardner relation. The corrected parameters are interpolated on a 3D grid (100m × 100m × 50m), with weights determined by the correlation strength (tensor element values) and the data quality level (looked up from the index table). The actual lithological field output is a 3D array, with each grid point storing the corrected density, resistivity, velocity values, and their uncertainties. The uncertainties are calculated from the data quality level and spatial interpolation error propagation.

[0037] Stress field inversion is performed on the tectonic-related components to infer the paleotectonic stress field based on present-day tectonic features. First, quantitative tectonic parameters are obtained by backtracking through the tectonic-related components: for each non-zero element in the tensor, a feature group is located using a feature index, and quantitative information such as fault displacement, fault dip angle, and fold curvature is extracted from the specific feature parameters of the group. These parameters were calculated from the feature data region of the original data block during feature extraction. The stress field governing equations adopt static equilibrium. The boundary conditions are set as follows: shear stress τ = μ × σ_n on the fault plane (μ is the friction coefficient 0.6, σ_n is the normal stress on the fault plane), and the principal stress direction on the fold axial plane is parallel to the axial plane. The finite element mesh is inherited from the real lithological field, and the elastic modulus E and Poisson's ratio ν of each element are calculated using the density and velocity of the lithological field: E = ρ × V_p ² ×(1+ν)(1-2ν) / (1-ν), where ν is determined by the ratio of V_p / V_s. The inversion employs constrained optimization, with the objective function being minΣ[(d_obs-d_calc)]. ² / σ_d ² ], where d_obs are the observation construction parameters (displacement, curvature, etc.), d_calc is the theoretical deformation generated by the stress field, and σ_d is the observation error (determined by data quality).

[0038] By superimposing the real lithological field and the paleotectonic stress field according to their physical coupling relationship, a comprehensive data field strength distribution map is generated. The field strength calculation formula is F = F_lithology + F_structure + F_coupling, where F_lithology is derived from the real lithological field (normalized lithological parameter gradient), F_structure is derived from the stress field (normalized value of maximum shear stress), and F_coupling is the coupling term, calculated using the following formula: , where k is the coupling coefficient (0.1-0.3). The density gradient is represented by σ_max, and the maximum principal stress is represented by σ_max. Field strength values ​​are normalized within the range of 0-1, and their physical meaning is a comprehensive index of geological activity intensity. Spatially, the field strength is defined in a three-dimensional grid, with the field strength value at each grid point integrating the lithological characteristics, tectonic stress, and their interaction at that point. The temporal dimension reflects the evolutionary process through multi-period overlay, with field strength maps from different periods showcasing the migration and evolution of geological activity. Visualization of the field strength distribution map uses volumetric rendering technology; color mapping represents the magnitude of the field strength, and transparency indicates data quality (high-quality data areas are opaque). High field strength areas (F>0.7) often correspond to geologically active zones and are favorable areas for mineral resource and oil and gas accumulation.

[0039] The generated field strength distribution map is subjected to equipotential line tracing to identify field strength gradient characteristics and locate relevant project groups with clear geological significance. First, the field strength gradient is calculated. The solution is obtained numerically at grid points using a central difference scheme. Equipotential surfaces are extracted using the MarchingCubes algorithm, with an equipotential value sequence F = [0.3, 0.4, 0.5, 0.6, 0.7, 0.8] to generate a series of nested equipotential surfaces. Equipotential lines are the intersections of equipotential surfaces with a specific plane (horizontal slice or vertical profile), extracted using a contour tracing algorithm. Gradient streamlines are traced from high-field-strength seed points along the steepest descent direction, with convergence regions identified as convergence centers. Seed points are selected as local maxima of field strength (F > 0.8 and the largest in the 26-neighborhood), typically identifying 20-50 seed points. Related project clusters are spatial regions with similar field strength patterns and connected gradient streamlines, identified through streamline density clustering. The clustering algorithm used is DBSCAN, with streamline density as a feature, a minimum of 10 streamlines, and a neighborhood radius of 500m. Each project group records the spatial extent (bounding box coordinates), average field strength F_avg, area ratio Area_ratio, and data quality.

[0040] Step S140: Establish a data pulsation monitoring network within the relevant project group, capture update frequency fluctuations through the data pulsation monitoring network to form a respiratory rhythm map, perform time-series correlation analysis on the respiratory rhythm map and geological feature groups to obtain active data sources, and perform interactive analysis on the active data sources to refine the collaborative triggering mechanism.

[0041] Establish a data pulsation monitoring network within the relevant project clusters. A distributed architecture is adopted, deploying monitoring nodes in a cloud-based data lake. Each node is responsible for tracking the status of 100-500 data blocks. The number of nodes is dynamically allocated based on the area ratio (Area_ratio) of the project cluster: 5-8 nodes for large project clusters with Area_ratio > 0.3, 3-5 nodes for medium-sized clusters with Area_ratio 0.1-0.3, and 1-3 nodes for small clusters with Area_ratio < 0.1. Monitoring indicators include: data access frequency (extracted from data lake access logs), data update events (new data block addition, changes in the quality level of existing data blocks), and changes in correlation relationships (updates of correlation tensors caused by the addition of new data). The monitoring cycle is set based on the average field strength (F_avg) of the project cluster: hourly monitoring for high field strength clusters (F_avg > 0.7), daily monitoring for medium field strength clusters (0.4-0.7), and weekly monitoring for low field strength clusters (< 0.4). Each monitoring node maintains a time-series database, with records formatted as {timestamp, block_id, event_type, event_value}. `event_type` includes three categories: access, update, and link_change. `event_value` represents the specific numerical value of the event (e.g., number of accesses, amount of updated data, number of new links). Data pulsation is defined as the event frequency per unit time, calculated as P(t) = N_events(t) / Δt, where N_events is the total number of events within the time window Δt. The deployment density of the monitoring network is optimized using the data quality of the project group: high-quality groups (Data_quality > 0.8) use intensive monitoring (node ​​coverage > 80%), medium-quality groups (0.5-0.8) use standard monitoring (coverage 50-80%), and low-quality groups (< 0.5) use sparse monitoring (coverage 30-50%).

[0042] A respiratory rhythm graph is generated by capturing update frequency fluctuations through a data pulsation monitoring network. First, the event sequence of each data block is statistically analyzed: access events are weighted at 0.3, update events at 0.5, and associated changes at 0.2. The overall frequency F_update = Σ(w_i × f_i), where w_i is the event weight and f_i is the frequency of each type of event (times / hour). Fluctuation analysis employs wavelet transform, decomposing the update frequency time series into multi-scale components. The mother wavelet is the Morlet wavelet, adapted to quasi-periodic signals. Wavelet coefficients display frequency fluctuation characteristics in the time-frequency domain, identifying typical rhythms such as daily (24-hour), weekly (7-day), and monthly (30-day) cycles. The respiratory rhythm graph is designed as a two-dimensional time-frequency graph, with the horizontal axis representing time (covering the most recent 90 days) and the vertical axis representing frequency components (cycles from hours to months). Colors represent fluctuation intensity (normalized to 0-1). An independent rhythm graph is generated for each project cluster, with key events (such as large-scale data updates and abnormal access peaks) marked on the graph. Rhythm feature extraction includes: dominant period T_main (the period corresponding to the frequency component with the highest energy), rhythm strength R_strength (the proportion of periodic component energy to total energy), and phase information φ (peak and trough times).

[0043] In some embodiments, the step of performing time-series correlation analysis on the respiratory rhythm map and the geological feature group to obtain active data sources includes: extracting a data update periodic spectrum from the respiratory rhythm map; extracting a time-series evolution pattern from the geological feature group; performing cross-correlation analysis on the update periodic spectrum and the time-series evolution pattern to generate a synchronization coefficient; and determining data sources with synchronization coefficients greater than a preset threshold as active data sources.

[0044] Data is extracted from the respiratory rhythm graph to update the periodic spectrum. Power spectral density of each periodic component is calculated by performing power spectral analysis on the wavelet coefficients of the respiratory rhythm graph. A peak detection algorithm is used to identify the main periods, retaining peaks with power exceeding 1.5 times the average value, typically identifying 3-5 dominant periods. Each periodic component records its period length T_i (in hours), relative intensity A_i (normalized power 0-1), and phase φ_i (0-2π), forming a periodic spectrum vector S_update=[(T_1, A_1, φ_1), (T_2, A_2, φ_2), ...]. Periodic stability is evaluated using a 30-day sliding window analysis; the coefficient of variation for stable periods is less than 0.2. The periodic spectrum at the project cluster level is obtained by weighted averaging of the periodic spectra of all data blocks within the cluster. The weights are determined by the quality level of the data block (looked up from the identifier index table) and its spatial location within the cluster (closer to the cluster center, the greater the weight).

[0045] Temporal evolution patterns are extracted from geological feature clusters. These clusters encompass feature types, spatial extent, temporal span, and specific feature parameters. Temporal evolution analysis extracts the time-varying patterns of different feature types. For tectonic feature clusters, the time series of fault activity intensity parameters (displacement rate mm / year) and stress field direction changes (degrees / million years) are analyzed. For lithological feature clusters, evolution curves of sedimentation rate (m / millennium) and lithofacies interface migration velocity are extracted. For fluid feature clusters, the dynamic processes of fluid pressure gradient and saturation changes are tracked. Evolutionary patterns employ piecewise linear fitting, segmenting the data at the moment of geological events (identified by abrupt changes in feature parameters). The slope k_i of each segment represents the rate of change during that period. Long-period analysis uses wavelet transform but with a geological timescale (millennium to million years) to extract periodic oscillations in geological processes. The evolutionary pattern vector S_geo = [(T'_1, A'_1, φ'_1), (T'_2, A'_2, φ'_2), ...], where T'_i is the geological period, A'_i is the amplitude (normalized), and φ'_i is the phase. The timescale mapping is achieved through the logarithmic transformation log(T_data) = a × log(T_geo) + b, and the coefficients a and b are calibrated according to the correspondence between known geological events and historical data updates.

[0046] Cross-correlation analysis was performed between the data update cycle spectrum and the geological time-series evolution model to identify the synchronous relationship between data activity and geological processes. First, time scale alignment was performed, transforming the geological cycle T' to the data time scale using a calibrated mapping relationship. The cross-correlation function was defined as R(τ) = Σ[S_update(t) × S_geo(t+τ)] / √[ΣS_update(t)]. ² (t)×ΣS_geo ² [(t)], where τ is the time delay (ranging from -30 days to +30 days). The synchronization coefficient is the maximum value of the cross-correlation function: Sync = max[R(τ)], and the delay τ_max at which the maximum value is reached is also recorded. Multi-scale correlation analysis calculates the synchronization coefficient for each pair of periodic components (T_i, T'_j) to obtain an M×N synchronization coefficient matrix [Sync_ij]. The preset threshold adopts an adaptive method: Threshold = μ_null + 2σ_null, where μ_null and σ_null are the mean and standard deviation of the null hypothesis distribution, with a typical threshold in the range of 0.3-0.4. Data blocks with synchronization coefficients exceeding the threshold are located by their identifier codes. These data blocks and their associated geological features constitute active data sources. Active data source records include: data block identifier code, synchronization coefficient value, associated geological feature type, optimal time delay τ_max, and synchronized periodic pairs (T_i, T'_j).

[0047] Deep interactive analysis is performed on the identified active data sources to extract collaborative triggering mechanisms. First, a causal network of active data sources is constructed, where nodes represent active data sources (represented by identifier codes), and directed edges represent causal relationships. Causal relationships are determined using the Granger causality test: a vector autoregressive model (VAR(p)) is built for the update time series of data sources i and j. If incorporating historical values ​​of i significantly reduces the prediction error of j (F-test p < 0.05), then a Granger causal relationship exists between i and j. Causal strength is defined as the proportion of prediction error reduction: G_ij = (MSE_j - MSE_j|i) / MSE_j, where MSE is the mean squared error. The causal network is represented as a directed graph, with edge weights representing Granger causality strength (range 0-1). Collaborative patterns are identified through network community detection. The Louvain algorithm is used to cluster data sources with close causal relationships into collaborative groups, with a modularity Q > 0.3 indicating a significant community structure. Each collaborative group represents synchronous or cascading updates of data sources within the group, typically containing 3-10 data sources. Triggering mechanism analysis identifies key nodes in the network: triggering sources (high out-degree, affecting multiple other nodes), relay nodes (high betweenness centrality, connecting different groups), and response nodes (high in-degree, influenced by multiple nodes). Propagation path analysis calculates the shortest path and average propagation delay from the triggering source to the response node. Cooperative triggering mechanisms are refined into rule form: IF (triggering condition) THEN (cascaded update sequence) WITH (time delay, confidence level). Triggering conditions are associated with specific geological events (such as "tectonic stress field mutation"), and cascaded sequences describe the chain update order of the data source. The final output includes 5-10 main cooperative triggering mechanisms, each mapping to a specific geological-data coupling process, forming a data-driven framework for understanding geological processes.

[0048] Step S150: Use the collaborative triggering mechanism to perform synchronous reorganization on the relevant project group to obtain the evolution sequence, perform curvature analysis on the evolution sequence to obtain the geological evolution trajectory, and extract the inflection point of the geological evolution trajectory to generate a trend vector.

[0049] Specifically, a collaborative triggering mechanism is used to synchronously reorganize related project groups to obtain evolutionary sequences. The collaborative triggering mechanism (IF-THEN-WITH rule) controls the entire processing flow: the triggering condition determines the data selection strategy; when "tectonic stress field mutation" is triggered, a short 7-day window is used to capture rapid changes, and when "deposition rate anomaly" is triggered, a long 90-day window is used to reflect slow evolution; the cascaded update sequence specifies the data fusion order, and corresponding data blocks are extracted from the cloud data lake according to the sequence, with the processing results of the preceding data source serving as the input weight adjustment factor for the subsequent data source; the time delay parameter (milliseconds to days) is directly mapped to the step size design of the sliding window to ensure the synchronous alignment of asynchronous data; the confidence level (0.3-0.9) not only participates in the fusion weight calculation (working together with the data quality level A 1.0 / B 0.8 / C 0.6 and the spatial decay factor), but also dynamically adjusts the anomaly detection threshold, with the high-confidence mechanism using a strict 3-standard-deviation approach and the low-confidence mechanism relaxing it to 2-standard-deviation. The entire synchronous reorganization process, driven by a mechanism, calculates the statistical characteristics and rate of change of density ρ, P-wave velocity Vp, and porosity φ for each window. The dominant spatial mode is extracted through empirical orthogonal function decomposition, and finally a structured evolution sequence E(t) is generated: {timestamp, [parameter mean], [rate of change], spatial modal coefficient, anomaly marker}. Each record is associated with its triggering mechanism identifier, realizing a mechanism-driven data-geological process mapping.

[0050] In some embodiments, performing curvature analysis on the evolutionary sequence to obtain the geological evolution trajectory includes: calibrating the evolutionary sequence on a geological timescale to form a geological age sequence; identifying sedimentary discontinuities and unconformities on the geological age sequence; performing segmented fitting with the sedimentary discontinuities and unconformities as boundaries to generate a stratigraphic sedimentation rate curve; and performing curvature analysis on the sedimentation rate curve to obtain the geological evolution trajectory.

[0051] Geological timescale calibration is performed on the evolutionary sequence to form a geological chronology. The time calibration is based on the geological principle of "interpreting the past through the present," employing the power law relationship T_geo = α × T_obs^β, where T_geo is the geological time (years), T_obs is the observation time (days), the coefficient α is set to 10000 to reflect differences in timescales, and the exponent β is set to 0.7 to represent nonlinear compression. These coefficients are calibrated by matching known geological events (such as historical earthquakes and volcanic activity) with current process rates. Calibration not only transforms the time axis but, more importantly, preserves the geological information within the evolutionary sequence. Depositional parameters are derived from the density change rate r_ρ and the porosity change rate r_φ: the compaction rate is calculated through porosity reduction, and the deposition rate v_s comprehensively considers density increase (new sediments) and compaction. In the calculation, the particle density is taken as the standard value of quartz, 2.65 g / cm³. ³The geological time series G(T) is constructed as a complete record: {geological time T_geo, [density ρ(T), velocity Vp(T), porosity φ(T)], sedimentation rate v_s(T), dominant spatial mode}, where each element has been transformed over time but retains its physical meaning. A variable resolution strategy is employed: the Quaternary (0-2.6 Ma) maintains millennial-level resolution to capture rapid changes, while the Neogene (2.6-23 Ma) decreases to tens of thousands of years to adapt to slower processes.

[0052] Identify key geological interfaces on the calibrated geological time series. Criteria for identifying sedimentary discontinuities: a sedimentary rate (v_s) consistently below 0.1 mm / ka for more than 10 ka, while porosity remains constant (no compaction), indicating the cessation of sedimentation. Angular unconformity criteria: a density jump exceeding 0.3 g / cm³. ³ Furthermore, velocity jumps exceeding 500 m / s reflect significant lithological differences between upper and lower strata, typically accompanied by long-term erosion. Parallel unconformities are characterized by abrupt velocity changes but gradual density variations, indicating sedimentary discontinuities without significant erosion. Abrupt change detection employs the Cumulative Summation (CUSUM) algorithm, capable of identifying abrupt change points in the gradual transition process, with a detection threshold set at 3 times the background noise level. Spatial continuity is verified through the dominant spatial mode, requiring interfaces to be traceable across the entire project group; isolated abrupt changes are considered local anomalies. Detailed records are kept for each identified interface, including: geological age (error range), interface type, abrupt changes in various parameters, spatial distribution, and inferred formation mechanism (tectonic uplift, sea-level drop, etc.).

[0053] The stratigraphic sedimentary rate curves are generated by segmenting and fitting the identified geological interfaces as natural segment boundaries. The segmentation strategy follows the inherent laws of geological processes, ensuring that each segment represents a relatively homogeneous sedimentary environment. Sedimentary rate data points for each segment are extracted from the geological time series G(T), and the fitting model is selected based on the data distribution characteristics and geological background. An exponential decay model is used for stable marine sedimentation to reflect a gradually slowing sedimentary process; a power-law model is used for rapid terrestrial deposition to adapt to varying sedimentary conditions; and piecewise linear or cubic spline models are used for transitional environments. Physical constraints are imposed on the fitting process: non-negative rates (negative sedimentation is not allowed), interface continuity (rates can jump at the interface between adjacent segments but are continuous in time), and total thickness conservation (the integral of the fitted curve equals the observed thickness). The fitting optimization uses the weighted least squares method, with weights set according to the uncertainty of the data points. Feature parameters are extracted for each segment: initial sedimentary rate v_0 (typical value 1-100 m / Ma), characteristic time scale (decay constant or transition time), average rate, and rate change trend. Inter-section comparative analysis identifies the evolution of sedimentary systems: a decrease in sedimentary velocity from high to low velocity may indicate a reduction in sediment source, while a sudden increase in velocity may reflect intensified tectonic activity.

[0054] A curvature analysis was performed on the generated sedimentation rate curves to obtain the geological evolution trajectory through curvature characteristics. A numerical differential method was used to first smooth the piecewise rate curves v_s(T) using cubic splines to eliminate numerical noise, and then the first derivative v_s'(T) and second derivative v_s''(T) were calculated. The curvature formula is κ(T) = |v_s''(T)| / [1 + v_s'(T)]. ² The curvature of the curve at each time point is quantified, and its physical meaning is the normalized value of the acceleration of the sedimentation rate change. High curvature sections (κ>κ_mean+2σ) correspond to periods of rapid change in the sedimentary environment, such as geological events like river diversion, marine transgression and regression, and tectonic uplift. Low curvature sections represent stable sedimentation periods, with slow and gradual rate changes. The temporal distribution of curvature shows obvious stages: curvature near discontinuities approaches infinity (rate abrupt change), while curvature during stable sedimentation periods approaches zero (rate constant). The geological evolution trajectory is constructed in the rate phase space, selecting three dimensions: sedimentation rate v_s(T), rate of change v_s'(T), and rate acceleration v_s''(T), forming a parameterized trajectory r(T)=[v_s(T), v_s'(T), v_s''(T)]. This representation method extends the one-dimensional rate curve into a three-dimensional dynamic trajectory, fully demonstrating the dynamic evolution characteristics of the sedimentation process. The movement patterns of trajectories in phase space have clear geological significance: spiral ascents indicate accelerated deposition, circular tracks correspond to periodic deposition, and straight segments indicate uniform deposition. The geometric characteristics of trajectories quantitatively describe evolutionary complexity: arc length integrals reflect the total change, curvature integrals quantify path complexity, and trajectory envelope volume characterizes the range of parameter variations. Stable attractors manifest as convergence points of trajectories, representing equilibrium states under specific depositional conditions; trajectories diverge near unstable points, indicating an impending transformation of the depositional system.

[0055] Inflection points are extracted from the geological evolution trajectory to generate trend vectors. Inflection point detection is performed on the velocity phase space trajectory r(T) = [v_s(T), v_s'(T), v_s''(T)], and the tangent vector dr / dT and acceleration vector d²r / dT² are calculated. The inflection point is defined using a geometric criterion: the magnitude of acceleration |d ² r / dT ² | is a local maximum point, or a location where the tangent vector changes direction by more than 60 degrees. Each type of inflection point corresponds to a specific sedimentary transition: inflection points of the v_s component indicate a reversal of the sedimentary rate trend (e.g., from acceleration to deceleration), inflection points of the v_s' component indicate a change in sedimentary acceleration (e.g., from uniform to accelerated), and inflection points of the v_s'' component reflect higher-order changes in the sedimentary process. The geological interpretation of inflection points is combined with the original rate curve: upward convex inflection points often correspond to a decrease in source material or a rise in the base level, while downward concave inflection points often indicate an increase in source material or accelerated tectonic subsidence. The verification algorithm is achieved by checking the trajectory torsion τ=(dr / dT×d ² r / dT² )·d ³ r / dT ³ / |dr / dT×d ² r / dT ² | ² High torsion confirms the true torsion of the space curve. Each inflection point records complete information: time T_i, phase space position r_i, and forward and backward tangent vectors v_i. - and v_i + The turning angle Δθ_i and the corresponding depositional rate value v_s(T_i) are used. The trend vector is defined as the normalized evolution direction between adjacent turning points: V_trend,j=(r_{i+1}-r_i) / |r_{i+1}-r_i|, representing the dominant evolutionary path in the rate phase space during that period. The direction of the vector in phase space has a clear meaning: pointing towards the high v_s indicates an accelerated depositional trend, and pointing towards the high |v_s'| indicates an increase in depositional instability. The spatial distribution of the project cluster is obtained by analyzing different locations separately. Each sub-region may have different evolutionary trajectories and trend vectors. The spatial vector field V(x, y, z) is constructed by interpolating the trend vectors at different locations, and its divergence and curl analysis identifies the regional depositional center and rotation mode. Temporal evolution is recorded through the trend vector sequence {(T_j, V_trend,j)}, showing the stage transitions of the depositional system. Statistical characteristics include: average trend direction (actively induced model), directional dispersion (evolutionary stability), and anisotropy of the vector field (spatial variability). The final output trend vector set quantitatively describes the directional characteristics of geological processes based on sedimentary rate evolution, completing a comprehensive analysis from rate curves to evolutionary trends.

[0056] Step S160: The trend vector and geological feature groups are cross-validated layer by layer to form a confidence gradient field. The confidence gradient field is used to trace the path to obtain the predicted path bundle. The predicted path bundle is then used to perform convergence analysis to deduce exploration hotspot areas.

[0057] Specifically, a confidence gradient field is formed by performing layer-by-layer cross-validation between the trend vector and geological feature groups. The trend vector consists of a time-stamped vector set {(T_j, V_trend, j)} and a spatial vector field V(x, y, z). The geological feature groups are organized hierarchically into three categories: structural features, lithological features, and fluid features. Layer-by-layer validation is first performed at the feature type level: the direction of the trend vector is compared with the fault strike and fold axis in the structural feature group, with an angle less than 30 degrees considered consistent; it is compared with the sedimentary interface dip in the lithological feature group to assess the rationality of the sedimentary trend; and it is matched with the migration direction of the fluid feature group to verify the response of the fluid system. The second layer of validation is at the spatial distribution level: the trend vector field V(x, y, z) is superimposed with the spatial range of the feature group to calculate the spatial overlap and directional consistency. The third layer of validation is at the temporal evolution level: the time series of the trend vector is compared with the time span of the feature group to ensure the rationality of the temporal logic of the evolution. Cross-validation employs fuzzy logic, with a consistency score C_ij = exp(-θ_ij / θ_0) × S_overlap × T_corr, where θ_ij is the angle between the vector and the feature, θ_0 = 30° is the feature angle, S_overlap is the spatial overlap (0-1), and T_corr is the temporal correlation coefficient. Confidence is calculated by integrating the three-layer validation results: Conf(x, y, z) = Σ(w_k × C_k), where C_k is the consistency score of the k-th layer validation, and w_k is the weight of each layer (feature type 0.5, spatial distribution 0.3, temporal evolution 0.2). Confidence values ​​are calculated on a 3D spatial grid with a resolution consistent with the fused data stream (100m × 100m × 50m). High-confidence regions (Conf > 0.7) indicate that the trend vector is supported by multiple geological features, while low-confidence regions may indicate cognitive bias or insufficient data. Confidence gradient. The gradient magnitude, calculated using the central difference method, reflects the spatial rate of change of confidence, while the gradient direction points in the direction of the fastest increase in confidence.

[0058] In some embodiments, the step of performing path tracing on the confidence gradient field to obtain a predicted path bundle includes: performing anisotropic diffusion on the confidence gradient field to generate a smooth gradient field; tracing high gradient ridges that satisfy formation continuity in the smooth gradient field; setting a search starting point on the high gradient ridges; and sorting the high gradient ridges by confidence starting from the search starting point to form a predicted path bundle.

[0059] Anisotropic diffusion is performed on the confidence gradient field to eliminate local noise while preserving important gradient features, generating a smooth gradient field suitable for path tracing. Anisotropic diffusion employs the Perona-Malik equation: The diffusion coefficient g(s) = 1 / (1+s) ² / K ²), where s is the gradient magnitude. K is the edge threshold, taken as 1.5 times the standard deviation of the gradient field. During diffusion, regions with small gradients (flat areas) experience strong diffusion for smoothing, while regions with large gradients (edges) experience suppressed diffusion to preserve characteristics. An iteration time step Δt = 0.1 ensures numerical stability, and 20-50 iterations achieve smoothing. The anisotropy of diffusion is achieved by constructing a diffusion tensor D, with weak diffusion along the gradient direction and strong diffusion perpendicular to the gradient direction, maintaining the sharpness of the ridges. The Neumann condition is used as the boundary condition to ensure the mass conservation of the gradient field. The diffused gradient field retains the main ridge and valley structures while eliminating spurious extrema caused by data noise. Smoothing the gradient field. This provides stable terrain for subsequent path tracking.

[0060] Tracing high-gradient ridges in a smooth gradient field involves identifying ridges that represent boundaries with rapidly changing confidence levels, often corresponding to important geological transition zones. A ridge is defined as a line connecting local maxima of the gradient field, satisfying two conditions: gradient magnitude... It reaches a local maximum in the direction perpendicular to the ridge line, in the gradient direction. The ridge line continuously changes along its length. Ridge tracking employs a prediction-correction algorithm: starting from a seed point, it predicts the next point along the direction perpendicular to the current gradient, and then corrects to the accurate ridge position through gradient ascent. The tracking step size is adaptively adjusted, decreasing in regions with high curvature (minimum 10m) and increasing in straight segments (maximum 100m). Branching strategy: when encountering a ridge bifurcation, all branches are recorded and tracked separately, forming a ridge network. Termination conditions include: reaching the boundary, gradient value below a threshold (0.1), and forming a closed loop. Each ridge line is recorded as a point sequence {P_i}, with attributes including the confidence value, gradient magnitude, and local curvature of each point. The topological relationship of the ridge lines is represented by a graph structure, with nodes being bifurcation points and endpoints, and edges being ridge segments. Typically, 20-50 major ridge lines are identified, covering the skeletal structure of the entire study area.

[0061] For example, tracing high gradient ridges that satisfy formation continuity in the smooth gradient field includes: setting constraints in the smooth gradient field, the constraints including formation dip angle constraints and fault displacement constraints; performing geodesic tracing based on the constraints to generate a candidate path set; and applying a formation thickness consistency check to the candidate path set to generate high gradient ridges.

[0062] First, set geological constraints in the smoothed gradient field. The formation dip constraint limits the local inclination of the path: |dz / ds| < tan(θ_max), where s is the path arc length, z is the vertical coordinate, and θ_max is the maximum regional formation dip (typically 15 - 45°), which is extracted from the structural features of the geological feature population. The fault displacement constraint deals with discontinuities: when the path encounters a fault (identified from the geological feature population), vertical jumps are allowed but horizontal displacements are restricted, with the displacement not exceeding the maximum throw of the fault. The path curvature constraint prevents unreasonable sharp turns: |dT / ds| < κ_max, where T is the tangent vector and κ_max is set according to the degree of formation folding. Then, geodesic tracing based on the constraints uses a constrained optimization method, with the objective function being to maximize the confidence integral of the path and the constraints being the above geological conditions. The solution uses dynamic programming, discretizes the three-dimensional space, and evaluates reachability and cost at each grid point. The candidate path set contains multiple suboptimal paths that satisfy the constraints, not just a single optimal solution, providing a range of prediction uncertainties. Finally, apply a formation thickness consistency test to the candidate path set and eliminate geologically unreasonable paths. The test principle is based on the lateral continuity of formation thickness: the thickness change of the same formation should be gradual over a short distance. The thickness is extracted by calculating the vertical distance between adjacent formation interfaces (obtained from the lithological features of the geological feature population) through which the path passes. The consistency measure uses the coefficient of variation: CV = σ_h / μ_h, where σ_h is the standard deviation of thickness and μ_h is the average thickness, and CV < 0.3 is considered consistent. Anomalous thickness changes require reasonable explanations: an increase in thickness is allowed in the core of a fold, and sudden changes are allowed near faults; otherwise, the path is considered incorrect. The paths passing the test retain their complete attributes: the path coordinate sequence, the sequence of formations traversed, the thickness change curve, the cumulative confidence, etc. The elimination rate is usually 30 - 50%, and finally 10 - 30 geologically reasonable high-gradient ridge lines are retained.

[0063] Starting points are set along high-gradient ridges, chosen at locations with clear geological significance and high confidence. The selection strategy considers multiple factors: priority is given to intersections of the ridge with known geological markers (such as exposed strata or well locations), as these locations offer the strongest constraints; secondly, points with local confidence maxima along the ridge represent the most reliable prediction locations; and thirdly, bifurcation and inflection points of the ridge are also considered, as these often correspond to key nodes in the geological system. Two to five starting points are set for each ridge, avoiding over- or under-density. The starting points inherit ridge information and are supplemented with: 3D coordinates, local strata, local geological features, confidence value, and search priority. High-gradient ridges are then sorted by cumulative confidence value, starting from the search starting points, forming an ordered bundle of prediction paths. Path confidence is calculated using a three-dimensional confidence field Conf(x, y, z) generated along the first segment of the path integral: Conf_path = ∫Conf(s)ds, where s is the path parameter, Conf(s) is the confidence value at point s on the path, and the integral is performed along the ridge path from the starting point to the ending point, reflecting the overall reliability of the area traversed by the entire path. The sorted path bundles are divided into three levels: high-confidence paths (Conf_path > 0.8) are the primary predictions, medium-confidence paths (0.5-0.8) are alternatives, and low-confidence paths (< 0.5) are for reference only. The path bundles are organized in a tree structure: the root node is the starting point, branches are extension paths in different directions, and leaf nodes are the path endpoints. Each level typically contains 5-10 paths, totaling 20-30 predicted paths, forming a predicted path bundle covering the main geological trends.

[0064] Convergence analysis is performed on the generated predicted path bundles to deduce the most likely exploration hotspot areas. First, the path density field is calculated: each path is spatially diffused using a Gaussian kernel function with a kernel width σ = 200 m, and all paths are superimposed to obtain the density field ρ_path(x, y, z). High-density regions represent the intersection points of multiple independent paths, indicating high prediction reliability. Convergence metrics include: path number density (number of paths per unit volume), directional consistency (average angle between path tangent vectors), and confidence concentration (proportion of high-confidence paths). Hotspot identification uses cluster analysis; connected regions in the density field greater than twice the average value are defined as candidate hotspots. Each candidate hotspot is scored using the following comprehensive score: Score = w1 × ρ_norm + w2 × Dir_consistency + w3 × Conf_concentration, where ρ_norm is the normalized path density (range 0-1), Dir_consistency is the consistency of path direction, and Conf_concentration is the proportion of high-confidence paths. Weights w1 = 0.4, w2 = 0.3, and w3 = 0.3 are used to balance the factors. The geological attributes of the hotspots are determined by backtracking path bundles: dominant strata (most frequently traversed by paths), structural location (anticline, syncline, fault block), and depth range (path Z-coordinate statistics). Spatial morphology analysis extracts the geometric features of the hotspots: volume, extension direction, depth-to-width ratio, etc., to guide exploration deployment. Finally, 5-10 exploration hotspot areas are output, each containing: center coordinates, influence range, comprehensive score, geological feature description, and suggested exploration scheme. Hotspots are prioritized according to their scores. Through convergence analysis of predicted paths, the transformation from trend prediction to exploration targets is completed.

[0065] Step S170: Perform spatial cluster analysis on exploration hotspot areas to construct a resource allocation matrix, generate a project priority sequence based on the resource allocation matrix, and output a multi-project collaborative management plan based on the project priority sequence.

[0066] In some embodiments, the step of constructing a resource allocation matrix by performing spatial cluster analysis on the exploration hotspot areas includes: performing spatial distribution analysis on the exploration hotspot areas to extract distribution features; performing density clustering based on the distribution features to form hotspot clusters; performing comprehensive evaluation on the hotspot clusters to generate allocation weights, the allocation weights including spatial distribution weights and resource demand weights; and constructing a resource allocation matrix based on the allocation weights.

[0067] Perform spatial distribution analysis on exploration hotspots, and extract distribution characteristics reflecting the spatial pattern of hotspots. The distribution analysis calculates global and local spatial statistics. Global characteristics include: the centroid of hotspots (the weighted center of all hotspots, with the weight being the comprehensive score), the distribution dispersion (the standard distance from each hotspot to the centroid), and the azimuth of the main axis (determined by the eigenvector of the coordinate covariance matrix, reflecting the dominant direction of hotspot distribution). The local characteristics use Ripley's K function to analyze the aggregation pattern of hotspots, where \(K(d)=\frac{A}{n}\) ² \(\times\sum_{i}\sum_{j}I(d_{ij}<d)\), where \(d\) is the distance scale for analysis, \(A\) is the area of the study area, \(n\) is the number of hotspots, \(I\) is the indicator function, and \(d_{ij}\) is the distance between hotspot \(i\) and \(j\). The deviation of the K function curve from the random distribution indicates spatial aggregation or dispersion. The nearest neighbor analysis calculates the distance from each hotspot to the nearest hotspot. A ratio \(R\) of the average nearest neighbor distance to the random expectation less than 1 indicates an aggregated distribution. Spatial autocorrelation is evaluated by Moran's I index, taking the comprehensive score of hotspots as the attribute value. A positive value indicates that hotspots with similar scores tend to aggregate. The distribution feature vector is organized as: [centroid coordinates \((x_c, y_c, z_c)\), dispersion \(\sigma\), azimuth of the main axis \(\theta\), peak value of the K function \(d_{peak}\), nearest neighbor index \(R\), Moran's I value], quantitatively characterizing the spatial pattern characteristics of hotspots.

[0068] Based on the extracted spatial distribution characteristics, use the density clustering algorithm to aggregate relevant hotspots into hotspot clusters. The density clustering selects the DBSCAN algorithm, which can identify clusters of any shape and handle noise points. The core parameters are set according to the distribution characteristics: the neighborhood radius \(\epsilon\) is taken as 1.5 times the average nearest neighbor distance (calculated from the nearest neighbor index \(R\)) to ensure that naturally aggregated hotspots are grouped into one cluster; when the dispersion \(\sigma\) is large, increase the value of \(\epsilon\) to adapt to the dispersed distribution. The minimum number of points is set to 2, allowing two adjacent hotspots to form a cluster. The distance metric combines spatial distance and attribute similarity: \(d_{combined}=\alpha\times d_{spatial}+\beta\times d_{attribute}\), where \(d_{spatial}\) is the shortest distance between hotspot boundaries, and \(d_{attribute}\) is the text similarity based on geological feature descriptions. The weights \(\alpha = 0.7\) and \(\beta = 0.3\) emphasize spatial proximity. When clustering, consider the azimuth of the main axis \(\theta\), relax the distance threshold along the main axis direction by 1.2 times, and keep it strict in the direction perpendicular to the main axis to adapt to the directional distribution of hotspots. For regions with a high Moran's I value, lower the clustering threshold to promote the aggregation of similar hotspots. The clustering result contains 3 - 6 hotspot clusters, and each cluster aggregates 2 - 4 original hotspots. The attributes of the clusters are synthesized from member hotspots: the spatial range takes the convex hull and expands it by 10% as a buffer, the comprehensive score is weighted and averaged according to the member scores, and the geological features take the union to retain all types. Isolated hotspots (noise points) remain independent, forming a total of 5 - 8 exploration units.

[0069] A comprehensive assessment of the formed hotspot clusters is conducted to generate allocation weights to guide resource allocation. The assessment is conducted from two dimensions: exploration value and implementation conditions, with each dimension containing multiple indicators. The exploration value dimension integrates: geological potential (comprehensive score of hotspots within the cluster), target scale (the total impact range reflects the resource magnitude), expected results (inferring exploration success rate based on geological characteristics), and strategic value (whether it is located in a key metallogenic belt or oil and gas basin). The implementation conditions dimension is assessed based on: topographic conditions (judging construction difficulty from elevation standard deviation), infrastructure (distance from the nearest road and power line), environmental constraints (reducing weight within the buffer zone of the protected area), and technological maturity (exploration experience with similar geological conditions). The spatial distribution weight W_spatial reflects the cluster's locational advantage in the overall layout, calculated using the formula W_spatial=exp(-d_center / d_0)×(1+n_neighbor / n_total), where d_center is the distance from the cluster to the centroid of all clusters, d_0 is the average cluster spacing, n_neighbor is the number of neighboring clusters within 2d_0, and n_total is the total number of clusters. Clusters located at the center and with multiple neighbors receive high weights, facilitating resource sharing and the dissemination of results. The resource requirement weight W_resource is determined by a combination of value score and implementation difficulty; clusters with high value and low difficulty receive high weights, ensuring maximum efficiency of resource investment. Ultimately, each hotspot cluster receives a two-dimensional configuration weight W = [W_spatial, W_resource], ranging from 0 to 1.

[0070] A refined resource allocation matrix is ​​constructed based on the configuration weights of hotspot clusters to quantitatively describe the resource needs of each cluster. The matrix is ​​designed as an M×N structure, where M represents the number of hotspot clusters and independent hotspots (5-8), and N represents the number of resource types, including 6-8 categories such as seismic exploration teams, electromagnetic survey teams, drilling teams, geological logging teams, data processing centers, and logistics support teams. Matrix elements reflect the demand intensity of the i-th cluster for the j-th type of resource, comprehensively considering basic demand, configuration weights, and resource characteristics. Basic demand is estimated based on the cluster's exploration plan: seismic exploration demand is proportional to the coverage area (square kilometers multiplied by the survey line density coefficient), drilling demand considers the target depth and geological complexity (depth multiplied by the number of boreholes and then by the difficulty coefficient), and manpower demand is calculated based on workload and operation cycle. Configuration weights integrate spatial weights and resource weights through a geometric mean; the demand of high-weight clusters is multiplied by a priority guarantee coefficient of 1.2-1.5. A resource sharing mechanism identifies spatially adjacent clusters (distance less than 10km) and marks the shareable resource types in the matrix, such as data processing centers and logistics bases. The time dimension is reflected through phased demand: high demand for geophysical resources in the early stages of exploration, shifting to drilling in the middle stages, and focusing on data processing in the later stages. Constraints ensure feasibility: total demand for all types of resources does not exceed available resources, specialized equipment cannot serve multiple clusters simultaneously, and technical personnel must meet specific professional requirements. Through constraint optimization, an M×N resource allocation matrix that satisfies all conditions is obtained, with each element representing a standardized demand intensity between 0 and 1.

[0071] Based on a resource allocation matrix, a project priority sequence is generated to scientifically determine the implementation order of exploration projects. Priority scoring integrates four key factors, using a weighted summation formula: P_i = w1 × V_i + w2 × E_i + w3 × T_i - w4R_i. Here, V_i is the value score, directly using the normalized value of the comprehensive score of hotspot clusters; E_i is the resource efficiency score, calculated by the reciprocal of the sum of elements in the i-th row of the matrix (a smaller sum indicates lower resource demand and higher efficiency); T_i is the timeliness score, considering the time span of existing geological data, with older data indicating higher timeliness; and R_i is the risk score, combining technical risk (target depth and geological complexity) and market risk. Weights w1=0.4, w2=0.3, w3=0.2, and w4=0.1 reflect a value-oriented approach while considering efficiency. Resource dependency analysis identifies bottleneck resources through matrix column scanning; projects requiring these resources must be scheduled at different times. Priority dependencies consider the progression of geological understanding; for example, regional geophysical exploration should precede detailed drilling. Seasonal windows constrain the operation time in certain regions; for example, avoiding winter in high-altitude and cold regions, and choosing the dry season in swampy regions. A multi-constraint sorting algorithm generates a project sequence that meets all conditions. The sequence is organized for batch implementation: the first batch consists of 2-3 high-priority projects with no resource conflicts; the second batch is dynamically adjusted based on the progress and new findings of the first batch, reserving a contingency window to handle unforeseen circumstances. Each project in the sequence clearly indicates the implementation period, resource requirements, preconditions, and expected outputs, forming an actionable project priority sequence.

[0072] Based on project priority sequences, a multi-project collaborative management plan is generated to achieve optimized resource allocation and effective collaboration among projects. The management plan comprises four core modules, comprehensively covering all aspects of project implementation. The resource scheduling plan is formulated based on priority sequences and allocation matrices, clearly defining the flow paths of resources between projects: geophysical equipment is planned with shortest routes to reduce empty-run costs; drilling teams are deployed in tiers for rapid relocation after completing high-priority projects. A progress coordination mechanism sets unified milestone nodes: completion of geophysical data acquisition, first mineral (oil and gas) discovery, and submission of phased results, etc. Each project regularly reports progress, and subsequent projects dynamically adjust their plans based on previous results. A results-sharing strategy maximizes the value of exploration information: all raw data is uploaded to a cloud data lake in real time, processed results and geological insights are shared promptly, newly discovered patterns are quickly applied to other projects, and technological innovations are promoted within the project cluster. A risk joint control system identifies and manages cross-project risks: technical risks are reduced through joint expert consultations and plan optimization; financial risks are controlled through project portfolios and phased investments; and market risks are diversified through diversified objectives. The final outputs of the solution include: a Gantt chart showing the timeline and resource allocation for multiple projects; a network diagram showing the dependencies between projects; an organizational chart clarifying management levels and responsibilities; and standard operating procedures to ensure consistent execution. Through systematic collaborative management, the entire process from data analysis, target selection, resource allocation to project implementation is optimized, achieving a complete closed loop for cloud-based management and analysis of geological exploration data.

[0073] To implement the above-described method embodiments, a cloud-based management and analysis method for geological exploration data is proposed to achieve the corresponding functions and technical effects. See also... Figure 2 , Figure 2 This diagram illustrates a structural block diagram of a cloud-based geological exploration data management and analysis device 200 provided in an embodiment of this application. For ease of explanation, only the parts relevant to this embodiment are shown. The cloud-based geological exploration data management and analysis device 200 provided in this embodiment includes: The data acquisition module 201 is used to collect monitoring data from various exploration sites and to perform quality grading on the monitoring data to generate a graded data queue. The cloud storage module 202 is used to perform edge preprocessing on the hierarchical data queue to form structured data blocks, and to perform dynamic storage optimization on the structured data blocks to obtain a cloud data lake; Project association module 203 is used to perform deep scanning on the cloud data lake to extract geological feature groups, perform multi-dimensional matching between the structured data blocks and the geological feature groups to obtain association tensors, construct a data field strength distribution map based on the association tensors, and perform equipotential line tracking on the data field strength distribution map to locate related project groups; Data monitoring module 204 is used to establish a data pulsation monitoring network within the relevant project group, capture update frequency fluctuations through the data pulsation monitoring network to form a respiratory rhythm map, perform time series correlation analysis on the respiratory rhythm map and the geological feature group to obtain active data sources, and perform interactive analysis on the active data sources to refine the collaborative triggering mechanism. The trend analysis module 205 is used to perform synchronous reorganization on the relevant project group to obtain the evolution sequence using the collaborative triggering mechanism, perform curvature analysis on the evolution sequence to obtain the geological evolution trajectory, and extract inflection points from the geological evolution trajectory to generate a trend vector. The hotspot prediction module 206 is used to perform layer-by-layer cross-validation between the trend vector and the geological feature group to form a confidence gradient field, perform path tracing on the confidence gradient field to obtain a predicted path bundle, and perform convergence analysis on the predicted path bundle to deduce the exploration hotspot area. The solution output module 207 is used to perform spatial cluster analysis on the exploration hotspot areas to construct a resource allocation matrix, generate a project priority sequence based on the resource allocation matrix, and output a multi-project collaborative management solution according to the project priority sequence.

[0074] The aforementioned cloud-based geological exploration data management and analysis device 200 can implement a cloud-based geological exploration data management and analysis method according to the above-described method embodiments. The options in the above method embodiments are also applicable to this embodiment, and will not be detailed here. The remaining content of this application's embodiments can be referred to the content of the above method embodiments, and will not be repeated in this embodiment.

[0075] like Figure 3 As shown, the third embodiment of the present invention also provides a computer device, including a memory 301, a processor 302, and a computer program stored in the memory 301 and executable on the processor 302, characterized in that the processor 302 executes the program to implement the steps of the cloud management and analysis method for geological exploration data described in the first embodiment of the present invention.

[0076] The above description is only a part or preferred embodiment of this application. Neither the text nor the drawings should limit the scope of protection of this application. All equivalent structural transformations made using the content of this application's specification and drawings under the overall concept of this application, or direct / indirect applications in other related technical fields, are included within the scope of protection of this application.

Claims

1. A cloud-based management and analysis method for geological exploration data, characterized in that, include: Collect monitoring data from each exploration site, and perform quality grading on the monitoring data to generate a graded data queue; The hierarchical data queue is preprocessed at the edge to form structured data blocks, and the structured data blocks are dynamically stored and optimized to obtain a cloud data lake. A deep scan is performed on the cloud data lake to extract geological feature groups. The structured data blocks are matched with the geological feature groups in multiple dimensions to obtain the correlation tensor. A data field strength distribution map is constructed based on the correlation tensor. Isopotential line tracing is performed on the data field strength distribution map to locate related project groups. A data pulsation monitoring network is established within the relevant project group. The data pulsation monitoring network is used to capture update frequency fluctuations to form a respiratory rhythm map. The respiratory rhythm map and the geological feature group are subjected to time series correlation analysis to obtain active data sources. Interactive analysis is performed on the active data sources to extract a collaborative triggering mechanism. The collaborative triggering mechanism is used to perform synchronous reorganization on the relevant project group to obtain the evolution sequence, curvature analysis is performed on the evolution sequence to obtain the geological evolution trajectory, and inflection point extraction is performed on the geological evolution trajectory to generate a trend vector; The trend vector and the geological feature group are cross-validated layer by layer to form a confidence gradient field. The confidence gradient field is used to trace the path to obtain the predicted path bundle. The predicted path bundle is used to perform convergence analysis to deduce exploration hotspot areas. Spatial clustering analysis is performed on the exploration hotspot areas to construct a resource allocation matrix. Based on the resource allocation matrix, a project priority sequence is generated, and a multi-project collaborative management scheme is output according to the project priority sequence.

2. The method according to claim 1, characterized in that, The step of dynamically optimizing the storage of the structured data blocks to obtain a cloud data lake includes: A data fingerprint encoding is applied to the structured data block to generate a unique identifier; A dynamic storage topology graph is constructed based on the unique identifier. Perform node density analysis on the dynamic storage topology to determine the optimal storage path; Distributed storage is performed via the optimal storage path to form a cloud data lake.

3. The method according to claim 1, characterized in that, The construction of the data field strength distribution map based on the correlation tensor includes: The correlation tensor is decomposed into lithological correlation components and structural correlation components; Apply formation compaction correction to the lithology-related components to generate a true lithology field; Stress field inversion is performed on the structurally related components to form a paleotectonic stress field; The actual lithological field is superimposed with the paleotectonic stress field to generate a data field strength distribution map.

4. The method according to claim 1, characterized in that, The step of performing time-series correlation analysis on the respiratory rhythm map and the geological feature groups to obtain active data sources includes: Data is extracted from the respiratory rhythm graph to update the periodic spectrum; Extract temporal evolution patterns from the aforementioned geological feature groups; The update cycle spectrum and the time-series evolution pattern are cross-correlation analysis to generate a synchronization coefficient. Data sources with synchronization coefficients greater than a preset threshold are identified as active data sources.

5. The method according to claim 1, characterized in that, The process of performing curvature analysis on the evolutionary sequence to obtain the geological evolution trajectory includes: The evolutionary sequence is calibrated on a geological timescale to form a geological chronology sequence; Identify sedimentary discontinuities and unconformities on the stated geological time sequence; Segmented fitting is performed using the sedimentary discontinuity and the unconformity as boundaries to generate a stratigraphic sedimentation rate curve; The geological evolution trajectory is obtained by analyzing the curvature of the sedimentation rate curve.

6. The method according to claim 1, characterized in that, The step of obtaining the predicted path bundle by path tracing the confidence gradient field includes: Anisotropic diffusion is performed on the confidence gradient field to generate a smooth gradient field; Tracing high-gradient ridges that satisfy formation continuity in the smooth gradient field; Set the search starting point on the high-gradient ridge line; The high-gradient ridges are sorted by confidence level starting from the search starting point to form a predicted path bundle.

7. The method according to claim 6, characterized in that, Tracing high-gradient ridges that satisfy formation continuity in the smooth gradient field includes: Constraints are set in the smooth gradient field, including formation dip angle constraints and fault displacement constraints. Based on the constraints, geodesic tracing is performed to generate a candidate path set; A formation thickness consistency check is applied to the candidate path set to generate high gradient ridges.

8. The method according to claim 1, characterized in that, The step of constructing a resource allocation matrix by performing spatial cluster analysis on the exploration hotspot areas includes: Spatial distribution analysis was performed on the aforementioned exploration hotspot areas to extract distribution characteristics; Based on the distribution characteristics, density clustering is performed to form hotspot clusters; The hotspot clusters are comprehensively evaluated to generate configuration weights, which include spatial distribution weights and resource demand weights. A resource allocation matrix is ​​constructed based on the configuration weights.

9. A cloud-based management and analysis device for geological exploration data, characterized in that, include: The data acquisition module is used to collect monitoring data from various exploration sites and to perform quality grading on the monitoring data to generate a graded data queue. The cloud storage module is used to perform edge preprocessing on the hierarchical data queue to form structured data blocks, and to perform dynamic storage optimization on the structured data blocks to obtain a cloud data lake; The project association module is used to perform deep scanning on the cloud data lake to extract geological feature groups, perform multi-dimensional matching between the structured data blocks and the geological feature groups to obtain association tensors, construct a data field strength distribution map based on the association tensors, and perform equipotential line tracing on the data field strength distribution map to locate related project groups. The data monitoring module is used to establish a data pulsation monitoring network within the relevant project group, capture update frequency fluctuations through the data pulsation monitoring network to form a respiratory rhythm map, perform time-series correlation analysis on the respiratory rhythm map and the geological feature group to obtain active data sources, and perform interactive analysis on the active data sources to refine the collaborative triggering mechanism. The trend analysis module is used to perform synchronous reorganization on the relevant project group using the collaborative triggering mechanism to obtain the evolution sequence, perform curvature analysis on the evolution sequence to obtain the geological evolution trajectory, and extract inflection points from the geological evolution trajectory to generate a trend vector. The hotspot prediction module is used to perform layer-by-layer cross-validation between the trend vector and the geological feature group to form a confidence gradient field, perform path tracing on the confidence gradient field to obtain a predicted path bundle, and perform convergence analysis on the predicted path bundle to deduce exploration hotspot areas. The solution output module is used to perform spatial cluster analysis on the exploration hotspot areas to construct a resource allocation matrix, generate a project priority sequence based on the resource allocation matrix, and output a multi-project collaborative management solution according to the project priority sequence.

10. A computer device, characterized in that, It includes a memory and a processor; the memory is used to store a computer program; the processor is used to execute the computer program and, when executing the computer program, implement the method as described in any one of claims 1 to 8.